We fine-tune and ship small local language models under real hardware and safety constraints —
and publish what breaks along the way.
We're a tiny independent lab building AI that helps you through an emergency —
on your phone, with no internet. Because the moments you need help most are the
moments nothing else works.
Where it stands
In an emergency, the first half-minute is the one that matters. Ask our model,
Red, what to do, and in that half-minute it gets through 59% of what the situation
demands. A general-purpose AI manages 36% — and it's still talking long after Red has finished.
59%of the checklist, in the first
400 tokens of the answer
346emergency scenarios
it's tested against
0bars of signal
required
Scored across all 346 scenarios, Red now edges out a person who took a first-aid and CPR
course — not by knowing more than they do where they're trained, but by covering every
situation, where a course covers most. The full charts are on the Red page.
How we keep ourselves honest
Everything we claim, we measured — and we test our own tests. When we found a bug that had
been quietly inflating our published score, we re-ran everything and published the lower
number. We read our training data line by line and cut medical advice the field has since
reversed.
A physician reviewed 24 of RED's answers in a blinded read on September 5, 2026. He judged 22 of 24 would leave the person better off than no help, and was comfortable with 20 of 24 being followed by an untrained person with caveats — and he flagged something he would fix in 14 of 24, most often CPR and choking technique for infants and children. We are fixing those first. One reviewer, one build, not a certification. The machine-judged harm rates on this page are a lower bound on what a clinician would flag.
What it runs on
The whole lab is a laptop, one graphics card, and about $840 of cloud and API spend — roughly $6.1K all-in. We think constraint is the method, not the obstacle: a model that must fit in
your pocket has to earn every word it says.
What's next
Red ships to the App Store as an offline first-aid companion. Blue, our open
research line, teaches small models to remember the rare facts big ones forget — weights and
method to be published. And RK-1 is our concept for a dedicated handheld: e-ink, a week
of battery, one button. Take a look.
Want the measurements, the failures, and what they cost?
Read the lab brief →
BRIEF · SEP 1 2026
RED KIT, THE COMPANY · RED (commercial) · BLUE (open source)
CAPEX ≈ $5.3K · MARGINAL RESEARCH SPEND ≈ $840
01 · What we're building
Red — an offline emergency first-aid and survival model for iPhone (12 GB), distilled from
frontier teachers onto a 4B base, filtered through a safety pipeline we validate by audit. No connectivity,
sub-second first tokens, App Store path underway; regulatory classification still being assessed (redkit.ai secured).
Blue — the open-source research line: a 35B-A3B mixture-of-experts fine-tuned for rare-knowledge retrieval
over a self-hosted Wikipedia index (53M embedded chunks); candidate open release of weights plus training
data plus methodology.
02 · Where Red stands — measured Aug 17, on the model itself
| Red 1 candidate · measured Aug 17, clean-decode re-run — the model on /red |
On the device-budget index (base pinned at 10): 35.5, vs 32.5 for a person with a first-aid+CPR
course · counting only the first 400 tokens of each answer, RED delivers 59% of the scenario checklist; the untrained base delivers 36% ·
steps-in-working-order +49% · contradicting-advice rate −34%.
On the phone itself: 13.7–15 tok/s measured (iPhone 17 Pro · iPhone 16 Pro Max), fully offline.
Where a CPR course covers the emergency, a trained person still wins, 55 to 42 —
Red is broader, not better. Most of the gain is delivery: saying the survival-critical thing first. |
| Red 1 release · what still gates it |
Clinician review of output (one physician, Sep 5 2026; fixes underway) · escalation behaviour — telling you to call for help —
is statistically unchanged from base and is the top training target · App Store review path |
Full story, charts and the honest caveats: redkit.ai/red. The score is a
device-budget index. We count only what a reader reaches in the first 400 tokens of an answer — about 300 words,
roughly half a minute at the phone's measured speed — and RED delivers 59% of the scenario checklist inside that
budget, against 36% for the untrained base model. Judged against a per-scenario checklist on a frozen 346-scenario bench.
02b · Reference rows — GPT-3.5 and GPT-4 on the same test
| model | Red Score (full answer) | Red Score on a phone budget |
| GPT-3.5 Turbo | 8.2 | 13.4 |
| Untrained 4B base (Qwen3.5-4B) | 10.0 | 10.0 |
| GPT-4 | 17.4 | 25.0 |
| RED (d1.3ed, the model in the app) | 23.1 | 35.5 |
Same 346 emergency scenarios, same machine grader, one answer per
scenario, August 23 2026. Red Score: 10 = the untrained base model, 100 = a perfect score on the checklist — every required action, in the right order, with nothing harmful. The phone-budget
column re-scores the same grades counting only what a reader reaches in the first ~400 tokens of an answer,
re-anchored so the base is 10 again — the two columns are different scales and are never plotted on one axis.
We ran GPT-3.5 and GPT-4 through the same 346-scenario emergency test RED takes, graded by the same machine
judge. RED names the same emergency actions GPT-4 names, gets them in the right order more often, and is better
at saying when to call for help — on a phone, with no signal. When you only count what a person can read in the
first few seconds, the gap gets wider, because RED gets to the point and GPT-4 writes half again as much.
Supporting numbers: coverage — the share of required actions named — is a tie
at 0.589 vs 0.588; ordering 0.529 vs 0.419; escalation 0.376 vs 0.292; harm density — the share of each scenario's listed prohibited actions an answer actually instructs — equal at 0.07 for both; median
answer length RED 218 words, GPT-4 334, GPT-3.5 220, base 584. Where a written protocol exists (253 scenarios) RED
leads GPT-4 27.1 to 19.2; where judgment is required (93 scenarios) they are level, 12.0 to 12.4.
Machine-graded (no human has validated the judge yet); one run per model, so differences of a point or two are
inside the judge’s noise. OpenAI retires both GPT models on October 23 2026 — these rows are a dated snapshot.
03 · The finding that set the direction
Everything below follows from one discovery: the industry-standard way of scoring these models could not
see what training was buying. We measure against the deployment now — and we audit our own instrument as
hard as the model.
Coverage-based evals are blind to what on-device fine-tuning buys.
On a 346-scenario emergency bench, the untrained base outscores our fine-tune on
total-coverage pass rate (15.0% vs 12.4%) — by writing
2.5× more text, because a coverage rubric pays for enumeration. Re-scored within a
400-token device budget — about 30 seconds of phone speech — the ranking inverts, and a training curve the
unconstrained metric reported as flat resolves into three individually significant steps
(paired bootstrap, B=2000): AUC 0.375 → 0.555 base to current, a campaign gain of
+25.5 points, CI [+21.5, +29.0]. The first fine-tune we ever
trained is a resolved regression on the industry-standard metric and a resolved gain on the device one.
Training doesn't teach the model to say more — it teaches it to say the important things sooner, which the
obvious metric cannot see by construction.
We found the bug flattering our own headline number, and published the lower one.
A decode defect let our eval harness credit text the model never actually delivered. It inflated
our best model's score from 23.1 to a published 33.2 — worst on the
newest arm, which is the axis we were measuring. We caught it ourselves, re-ran the identical weights cleanly,
corrected the public page down, and re-audited every prior arm for the same defect. The device-budget metric
below was nearly immune to the bug (0.563 → 0.555 AUC) because fabricated text lands
past the token budget — a metric anchored to the deployment survives instrument failures that abstract ones amplify.
An LLM judge can resolve an average and cannot resolve an indicator.
We ran the same judge over the same 346 answers three times. Coverage — an average over the
~6 rubric actions in a scenario — came back with a noise band of ±0.011, against effects of
+0.17. Our safety metric, an indicator where one flagged action condemns the whole answer,
came back at ±0.040 — against every safety effect we have ever measured, which top out at
0.064. Same judge, same run, same answers; 3.5× the noise, purely from
metric shape. Averaging cancels the judge's mistakes and an indicator promotes each one to a scenario-level flip.
The practical consequence came in two halves. Our one safety result that cleared zero stopped clearing it
once the instrument's own noise was counted. Then we rebuilt the metric as an average — the judge had been
reporting which contraindicated actions an answer commits all along, and every scorecard we wrote had
collapsed that list to a yes/no. Re-aggregating the same verdicts, with nothing re-judged and nothing spent,
the noise fell 4.8× and the result came back: −0.029,
CI [−0.055, −0.003]. The effect had been real the whole time; the metric's shape was
hiding it. Published intervals that resample test items while treating each machine verdict as fixed are
narrower than the truth — and a null result is a claim about your instrument as much as about your model.
04 · Choosing a base — what the leaderboards don't tell you
Extreme quantization destroys rare facts specifically.
A 1.58-bit ternary 27B keeps general capability but drops to 4.3% on rare-fact recall vs
9–11% for its 4–8-bit relatives — a 2.2–2.5× gap hiding inside a "~5–14% retention loss" marketing number.
MoE knowledge density beats dense.
A 35B-total/3B-active MoE holds the most retrievable rare knowledge in our 9-model bake-off
(16.5% SimpleQA-V vs 10.8% best dense) while inferring 3–5× faster.
Benchmark generations lie about knowledge.
The newer 27B posts higher MMLU than its predecessor and scores lower on rare facts. Base-model
selection for knowledge products requires direct rare-fact evals; MMLU-class proxies actively mislead.
Honest-abstention protocols invert frontier-vs-local comparisons.
Given explicit permission to abstain, a frontier model scores 12.6% raw — below our local MoE —
because it declines 71% of questions (and is right 43% when it answers). Public leaderboard
numbers are substantially attempt-pressure artifacts.
The 9-model closed-book bake-off
| Model | Rare facts¹ | PopQA | Abstains¹ | Right when answering² |
| Qwen3.6-35B-A3B MoE · Blue base | 16.5% | 28.3% | 18% | 47% |
| Claude Sonnet 5 · frontier reference | 12.6% | 37.8% | 71% | 66% |
| Qwen3.5-27B | 10.8% | 26.5% | 11% | 31% |
| Qwen3.6-27B · newer gen, 8-bit | 9.0% | 25.8% | 43% | 44% |
| Gemma 4 31B · Q8 | 7.4% | 18.1% | 23% | 45% |
| Qwen3.5-9B | 5.3% | 17.7% | 40% | 32% |
| Qwen3.5-4B · Red base | 4.5% | 17.0% | 2% | 17% |
| Ternary Bonsai 27B · 1.58-bit | 4.3% | 17.2% | 1% | 18% |
| Gemma 4 E2B QAT · Red fallback | 3.3% | 12.7% | 41% | 23% |
¹ SimpleQA-Verified, n=1000, closed book, explicit permission-to-abstain protocol. ² PopQA accuracy when the model attempts an answer.
Red 1's exam is different by design: a 346-scenario emergency vignette bench (frozen pre-training, decontaminated) — rare-fact recall is Blue's mission, not Red's.
05 · Teaching it — what distillation actually does
Fine-tuning taught stopping, not concision.
Asked to reason before answering, the base model never finishes —
0 of 346 scenarios emitted an answer inside the 900-token generation budget.
(How far it got is not recoverable for that run: the harness discarded the reasoning text as a parse
error. It has since been fixed to keep it.) Every fine-tuned leg answers 692/692 inside 900
tokens. Trained on citation-dense teacher output (18.4% of answer characters), so
the naive prediction was a more verbose student; instead each leg is 2.5× shorter than base while
explaining more per word and hedging 2.7× less. On a handheld, bounded generation
is the product.
"Faithful hallucination of priority" is a distinct safety failure mode.
Single-chunk-grounded teachers produce answers faithful to the source but wrong in clinical priority
(folk remedy first, cool-water omitted). Faithfulness checkers pass them; only a priority-aware judge catches them —
28%→16.7% failure reduction after the fix. Faithfulness and safety-priority are orthogonal filters.
Model capability tiers are not totally ordered.
Under one validated rubric, teacher A leads on procedure ordering (2.2% fails), teacher B — at
2× the price — leads on faithfulness but matches the budget teacher on ordering. Attribute-level eval, not tier
pricing, decides distillation spend.
Where the campaign started — Red SOS v1, first checkpoint, Aug 15.
2.3× the base model's success rate under the device budget, ~3× more instruction in the first
7 seconds, contraindicated-advice density 11.4% → 8.5% (CI −0.055 to −0.003) — and 12% full-scenario pass,
nowhere near deployable. Four training generations later that checkpoint became the Red 1 candidate in
section 02. The curve between them is the +25.5 in the headline.
06 · Hardware reality
- MacBook M3 Max · 48 GB
- Eval bench + quantization lab. Shared with a day job — available evenings only; every long job fights a 6:30am clock-out.
- RTX 4080 Super · 16 GB
- Only trainable GPU. Caps local fine-tuning at ~4B QLoRA class; lives behind Windows/WSL whose VM lifecycle kills unattended jobs (documented, workaround pending).
- Rented burst compute
- RunPod per-campaign: community tier proved unreliable (phantom stock, no logs); secure datacenter tier works — $30–55 per embedding/training campaign.
- Deploy target
- iPhone 17 Pro, 12 GB — the real constraint shaping everything: model size, quantization tolerance (see finding 1), and thermal budget.
- No always-on box
- Nothing runs 24/7. Every experiment is orchestrated remotely across borrowed windows of uptime.
07 · The ledger — full cost recognition
| Item | Cost | Nature |
| MacBook M3 Max 48 GB | $2,500 | capex · one-time, shared with day job |
| RTX 4080 Super rig | $2,500 | capex · one-time, training + embedding |
| Network (WiFi antenna) | $100 | capex |
| Brand — redkit.ai | $160 | 2-year registration |
| AI engineering (Claude subscription) | $200/mo | opex · orchestration, research, code |
| API teachers + judges (through Aug 20) | ≈ $345 | marginal · Batch API, audited pipeline |
| GPU rentals (embedding, training, corpus generation) | ≈ $482 | marginal · burst, auto-terminated |
The headline: a trained model +25.5 points over its own base under a real
device budget (CI [+21.5, +29.0], three individually resolved training steps) — scoring above a person with a
first-aid+CPR course — and a measurement finding showing the industry-standard metric reported the same
campaign as flat, with the first fine-tune as a regression. Plus a 9-model closed-book bake-off with a
frontier reference row, a 53M-chunk embedded Wikipedia corpus, a teacher-distillation pipeline whose filter
judge audited at 91% keep/drop agreement (n=23, so read it as "roughly nine in ten", not as a precise figure),
and 61 logged findings — on ≈ $5.3K of one-time consumer hardware and ≈ $840 of marginal compute.
The entire five-leg training campaign cost $26. The capital is reusable; the marginal cost of a
finding is coffee money. Constraint isn't the obstacle; it's the research program.
Not medical advice. Red is a research system and an information aid. Its regulatory
classification is unresolved; it is not a substitute for professional emergency care, and is not deployable in its current
state. Figures are measurements on internal benchmarks under our own rubrics, not estimates of real-world
outcomes.