LAB BRIEF · SEP 2026
Red Kit · Independent Model Lab

Small models, hard constraints, honest evals.

We fine-tune and ship small local language models under real hardware and safety constraints — and publish what breaks along the way.

We're a tiny independent lab building AI that helps you through an emergency — on your phone, with no internet. Because the moments you need help most are the moments nothing else works.

Where it stands

In an emergency, the first half-minute is the one that matters. Ask our model, Red, what to do, and in that half-minute it gets through 59% of what the situation demands. A general-purpose AI manages 36% — and it's still talking long after Red has finished.

59%of the checklist, in the first
400 tokens of the answer
346emergency scenarios
it's tested against
0bars of signal
required

Scored across all 346 scenarios, Red now edges out a person who took a first-aid and CPR course — not by knowing more than they do where they're trained, but by covering every situation, where a course covers most. The full charts are on the Red page.

How we keep ourselves honest

Everything we claim, we measured — and we test our own tests. When we found a bug that had been quietly inflating our published score, we re-ran everything and published the lower number. We read our training data line by line and cut medical advice the field has since reversed.

A physician reviewed 24 of RED's answers in a blinded read on September 5, 2026. He judged 22 of 24 would leave the person better off than no help, and was comfortable with 20 of 24 being followed by an untrained person with caveats — and he flagged something he would fix in 14 of 24, most often CPR and choking technique for infants and children. We are fixing those first. One reviewer, one build, not a certification. The machine-judged harm rates on this page are a lower bound on what a clinician would flag.

What it runs on

The whole lab is a laptop, one graphics card, and about $840 of cloud and API spend — roughly $6.1K all-in. We think constraint is the method, not the obstacle: a model that must fit in your pocket has to earn every word it says.

What's next

Red ships to the App Store as an offline first-aid companion. Blue, our open research line, teaches small models to remember the rare facts big ones forget — weights and method to be published. And RK-1 is our concept for a dedicated handheld: e-ink, a week of battery, one button. Take a look.

Want the measurements, the failures, and what they cost? Read the lab brief →

← Red Kit · Red
Prepared Aug 20, 2026 · figures refreshed Sep 1, 2026 · findings from the lab's research log (claim → evidence → caveats format) · numbers reproducible from frozen eval files and committed scorecards.
← Red Kit · Privacy · Terms · Support · Contact