← Red Kit RED · MODEL LINE · AUG 2026
Red

The help you need is offline, in your pocket, before the ambulance arrives.

Red is a first-aid and survival model small enough to run on a phone with no signal, no server, and no account. It is built for the worst ten minutes of someone's life — the ones that happen before help gets there.


Why offline is the whole product

Emergencies are not evenly distributed across cell coverage. They happen in basements and canyons, on trails and back roads, during storms that took the towers down, in the exact places a cloud assistant is a blank screen.

Everything Red knows sits on the device. It works in airplane mode, in a dead zone, on a dead network — and nothing you ask it ever leaves your hand.

What we optimise for: the first hundred words

Nobody reads to the end of a long answer while someone is bleeding. What matters is how much correct, usable instruction arrives in the first few seconds.

So that is what we measure. Every answer is scored against the steps a clinician rubric says that scenario requires, at each length budget — how much did it get right by 100 words, by 200, by 400.

100 words
9.4%
27.5%
200 words
18.5%
44.1%
400 words
36.2%
53.5%
900 words
53.7%
53.7%
the model we started from Red, current build
Both end up in the same place. Red gets there in under half the words — and delivers nearly three times as much correct instruction in the first hundred.
2.9×more correct instruction in the first 100 words
+35%across the whole length-budget curve
692/692 answers finished inside budget. Asked to reason first, the base model finished none of them

Measured on a frozen bench of 346 emergency scenarios, each with a written rubric of the actions that scenario requires.

Three lines moving the right way

Across four successive versions of the training data — same model, same test, same settings, only the data changing — two of these improved at every single step.

Advice that contradicts the situation↓ 20% better
.292 .234
fell at every single steplower is better
Padding — words that are not instructions↓ 27% better
.117 .085
fell at every single steplower is better
Getting the steps in the right order↑ 21% better
.332 .402
rose overall, though not at every stephigher is better

Order matters more than it sounds. Stopping heavy bleeding before you check anything else is not a stylistic preference — it is the difference between the steps working and not. Our best result on ordering, from a later build, reaches .454.

The versions, and what each one fixed

Every version starts from the same public base model and is graded on the same frozen test. Only the training material and the training settings change — which is what makes the comparison mean anything.

d1.012,810 examples

Teach it the job

The first real training run, on emergency and field-medicine material. It established the core win immediately: the trained model puts the important step first instead of building up to it.

d1.1references trimmed

Stop reciting, start instructing

The model was spending its very limited output quoting long source references back at the user. We cut them by 77% of their characters. The answers got shorter and carried more actual instruction.

d1.213,128 examples testing

Audit the textbook

We put every one of 12,713 training answers under a machine reader looking for advice that was wrong for the situation, with a stronger model confirming each flag. 533 came out. Then we wrote 884 new examples for children and severe allergic reactions. Most teams grade the model. We grade what the model was taught.

d1.3three variants training now

Teach it to call for help

Almost all first-aid literature is written for people who are on their own — field medics, expedition crews. Most real emergencies are not like that: there is usually a phone, and using it is the single highest-value thing a bystander does. We measured that gap and built for it. Three versions are training right now — one teaching the reflex directly, one rebalancing the source material, one doing both — so we learn which lever actually moves it.

Four things we do that most teams skip

That is the argument for a small lab building this. Nobody is going to audit emergency-medicine training data line by line at scale. We will.

Status

Red is in development, App Store path underway. Every figure on this page comes from a frozen bench of 346 scenarios with written clinical rubrics, and none of it is a substitute for calling emergency services — Red is built for the minutes before they arrive, and for the places they can't be reached.