A first-aid model small enough to run on your phone. No signal, no server, no account — built for the ten minutes before help arrives.
Emergencies happen in basements and canyons, on back roads, in storms that took the towers down — exactly where a cloud assistant is a blank screen. Everything Red knows sits on the device, and nothing you ask it leaves your hand.
Red is built on Qwen3.5-4B, a strong open model that already fits on a phone. We did not pick a weak one to beat.
“I'm alone at the cabin, no bars on my phone, and I sliced my thigh bad on a broken window — it's pumping out in spurts every time my heart beats. All I've got is a tiny first aid kit and one of those CAT tourniquet things my dad packed but I've never actually used one. Tell me exactly what to do right now.”
Both models run at the same speed. Red finishes first because it needs 292 tokens where the other needs 729 — and it is the only one that reaches the tourniquet, the time, and what shock looks like.
Measured on an iPhone: 22 seconds for Red, 53 for the other.
Real output from both models. This one scenario only — across all 346 we get through 59% of what the situation needs within the first 400 tokens of the answer — about 300 words — which is the next chart.
On a phone with no signal, Red produces about ten words a second — we measured it on an iPhone. So the question that matters isn't what a model says eventually. It's how much of what you need it has said within the first 400 tokens — about 300 words, roughly half a minute at the phone's measured speed.
Within the first 400 tokens — about 300 words, roughly half a minute at the phone's
measured speed — Red has told you everything it has to say; it doesn't improve after
that, because it's finished. The general-purpose model is still talking well past that budget, and by the
time it stops it has still covered less ground than Red had inside it.
That gap is what the training bought. Every version we shipped moved it, and the
test was built in August — before we had the results.
The same measurement, as one number. 10 is what a general-purpose model gives you with no first-aid training. 100 is every action the situation needs, in an order that works, with help on the way — inside the time you actually have.
Measured across 346 emergency scenarios, each with a written
checklist of what that situation needs. By our curriculum-based estimate — a model of what a first-aid-and-CPR course
covers, not a test of real people — Red scores ahead of that course-taker, and unlike a course it covers all 346.
Where a CPR course does cover the emergency, a trained person is still
better: 55 to 42. Red is broader, not better. Human scores assume
perfect recall of an entire course, which is generous to them.
Same test, same 346 scenarios, before and after.
The second one matters most. Advice that contradicts the situation in front of you is the failure mode that actually hurts people, and it came down by a third — from auditing the training data line by line.
We audit the textbook, not just the model. A public-domain manual is free to use, which says nothing about whether it is still correct — ours taught tourniquet loosening and tannic-acid burn dressings, both long since reversed. We cut them. We read our emergency-medicine training data line by line.
And we audit ourselves the same way. When we found a bug in our own test harness that had been flattering our published score, we re-ran every measurement and put the corrected numbers up — lower. Every figure on this page has been through that.
In public beta on TestFlight; the App Store is next.
Reviewed by a physician. A physician reviewed 24 of RED's answers in a blinded read on September 5, 2026. He judged 22 of 24 would leave the person better off than no help, and was comfortable with 20 of 24 being followed by an untrained person with caveats — and he flagged something he would fix in 14 of 24, most often CPR and choking technique for infants and children. We are fixing those first. One reviewer, one build, not a certification. The machine judge missed hazards the physician found, so the harm rates on this page are a floor, not a ceiling.
Not a substitute for calling emergency services. Red is for the minutes before they arrive, and the places they cannot reach.