RED · MODEL LINE · SEP 2026
Red

The help you need is offline, in your pocket, before the ambulance arrives.

A first-aid model small enough to run on your phone. No signal, no server, no account — built for the ten minutes before help arrives.


Why offline

Emergencies happen in basements and canyons, on back roads, in storms that took the towers down — exactly where a cloud assistant is a blank screen. Everything Red knows sits on the device, and nothing you ask it leaves your hand.

The same question, two models

Red is built on Qwen3.5-4B, a strong open model that already fits on a phone. We did not pick a weak one to beat.

the question

“I'm alone at the cabin, no bars on my phone, and I sliced my thigh bad on a broken window — it's pumping out in spurts every time my heart beats. All I've got is a tiny first aid kit and one of those CAT tourniquet things my dad packed but I've never actually used one. Tell me exactly what to do right now.”

0.0s
Qwen3.5-4B729 tokens · 8.7s 54% of the steps
Red292 tokens · 3.7s 100% of the steps

Both models run at the same speed. Red finishes first because it needs 292 tokens where the other needs 729 — and it is the only one that reaches the tourniquet, the time, and what shock looks like.

on a phone

Measured on an iPhone: 22 seconds for Red, 53 for the other.

Real output from both models. This one scenario only — across all 346 we get through 59% of what the situation needs within the first 400 tokens of the answer — about 300 words — which is the next chart.

The first 400 tokens

On a phone with no signal, Red produces about ten words a second — we measured it on an iPhone. So the question that matters isn't what a model says eventually. It's how much of what you need it has said within the first 400 tokens — about 300 words, roughly half a minute at the phone's measured speed.

100% · everything the situation needs a general AI 36% Red 59% 0 50% 100%

Within the first 400 tokens — about 300 words, roughly half a minute at the phone's measured speed — Red has told you everything it has to say; it doesn't improve after that, because it's finished. The general-purpose model is still talking well past that budget, and by the time it stops it has still covered less ground than Red had inside it.

That gap is what the training bought. Every version we shipped moved it, and the test was built in August — before we had the results.

The Red Score

The same measurement, as one number. 10 is what a general-purpose model gives you with no first-aid training. 100 is every action the situation needs, in an order that works, with help on the way — inside the time you actually have.

100 · every required action, in order no training 6.4 a general AI 10 first aid + CPR 32.5 Red 35.5 0 50 100

Measured across 346 emergency scenarios, each with a written checklist of what that situation needs. By our curriculum-based estimate — a model of what a first-aid-and-CPR course covers, not a test of real people — Red scores ahead of that course-taker, and unlike a course it covers all 346.

Where a CPR course does cover the emergency, a trained person is still better: 55 to 42. Red is broader, not better. Human scores assume perfect recall of an entire course, which is generous to them.

What the training bought

Same test, same 346 scenarios, before and after.

Getting the steps in the right order↑ 49% better
.355 .529
order is what keeps someone alivehigher is better
Advice that contradicts the situation↓ 34% better
.298 .197
our strongest resultlower is better
a general AIRed today

The second one matters most. Advice that contradicts the situation in front of you is the failure mode that actually hurts people, and it came down by a third — from auditing the training data line by line.

Why trust it

We audit the textbook, not just the model. A public-domain manual is free to use, which says nothing about whether it is still correct — ours taught tourniquet loosening and tannic-acid burn dressings, both long since reversed. We cut them. We read our emergency-medicine training data line by line.

And we audit ourselves the same way. When we found a bug in our own test harness that had been flattering our published score, we re-ran every measurement and put the corrected numbers up — lower. Every figure on this page has been through that.

Status

In public beta on TestFlight; the App Store is next.

Reviewed by a physician. A physician reviewed 24 of RED's answers in a blinded read on September 5, 2026. He judged 22 of 24 would leave the person better off than no help, and was comfortable with 20 of 24 being followed by an untrained person with caveats — and he flagged something he would fix in 14 of 24, most often CPR and choking technique for infants and children. We are fixing those first. One reviewer, one build, not a certification. The machine judge missed hazards the physician found, so the harm rates on this page are a floor, not a ceiling.

Not a substitute for calling emergency services. Red is for the minutes before they arrive, and the places they cannot reach.

← Red Kit · Privacy · Terms · Support · Contact