This is a follow-up to Small Models Decide, Code Does, LLMs Think.

In that post I argued that the model is only a third of the answer: small models decide, code does, and LLMs think, only when they have to. It was about taking model calls out of a browser agent. This time I wanted the same split in the most common LLM interface there is: a chat box.

Most chat messages don’t need a large language model. “What’s the Wi-Fi password?” has an answer, and it’s on a sticky note on your fridge. “Where did I park?” is a photo you took of a sign. Sending those to a 3 GB generative model is like asking a novelist to read you a phone number.

System 1 + 2 is that idea as a demo: a fast model decides, message by message, whether a memory already answers you or Gemma 4 is needed. Everything runs in the browser tab on WebGPU.

Try it: huggingface.co/spaces/shreyask/system-one-two

Three models, two routes

How a message is routed: memory search, System One, then the memory itself or Gemma 4

RoleModelJob
MemoryEmbeddingGemma 2Notes, photos and voice memos in one 768-d space
System OneStrands Decider 2B via open-jevOne typed decision per message: memory or large model
System TwoGemma 4 E2BDescribes each dropped photo or memo once; answers everything else

Each message is embedded and matched against the fridge. If the best match is weak, it goes straight to Gemma. Otherwise System One reads the message and the top three memories and answers one question: can a memory answer this message?

If it says yes with enough confidence, the reply is the memory itself: the note, the photo, the voice memo. Nothing is generated, so nothing can be hallucinated. If it says no, Gemma 4 writes the answer with those memories as context.

A meter counts how often Gemma was actually needed. Every reply shows the path it took, memory search, then System One, then memory or Gemma, with measured times.

Where did I park? answered from a photo on the fridge, no LLM call

Similarity isn’t a decision

This surprised me most, and it’s the reason System One exists.

In my first run, “Where did I park?” matched the parking photo with a cosine of 0.737. “Tell me a joke about parking garages” matched the same photo at 0.740. “What’s the Wi-Fi password?” and “How do I change the Wi-Fi password on my router?” both land on the Wi-Fi note, at 0.79 and 0.71.

Similarity tells you what a message is about. It doesn’t tell you whether a stored fact answers it. No threshold separates “where did I park” from “joke about parking”, because the difference isn’t in the topic, it’s in the intent. That judgment is exactly what a decision model is for.

Finding a System One that could tell

I wrote a probe before building any UI: 40 cases (20 lookups, 20 that need Gemma, many of them look-alike traps), plus 20 held-out cases I wrote without looking at scores. The bar: zero wrong memory answers. Showing a related-but-wrong memory is worse than waiting a few seconds for Gemma.

I started with GLiNER2.5-Decide, the 0.5B model that plays Chrome dino perfectly in Jevosaurus. Here it topped out. Across 15 wordings (different option descriptions, top-1 vs top-3 memories, an intent-only question, even a with/without-memories contrast) its best AUC was 0.93. At a cut-off with no wrong memory answers, it answers only 12 of 20 lookups.

Same question and cases on the other open decision models, each with its cut-off set for zero wrong memory answers:

ModelSizeLookups answeredOverall
GLiNER2.5-Decide0.5B12/2032/40
Decision 2.0 Kai0.6B14/2034/40
Decision 2.0 Sol2B17/2037/40
Strands Decider 2B2B18/2038/40

Two small changes helped:

  • Calling the options memory and large model instead of memory and gemma. Decision models read labels, and “gemma” means nothing to them.
  • Naming lists, calculations and opinions in the option descriptions, so “What’s 15% of 11 dollars?” stops looking like a lookup.

One lesson carried over from Jevosaurus: decision models read words, not numbers. System One never sees a similarity score. It sees this:

Message: Where did I park?
Memory 1 (strong match, photo): This is a parking sign indicating a specific parking spot...
Memory 2 (strong match, photo): This is a baggage tag for a flight...
Memory 3 (strong match, voice memo): Renewed the car insurance by Friday...

How it decided: the exact state System One read

The fridge

Memory is a fridge door. Notes are sticky notes, photos are polaroids, voice memos are little tape magnets. When a memory answers, it lifts off the fridge. When you drop a photo, Gemma 4 describes it once (“This is a boarding pass for an airplane flight… Seat 22A, Gate B7”). From then on it’s searchable and answerable without Gemma.

A dropped boarding pass, answered from memory

Still ask Gemma puts both systems side by side, each with its own timer:

System 1 vs System 2, side by side

The honest part

It isn’t a 100× speedup. A memory answer takes about 0.4–0.5 s in the browser: about 50 ms of search plus about 300 ms of System One. Gemma answers a short factual question in about 1.5 s, and an explanation in 6–7 s. The win is that the 3 GB model isn’t in the loop for questions you’ve already answered, and that the memory route also works on phones, where Gemma 4 can’t run at all. Phones get the 345 MB GLiNER2 build with a cautious cut-off.

Hosted Jev is better. I ran the same 60 cases, with the same state and question, against TypeSafe’s hosted Jev (jev-1.13.0):

System OneWhere it runsCasesHeld outWrong memory answersLatency
Strands Decider 2BYour GPU (WebGPU)37/4017/202 (held out)~300 ms
Jev (hosted)TypeSafe’s API40/4020/20091 ms round trip

Even over the network, Jev is faster. Its separation is clean: every lookup scored 0.68 or higher, and every trap 0.48 or lower.

So why does the demo use Strands? Because of what’s on the fridge: Wi-Fi passwords, gate codes, your landlord’s number, a photo of your boarding pass. In this demo none of that leaves the tab, not the photos, not the voice memos, not the questions. If your memories can go to an API, use Jev. If they can’t, a 2B open model on WebGPU gets you most of the way, and it’s free.

The two held-out misses are near-misses:

  • “Is 3815 a good bike lock combination?” The note contains 3-8-1-5, but the question asks for an opinion.
  • “What’s the deadline to ship onboarding v2?” The whiteboard names the task but gives no deadline.

When that happens you see the related memory with a one-click Still ask Gemma.

Things that bit me

  • WebGPU kernels compile per input length. The first question took 4 s; the second, 50 ms. The embedder now warms up across typical question lengths at load, so your first question isn’t the one that pays.
  • Vite only bundles a worker written as new Worker(new URL("…", import.meta.url)) in one expression. Build the URL elsewhere and the production build ships no workers. Dev works fine, so nothing warns you. I caught it in a preview build minutes before deploying.
  • Chrome’s scrollIntoView now returns a Promise. An expression-bodied useEffect(() => el.scrollIntoView()) hands React a Promise as its cleanup, and the app crashes with destroy is not a function.
  • Background tabs are throttled hard. A 25-second answer that looked like a bug was a hidden tab.
  • Captions matter more than the router. For a plain “LEVEL P3 SPOT 42” card, Gemma wrote “a text graphic”, and System One rightly refused to treat it as the answer to “where did I park”. Asking Gemma to say what the photo is first fixed it.
  • Test the phone path on purpose. A review found that on phones, any message that needed Gemma waited forever for a model that never loads. Now it says so.

Fast thinking where it’s enough, slow thinking only where it’s needed. Don’t wake the big model.