This is a follow-up to Small Models Decide, Code Does, LLMs Think.
In that post I argued that the model is only a third of the answer: small models decide, code does, and LLMs think, only when they have to. It was about taking model calls out of a browser agent. This time I wanted the same split in the most common LLM interface there is: a chat box.
Most chat messages don’t need a large language model. “What’s the Wi-Fi password?” has an answer, and it’s on a sticky note on your fridge. “Where did I park?” is a photo you took of a sign. Sending those to a 3 GB generative model is like asking a novelist to read you a phone number.
System 1 + 2 is that idea as a demo: a fast model decides, message by message, whether a memory already answers you or Gemma 4 is needed. Everything runs in the browser tab on WebGPU.
Try it: huggingface.co/spaces/shreyask/system-one-two
Three models, two routes
| Role | Model | Job |
|---|---|---|
| Memory | EmbeddingGemma 2 | Notes, photos and voice memos in one 768-d space |
| System One | Strands Decider 2B via open-jev | One typed decision per message: memory or large model |
| System Two | Gemma 4 E2B | Describes each dropped photo or memo once; answers everything else |
Each message is embedded and matched against the fridge. If the best match is weak, it goes straight to Gemma. Otherwise System One reads the message and the top three memories and answers one question: can a memory answer this message?
If it says yes with enough confidence, the reply is the memory itself: the note, the photo, the voice memo. Nothing is generated, so nothing can be hallucinated. If it says no, Gemma 4 writes the answer with those memories as context.
A meter counts how often Gemma was actually needed. Every reply shows the path it took, memory search, then System One, then memory or Gemma, with measured times.

Similarity isn’t a decision
This surprised me most, and it’s the reason System One exists.
In my first run, “Where did I park?” matched the parking photo with a cosine of 0.737. “Tell me a joke about parking garages” matched the same photo at 0.740. “What’s the Wi-Fi password?” and “How do I change the Wi-Fi password on my router?” both land on the Wi-Fi note, at 0.79 and 0.71.
Similarity tells you what a message is about. It doesn’t tell you whether a stored fact answers it. No threshold separates “where did I park” from “joke about parking”, because the difference isn’t in the topic, it’s in the intent. That judgment is exactly what a decision model is for.
Finding a System One that could tell
I wrote a probe before building any UI: 40 cases (20 lookups, 20 that need Gemma, many of them look-alike traps), plus 20 held-out cases I wrote without looking at scores. The bar: zero wrong memory answers. Showing a related-but-wrong memory is worse than waiting a few seconds for Gemma.
I started with GLiNER2.5-Decide, the 0.5B model that plays Chrome dino perfectly in Jevosaurus. Here it topped out. Across 15 wordings (different option descriptions, top-1 vs top-3 memories, an intent-only question, even a with/without-memories contrast) its best AUC was 0.93. At a cut-off with no wrong memory answers, it answers only 12 of 20 lookups.
Same question and cases on the other open decision models, each with its cut-off set for zero wrong memory answers:
| Model | Size | Lookups answered | Overall |
|---|---|---|---|
| GLiNER2.5-Decide | 0.5B | 12/20 | 32/40 |
| Decision 2.0 Kai | 0.6B | 14/20 | 34/40 |
| Decision 2.0 Sol | 2B | 17/20 | 37/40 |
| Strands Decider 2B | 2B | 18/20 | 38/40 |
Two small changes helped:
- Calling the options
memoryandlarge modelinstead ofmemoryandgemma. Decision models read labels, and “gemma” means nothing to them. - Naming lists, calculations and opinions in the option descriptions, so “What’s 15% of 11 dollars?” stops looking like a lookup.
One lesson carried over from Jevosaurus: decision models read words, not numbers. System One never sees a similarity score. It sees this:
Message: Where did I park?
Memory 1 (strong match, photo): This is a parking sign indicating a specific parking spot...
Memory 2 (strong match, photo): This is a baggage tag for a flight...
Memory 3 (strong match, voice memo): Renewed the car insurance by Friday...

The fridge
Memory is a fridge door. Notes are sticky notes, photos are polaroids, voice memos are little tape magnets. When a memory answers, it lifts off the fridge. When you drop a photo, Gemma 4 describes it once (“This is a boarding pass for an airplane flight… Seat 22A, Gate B7”). From then on it’s searchable and answerable without Gemma.

Still ask Gemma puts both systems side by side, each with its own timer:

The honest part
It isn’t a 100× speedup. A memory answer takes about 0.4–0.5 s in the browser: about 50 ms of search plus about 300 ms of System One. Gemma answers a short factual question in about 1.5 s, and an explanation in 6–7 s. The win is that the 3 GB model isn’t in the loop for questions you’ve already answered, and that the memory route also works on phones, where Gemma 4 can’t run at all. Phones get the 345 MB GLiNER2 build with a cautious cut-off.
Hosted Jev is better. I ran the same 60 cases, with the same state and question, against TypeSafe’s hosted Jev (jev-1.13.0):
| System One | Where it runs | Cases | Held out | Wrong memory answers | Latency |
|---|---|---|---|---|---|
| Strands Decider 2B | Your GPU (WebGPU) | 37/40 | 17/20 | 2 (held out) | ~300 ms |
| Jev (hosted) | TypeSafe’s API | 40/40 | 20/20 | 0 | 91 ms round trip |
Even over the network, Jev is faster. Its separation is clean: every lookup scored 0.68 or higher, and every trap 0.48 or lower.
So why does the demo use Strands? Because of what’s on the fridge: Wi-Fi passwords, gate codes, your landlord’s number, a photo of your boarding pass. In this demo none of that leaves the tab, not the photos, not the voice memos, not the questions. If your memories can go to an API, use Jev. If they can’t, a 2B open model on WebGPU gets you most of the way, and it’s free.
The two held-out misses are near-misses:
- “Is 3815 a good bike lock combination?” The note contains 3-8-1-5, but the question asks for an opinion.
- “What’s the deadline to ship onboarding v2?” The whiteboard names the task but gives no deadline.
When that happens you see the related memory with a one-click Still ask Gemma.
Things that bit me
- WebGPU kernels compile per input length. The first question took 4 s; the second, 50 ms. The embedder now warms up across typical question lengths at load, so your first question isn’t the one that pays.
- Vite only bundles a worker written as
new Worker(new URL("…", import.meta.url))in one expression. Build the URL elsewhere and the production build ships no workers. Dev works fine, so nothing warns you. I caught it in a preview build minutes before deploying. - Chrome’s
scrollIntoViewnow returns a Promise. An expression-bodieduseEffect(() => el.scrollIntoView())hands React a Promise as its cleanup, and the app crashes withdestroy is not a function. - Background tabs are throttled hard. A 25-second answer that looked like a bug was a hidden tab.
- Captions matter more than the router. For a plain “LEVEL P3 SPOT 42” card, Gemma wrote “a text graphic”, and System One rightly refused to treat it as the answer to “where did I park”. Asking Gemma to say what the photo is first fixed it.
- Test the phone path on purpose. A review found that on phones, any message that needed Gemma waited forever for a model that never loads. Now it says so.
Links
- Demo: System 1 + 2 on Hugging Face Spaces. Desktop Chrome with WebGPU gets the full experience: about 4.5 GB of models on first load, cached after. Ask “Where did I park?”, “What’s the Wi-Fi password?”, then “Explain how WebGPU works”, and watch the meter.
- Models: EmbeddingGemma 2, Strands Decider 2B, GLiNER2.5-Decide mobile, Gemma 4 E2B
- Built with: Transformers.js and open-jev
- Related: Small Models Decide, Code Does, LLMs Think and Jevosaurus
Fast thinking where it’s enough, slow thinking only where it’s needed. Don’t wake the big model.