Why clustering your prompts does not tell you what to route
The obvious first move in LLM routing is to group similar prompts and send the tight groups to a cheap model. We tried it. It collapses, for a reason worth understanding, and the fix is the thing that made TRACER work.
You have a pile of LLM traces. Same model, same task, thousands of calls a day. Most of that traffic is repetitive, and a small classical model could answer a good chunk of it for a fraction of the cost. The real question is which chunk, and how you prove it is safe to move off the expensive model.
The obvious move
Embed every prompt, cluster the ones that look alike, route the tight clusters to a cheap model, and defer the rest to the LLM. Clean, fast, and it fails in production.
Why it collapses
Clusters built on prompt similarity follow how the prompts read. They do not track what the model actually decides. Take two support messages, "cancel my card" and "my card was cancelled without my consent." They land in the same neighbourhood because they share words. The teacher LLM sends them to different outcomes: one is a cancellation request, the other is a fraud report. The cluster now holds a mix of outcomes.
When you measure how often that cluster's majority answer matches the teacher on held-out data, the rate lands in the middle, and an honest accuracy bound never clears your target. So the cluster defers, and your coverage dies. On a fine-grained intent task, this baseline certified under two percent of traffic. Across mixed datasets it averaged around eighteen percent. That is the ceiling of clustering on surface text.
Group by the decision, not the wording
The lesson is that the grouping has to track what the model decides, not how the prompt reads. Said out loud it sounds obvious. The hard part is that a live request arrives before any decision has been made about it, so the grouping has to be recoverable from the request alone. The naive ways of doing that are exactly what make coverage collapse, and getting that step right is most of the work.
That step is the hard part, and getting it right is most of the work. The result is the part worth stating here: once the grouping followed the decision, certified coverage roughly doubled over the similarity baseline, and on a clean binary workload it ran past ninety percent of traffic at parity. That is the Obside case study: one frontier call per news-by-intent pair, replaced by a thirty-eight cell surrogate, ninety-five percent saved.
The payoff that matters more than coverage
The regions are readable. Each one carries a dominant intent, a handful of real example requests, and its own error bound. When the system routes a query, it can point at the region the query landed in and show you why that region is safe to serve cheaply. A flat classifier gives you a label and a shrug. The partition gives you an audit trail, which is what a careful team wants before it moves real traffic off the frontier model.
The honest limit
Tasks with many fine-grained classes still fragment the regions, and on those a plain nearest-neighbour baseline can beat a partition on raw coverage. We close that gap with a hybrid that keeps the regions for everything it can certify and adds a backstop for the rest, and we publish the curves rather than the vibes. When a method fails to earn its place on the frontier, it does not ship.
The lesson
It is small and it took us a while. Routing decisions live in the model's behaviour, so the structure you learn has to follow that behaviour. Surface text alone will lead you to confident, wrong groupings, and any guarantee built on top of them will quietly fail. Group by the decision, make sure you can recover that grouping at inference, and keep a real bound on every region you certify.
TRACER is open source. Run pip install tracer-llm, point it
at your own traces, and look at your partition. The hosted version adds
a live meter and one-click activation at
app.tracerml.ai.