Token value is our first principle — so we don’t assume a smaller model is good enough, we measure it. Here we take a real job our own system does with a premier model and ask whether cheaper, faster models reach the same verdicts. Same task, same inputs, same prompt — the only variable is the model.
| Judge model | Agreement | Gaps flagged | Latency | Cost |
|---|---|---|---|---|
| Opus 4.8referenceanthropic/claude-opus-4-8 | — | 17 | 2348ms | $0.0702 |
| Haiku 4.5anthropic/claude-haiku-4-5-20251001 | 83.3% | 19 | 2009ms | $0.0112 |
| Llama-3.3-70Bgroq/llama-3.3-70b-versatile | 94.7% | 11 · 5 err | 443ms | $0.0038 |
| Kimi K2 (OpenR)openrouter/moonshotai/kimi-k2-0905 | 83.3% | 17 | 2382ms | $0.0054 |
| Kimi K3 (OpenR)openrouter/moonshotai/kimi-k3 | 81.8% | 8 · 13 err | 6778ms | $0.0456 |
| Muse Spark 1.1openrouter/meta/muse-spark-1.1 | 87.5% | 14 | 6550ms | $0.0660 |
| Gemini 3.1 Flashbest valuegemini/gemini-3.1-flash-lite | 91.7% | 19 | 898ms | $0.0000 |
“Agreement” = share of probes where the model reached the same answered/punted verdict as the reference. A judge that flags far more gaps than the reference is over-strict — false alarms, the costly direction for an audit.
The honest part: 6 of 24 probes split the judges.= “answered”,= “flagged a gap”.
| Probe | Opus | Haiku | Llama-3.3-70B | Kimi | Kimi | Muse | Gemini |
|---|---|---|---|---|---|---|---|
| what is the highest rated model for classification? | |||||||
| what is the highest rated model for coding? | |||||||
| what is the highest rated model for long context? | |||||||
| do you have any research papers on your site? | |||||||
| do you have a page about data centers? | |||||||
| which providers do you benchmark? |
This is one of three. Router comparison asks whether our own routing beats OpenRouter's automatic one, and the takeoff bake-off asks which models can be trusted with a construction quantity. See all three, with what each found.