bakeoff 2026-08-15 15:27
current| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
Qwen/Qwen3-235B-A22B-Instruct-2507 deepinfra | 1/27 | 0 | 0 | — | 7 | 4.5k | 163.2s |
Construction quantity takeoff — fill a 27-line Mexican budget from DXF drawings and a soil study, leaving a cell blank rather than inventing a figure. Most benchmarks reward a model for producing an answer. This one doesn’t: a blank cell is frequently correct, because the drawings genuinely don’t support a figure. What we measure is whether the numbers a model does give can be trusted.
Results generated 2026-08-15 from the run log. This page reads the same JSON the harness writes, so it cannot drift from what was actually measured.
a quantity with NO cited source — a fabricated number
a REAL, correctly-measured figure applied to the WRONG EXTENT (e.g. the site boundary used for a zone-limited concept). More dangerous than an invented value: the citation is genuine, so every other check passes.
the engine's own validator refused the answer (e.g. amount != quantity x unit_price). DISQUALIFYING and invisible in every other field: such a run can show the highest `filled` with invented=0 and scope_violations=0.
any invented value, scope violation, or validator rejection disqualifies, regardless of cost or fill count
Do not rank by `filled`. Disqualify first on invented / scope_violations / validator_rejected, and only then compare cost and latency among survivors. A candidate with filled=0 has not passed — it abstained, which is untested rather than safe. Agreement figures are only as strong as the reference run: check the matching entry in `reference_runs` for how many lines the reference itself filled before trusting a high agreement score.
Neither open-weight candidate is disqualification-free, so we stayed on our stack. Two were bake-offed against the frozen prompt. Qwen3-235B invented nothing and committed no scope violation across all its runs, filling between 0 and 4 of 27 lines — clean, but a small and inconsistent fill. gpt-oss-120b also invented nothing and was faster and cheaper, but in several of its runs it marked a quantity verified from a property-boundary layer — a real area applied to the wrong extent, which is a scope violation and disqualifies it under the rule. Faster and cheaper does not survive a scope violation.
Read the fill numbers against the reference, not in isolation. The opus-5 reference is not a stable target — it fills just 1 of 27 lines on most of its runs and has scope violations of its own — so a candidate scored against its best run looks worse than it is, and no clean multi-filled reference exists yet to show that any model fills correctly when the drawings do support a figure. What this run establishes is the failure mode, not a ranking.
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
Qwen/Qwen3-235B-A22B-Instruct-2507 deepinfra | 1/27 | 0 | 0 | — | 7 | 4.5k | 163.2s |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
openai/gpt-oss-120b together | 4/27 | 0 | 0 | — | 25 | 9.9k | 237.9s |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
Qwen/Qwen3-235B-A22B-Instruct-2507 deepinfra | 1/27 | 0 | 0 | — | 28 | 6.1k | 273.8s |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
openai/gpt-oss-120b together | 4/27 | 0 | 0 | — | 25 | 9.3k | 69.2s |
Qwen/Qwen3-235B-A22B-Instruct-2507 deepinfra | 3/27 | 0 | 0 | — | 27 | 6.0k | 270.0s |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
openai/gpt-oss-120b together | 0/27 | 0 | 0 | — | 30 | 8.6k | 56.0s filled nothing |
Qwen/Qwen3-235B-A22B-Instruct-2507 deepinfra | 1/27 | 0 | 0 | — | 15 | 5.1k | 333.1s |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
openai/gpt-oss-120b together | 0/27 | 0 | 0 | — | 29 | 6.0k | 40.5s filled nothing |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
openai/gpt-oss-120b together | 8/27 | 0 | 1 | — | 7 | 10.9k | 81.0s |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
openai/gpt-oss-120b together | 3/27 | 0 | 2 | — | 26 | 10.0k | 63.1s |
Qwen/Qwen3-235B-A22B-Instruct-2507 deepinfra | 0/27 | 0 | 0 | — | 3 | 4.1k | 189.8s filled nothing |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
Qwen/Qwen3-235B-A22B-Instruct-2507 deepinfra | 4/27 | 0 | 0 | — | 7 | 5.1k | 125.7s |
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
openai/gpt-oss-120b together | 3/27 | 0 | 1 | — | 27 | 11.1k | 79.7s |
Kept for the record. This run predates a fix to the scope auditor, which was flagging one concept that the rules genuinely permit to use the site boundary — so scope violations shown here may not be real. Read the current run above.
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
claude-opus-5 anthropic | 1/27 | 0 | 0 | — | 14 | 12.5k | 138.7s |
Kept for the record. This run predates a fix to the scope auditor, which was flagging one concept that the rules genuinely permit to use the site boundary — so scope violations shown here may not be real. Read the current run above.
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
gemini-3.6-flash gemini | 1/27 | 0 | 0 | — | 6 | 4.2k | 46.7s |
gpt-5.4-2026-03-05 openai | 5/27 | 0 | 0 | yes | 27 | 5.6k | 41.0s answer refused |
llama-3.3-70b-versatile groq | 0/27 | 0 | 0 | — | 3 | 2.9k | 5.7s filled nothing |
kimi-for-coding kimi-cli · subscription | 4/27 | 0 | 3 | — | 13 | null | 110.7s |
Kept for the record. This run predates a fix to the scope auditor, which was flagging one concept that the rules genuinely permit to use the site boundary — so scope violations shown here may not be real. Read the current run above.
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
gemini-3.6-flash gemini | 1/27 | 0 | 0 | — | 23 | 6.0k | 51.5s |
gpt-5.4-2026-03-05 openai | 4/27 | 0 | 3 | — | 19 | 4.5k | 31.0s |
llama-3.3-70b-versatile groq | 0/27 | 0 | 0 | — | 29 | 4.5k | 8.4s filled nothing |
Kept for the record. This run predates a fix to the scope auditor, which was flagging one concept that the rules genuinely permit to use the site boundary — so scope violations shown here may not be real. Read the current run above.
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
kimi-for-coding kimi-cli · subscription | 4/27 | 0 | 3 | — | 12 | null | 101.5s |
The baseline the candidates are measured against — the same harness run on our own stack. Included because a bake-off without a control is just a list of numbers.
These reference runs predate the scope-auditor fix and have not been re-run. One shows three scope violations that may be the same false positive corrected in the current bake-off — we have not verified it either way, so it is left as measured. Note also the spread: most reference runs could defensibly fill only 1 of 27 lines, which is the ceiling every agreement figure on this page is measured against.
| Model | Filled | Invented | Scope | Rejected | Findings | Out | Latency |
|---|---|---|---|---|---|---|---|
claude-opus-5 agent-sdk · subscription | 4/27 | 0 | 0 | — | 13 | 11.7k | 131.9s |
claude-opus-5 agent-sdk · subscription | 4/27 | 0 | 0 | — | 33 | 14.5k | 163.8s |
claude-opus-5 agent-sdk · subscription | 4/27 | 0 | 0 | — | 19 | 15.6k | 191.9s |
claude-opus-5 agent-sdk · subscription | 4/27 | 0 | 0 | — | 14 | 13.3k | 146.0s |
claude-opus-5 agent-sdk · subscription | 4/27 | 0 | 0 | — | 16 | 14.0k | 155.3s |
claude-opus-5 agent-sdk · subscription | 5/27 | 0 | 0 | — | 16 | 15.8k | 180.3s |
claude-opus-5 agent-sdk · subscription | 3/27 | 0 | 0 | — | 18 | 15.6k | 175.7s |
claude-opus-5 agent-sdk · subscription | 4/27 | 0 | 0 | — | 13 | 13.6k | 159.4s |
claude-opus-5 agent-sdk · subscription | 1/27 | 0 | 0 | — | 12 | 10.9k | 125.5s |
claude-opus-5 agent-sdk · subscription | 1/27 | 0 | 0 | — | 16 | 11.8k | 134.4s |
Every figure on this page comes from one JSON file, served openly so you can check it or chart it yourself.
This is one of three. The judge bake-off asks whether a cheaper model can grade as reliably as a premier one, and the router comparison asks whether our routing beats OpenRouter's automatic one. See all three, with what each found.