Live benchmark results across all tested AI models. Sortable by cost, latency, throughput.
Free models (poolside/laguna, nvidia/nemotron) now match premium tier on all metrics—rethink your vendor mix.
Two free models dominate the benchmark: poolside/laguna-xs-2.1 and nvidia/nemotron-3-ultra-550b-a55b both achieve 100% pass rates with zero cost. Groq's openai/gpt-oss-120b delivers fastest latency across ping (283ms) and reasoning (418ms) tests, while gemini-2.5-flash-image failed JSON validation and anthropic claude-opus-4-6 failed tool_use. No regressions detected. Most surprising: multiple free-tier models match or exceed premium models on all metrics.
Two free models define the cost-optimal frontier: openrouter/poolside/laguna-xs-2.1 and openrouter/nvidia/nemotron-3-ultra-550b-a55b both achieve 100% pass rates at zero cost. The laguna model averages 676ms latency, while nemotron averages 1436ms. Both outperform paid competitors on cost-per-quality basis. Groq's openai/gpt-oss-120b (pay-per-token at $0.15/$0.75 per 1M tokens) offers fastest latency for time-sensitive tasks despite minor cost.
For context retrieval, route to groq/openai/gpt-oss-120b (586ms, 100% pass). For speed/throughput, use groq/openai/gpt-oss-120b (425 tokens/sec, 1779ms end-to-end). For structured JSON output, route to groq/openai/gpt-oss-120b (485ms, 100% pass). For tool-use requiring function calls, use fireworks/accounts/fireworks/models/kimi-k2p6 (1738ms, 100% pass) or groq/openai/gpt-oss-120b (293ms, 100% pass). For cost-constrained workloads with no latency SLA, use openrouter/poolside/laguna-xs-2.1 (free, 100% pass rate).
Poolside Laguna XS 2.1 (openrouter, free tier) and NVIDIA Nemotron 3 Ultra 550B (openrouter, free tier) both achieved 100% pass rates across all tests. This shatters the assumption that premium models are necessary—developers can now achieve full compliance with zero cost. Most surprising: qwen/qwen3-vl-8b-thinking (openrouter) delivered fastest median latency of 323ms across all tests, beating groq's 283ms ping by being consistently fast across diverse workloads.
FREE model with 100% pass rate — strong default for cost-sensitive workloads
FREE model with 100% pass rate — strong default for cost-sensitive workloads
lowest median latency across all tests at 323ms
Migrate all cost-sensitive workloads to openrouter/poolside/laguna-xs-2.1 (free, 100% pass rate, 676ms avg latency) or openrouter/nvidia/nemotron-3-ultra-550b-a55b (free, 100% pass rate, 1436ms avg latency). For latency-critical paths, use groq/openai/gpt-oss-120b (283ms ping, 418ms reasoning, 425 tps throughput). Retire together/arize-ai/qwen-2-1.5b-instruct (57% pass rate, multiple failures on context and tool_use).