What's new
27m ago Recomputed internal benchmark aggregates Refreshed the remediation layer; public rankings still require row-level provenance and current dates. 27m ago Attempted Chatbot Arena collection Collection attempt only; cached scores and transport success do not verify measurements. 27m ago Updated speed measurements Refreshed output speed and latency references for tracked models. 27m ago Pulled latest OpenRouter price index Updated comparison data for providers and routed model endpoints. 27m ago Validated official pricing snapshots Rechecked provider pricing pages against the comparison database. 9h ago ‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop The Verge featured in the latest daily brief. 9h ago An Anthropic AI model sent a false homicide tip to Philadelphia police TechCrunch featured in the latest daily brief. 9h ago Jobs market snapshot refreshed 1,020 open roles across 10 tracked companies. 27m ago Recomputed internal benchmark aggregates Refreshed the remediation layer; public rankings still require row-level provenance and current dates. 27m ago Attempted Chatbot Arena collection Collection attempt only; cached scores and transport success do not verify measurements. 27m ago Updated speed measurements Refreshed output speed and latency references for tracked models. 27m ago Pulled latest OpenRouter price index Updated comparison data for providers and routed model endpoints. 27m ago Validated official pricing snapshots Rechecked provider pricing pages against the comparison database. 9h ago ‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop The Verge featured in the latest daily brief. 9h ago An Anthropic AI model sent a false homicide tip to Philadelphia police TechCrunch featured in the latest daily brief. 9h ago Jobs market snapshot refreshed 1,020 open roles across 10 tracked companies.
Snapshot 12 Jun 2026

Floor Capability - historical benchmark snapshot

10 snapshot rows How these scores work →

Historical evidence only. These provisional scores were captured on 12 June 2026 and must not be read as a current model ranking.

Score

0-100 saturation. 100 = the lane is complete for that model; a saturated lane gets retired or hardened.

Evidence %

Source coverage behind the score, not model quality.

DSWE

Agentic coding with cost, time, and output tokens, from DeepSWE.

Confidence

Evidence strength, on the ? beside each score. Never another model grade.

10 historical rows - snapshot from 12 June 2026 Open the full Meta Benchmark Hub →

Pulse - current table

4ranked public LLM rows

0open-weight rows

1rows with speed data

Value leader: Claude Sonnet 4.5 (batch)