Open-Source LLM Benchmarks vs Closed Frontier Models

Analysis
Smallest LLMs improving faster than SOTA?
Closed AI is 18 to 26 points ahead, but the smallest open models are improving faster than anything else on the board: 2.73 points a month against 1.83. The statistics behind this dashboard, done properly.

SWE-bench Verified over time

Agentic issue resolution (%). Thin lines connect each class's releases (extreme outliers omitted); dashed trendlines: one per closed lab and one per open size class, least-squares fit to every point in the series (extreme outliers excluded) and extended past today. Hollow points: approximate release date. Hover the list to locate a model.

SWE-rebench: decontaminated resolved rate over time

Continuously refreshed with post-release GitHub tasks (harder than SWE-bench Verified; honest for open models). SWE-rebench draws a fresh task batch monthly and never re-runs older models, so raw leaderboard scores are measured on different tasks and cannot be compared directly. Batch difficulty is fitted out below and every model is placed on one reference batch. Hollow points: never ran the reference batch, score estimated from the fit.

BFCL v4: tool calling over time

Berkeley Function Calling overall (%), live leaderboard, best variant per model.

LiveCodeBench over time

Contamination-free code generation (%); live leaderboard plus wiki-curated edge checkpoints.