Open-Source LLM Benchmarks vs Closed Frontier Models
SWE-bench Verified over time
Agentic issue resolution (%). Thin lines connect each class's releases (extreme outliers omitted); dashed trendlines: one per closed lab and one per open size class, least-squares fit to every point in the series (extreme outliers excluded) and extended past today. Hollow points: approximate release date. Hover the list to locate a model.
SWE-rebench: decontaminated resolved rate over time
Continuously refreshed with post-release GitHub tasks (harder than SWE-bench Verified; honest for open models). SWE-rebench draws a fresh task batch monthly and never re-runs older models, so raw leaderboard scores are measured on different tasks and cannot be compared directly. Batch difficulty is fitted out below and every model is placed on one reference batch. Hollow points: never ran the reference batch, score estimated from the fit.
BFCL v4: tool calling over time
Berkeley Function Calling overall (%), live leaderboard, best variant per model.
LiveCodeBench over time
Contamination-free code generation (%); live leaderboard plus wiki-curated edge checkpoints.