LiveBench is a public LLM benchmark and leaderboard that limits test-set contamination by rotating in new questions, and that scores every answer against objective ground truth instead of an LLM judge.
It is this category’s checkable number: the questions, the model answers, and the judgments are all published, so any score on it can be recomputed from scratch.
What it is #
An open-source benchmark (github.com/LiveBench/LiveBench, 1,339 stars as of 2026-10-06) with a leaderboard at livebench.ai, created by Colin White and collaborators including Yann LeCun, Tom Goldstein, and Micah Goldblum, and maintained day to day by Abacus.AI engineers. The design (an ICLR 2025 Spotlight paper) calls for monthly question releases drawn from recent sources (math competitions, arXiv papers, news articles, movie synopses), a verifiable ground-truth answer per question, and about 1,000 questions across 6 categories and 18 tasks. An agentic coding category joined in May 2025: models work through repository tasks in Docker containers using the Mini-SWE-Agent harness at a 250-step limit. The evaluation offer is unusually direct: open a GitHub issue or email the team, and they will run your model. All questions, model answers, and model judgments are published on Hugging Face.
Status #
Active, but its defining cadence has slipped: the repository was pushed 2026-09-29 (adding GPT-6.1 Sol and Grok 4.7 configurations under a config gate that requires dated pricing citations), yet the changelog’s newest question release is 2026-01-08, and releases that came monthly through 2024 have arrived in bursts since. That gap matters because the September-October release wave (GPT-6, Claude Opus 5.5, Gemini 4 Argon) shipped onto a question set that has been frozen for nine months. Community footprint is thin on HN (the launch thread drew 6 points and 0 comments) even though model cards and release posts cite LiveBench scores routinely. The leaderboard site is a client-rendered app, so automated fetchers see only the title; the fetchable surfaces are the repository, the changelog, and the Hugging Face datasets.
Strengths #
- Objective ground-truth scoring avoids the judge biases the project itself documents (its changelog cites GPT-4-Turbo pass/fail judgments erroring on up to 46% of hard reasoning and math answers).
- Full transparency: questions, answers, and judgments are downloadable, and the whole run pipeline is public.
- The agentic coding category tests models in Dockerized repositories across multiple turns, not single-shot prompts.
Cautions #
- The contamination-limiting premise depends on rotation speed, and rotation has slowed: the longer a question set freezes, the more every score on it measures memorization instead of capability.
- The authors themselves downgraded the claim, retitling the paper from Contamination-Free (v1, 2024) to Contamination-Limited (v2, 2025), and the November 2025 changelog entry says rotation only resolved saturation and contamination “to some extent”.
- Abacus.AI, a commercial ML-platform vendor, runs the day-to-day operation, so the benchmark also functions as research marketing for that company.
- The README carries a stale release note (calling 2025-04-25 the current release), marks local-model inference unmaintained, and warns the agentic track can need up to 150GB of Docker images.
Pricing #
Free: the repository, the questions, the answers, the judgments, and the leaderboard are public, and the team evaluates requested models at no charge. No paid tier exists, so no price history applies.
Compared to #
- Artificial Analysis: the operator’s private evals with closed question sets; LiveBench publishes everything it scores.
- LLM Stats: aggregates other benchmarks’ published numbers; LiveBench generates its own evidence.
- AI Release Tracker: logs what shipped; LiveBench scores what ships.
Bottom line #
Recommended for an objective, auditable current-capability read you can rerun end to end, especially across open-weight models; not for preference, spend, or trend questions. My disagreeable claim: the rotation gap is the number that matters most in this note, because a frozen contamination-limiting benchmark quietly converts into a static one, and a static one is just a saturated benchmark with better manners.
Changes #
- 2026-10-06 - Created.
See also #
- Artificial Analysis - the closed-evals counterpart to a fully published benchmark
- LLM Stats - the aggregator that consumes benchmark outputs like LiveBench’s
- AI Release Tracker - the release stream that dates every model on the leaderboard
- Trackers and Leaderboards Feature Matrix - the category comparison this note joins
- Model Selection for Coding Tasks - where benchmark scores land when a model choice gets made
References #
https://api.github.com/repos/LiveBench/LiveBench - repository metadata: 1,339 stars, pushed 2026-09-29, license listed as Other (fetched 200, 2026-10-06)
https://raw.githubusercontent.com/LiveBench/LiveBench/main/README.md - design, 6 categories and 18 tasks, the evaluation offer, the stale release note, the unmaintained local-inference path, and the 150GB agentic warning (fetched 200, 2026-10-06)
https://raw.githubusercontent.com/LiveBench/LiveBench/main/changelog.md - release history: monthly through 2024, agentic category 2025-05-30, Mini-SWE-Agent switch 2025-10-03, newest question set 2026-01-08, and the 46% GPT-4-Turbo judge-error figure (fetched 200, 2026-10-06)
https://arxiv.org/abs/2406.19314 - the paper: ICLR 2025 Spotlight, the 18-author list including LeCun, and the Contamination-Free to Contamination-Limited title change (fetched 200, 2026-10-06)
https://api.github.com/repos/LiveBench/LiveBench/commits?per_page=6 - 2026 activity: GPT-6.1 Sol and Grok 4.7 configs merged with dated pricing citations under the config gate (fetched 200, 2026-10-06)
https://hn.algolia.com/api/v1/search?query=livebench&tags=story - the thin HN footprint: launch thread at 6 points, 0 comments (fetched 200, 2026-10-06)
https://livebench.ai/ - the leaderboard surface, recorded as a client-rendered app that returned only its title to this fetch (fetched 200, 2026-10-06)