<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>human-feedback on tomrochette.com</title>
    <link>https://tomrochette.com/tags/human-feedback/</link>
    <description>Recent content in human-feedback on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Thu, 24 Sep 2026 11:16:33 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/human-feedback/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>LMArena</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>evaluation</category><category>leaderboards</category><category>human-feedback</category>
      <description>&lt;p&gt;LMArena ranks AI models by blind human preference votes: two anonymous models answer the same prompt, you pick the winner, and Bradley-Terry-style statistics turn millions of those picks into leaderboards spanning text, image, video, vision, search, web development, and agents.&#xA;Facts below verified as of 2026-09-24; lmarena.ai is a client-rendered app, so vote and model counts beyond the founding paper&amp;rsquo;s figures could not be read from its HTML.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;It is the field&amp;rsquo;s mood ring: the most cited signal of which model people prefer, and the easiest leaderboard in existence to game.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A website (lmarena.ai), a family of leaderboards (Agent Overall, Text, WebDev, Image, Video, Vision, Document, Search), a WebDev arena (web.lmarena.ai), a blog, and open methodology repositories, run by Arena Intelligence Inc., the company that grew out of the UC Berkeley and LMSYS Chatbot Arena project.&#xA;The founding paper (arXiv:2403.04132, March 2024) describes the pairwise crowdsourcing method and 240K+ votes at the time; the leaderboard methodology source is published as the &lt;code&gt;arena-rank&lt;/code&gt; repository, pushed August 2026.&#xA;The company raised $100M at a $600M valuation in May 2025, led by Andreessen Horowitz and UC Investments, and labs including OpenAI, Google, and Anthropic partner with it to put flagship models in front of voters.&#xA;Recent product motion includes AutoEval scores added to the leaderboards (to complement slowly collected human votes), agent leaderboard categories with task costs (August 2026), and a HarnessTax research post (September 2026).&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;The most-cited leaderboard in the field: a critical paper describing Chatbot Arena as &amp;ldquo;the go-to leaderboard for ranking the most capable AI systems&amp;rdquo; is itself the best evidence of that status.&#xA;The blog posts within days of verification, the arenas run continuously, and HN threads routinely open with its rankings as the premise.&#xA;The company is well capitalized and has converted the academic project into a venture-scale business, which is also the source of its hardest questions.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Preference at scale is a measurement no benchmark suite replaces: it captures whatever makes people pick one answer over another, including style, format, and thoroughness.&lt;/li&gt;&#xA;&lt;li&gt;The methodology code and the voting procedure are public, so the statistics are checkable even when the data pipelines are not.&lt;/li&gt;&#xA;&lt;li&gt;The arena expansion (WebDev, agents, image, video) follows usage: it measures the surfaces people actually use models on.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The gaming record is documented: Meta&amp;rsquo;s Llama 4 Maverick episode (an &amp;ldquo;experimental chat version&amp;rdquo; tested on the arena that differed from the shipped model) forced a policy update, and LMArena&amp;rsquo;s own statement conceded &amp;ldquo;Meta&amp;rsquo;s interpretation of our policy did not match what we expect from model providers&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;The Leaderboard Illusion paper (arXiv:2504.20879) documents &amp;ldquo;undisclosed private testing practices&amp;rdquo; that &amp;ldquo;benefit a handful of providers&amp;rdquo;, counting 27 private Meta variants tested before the Llama 4 release and sampling-rate asymmetries favoring closed models.&lt;/li&gt;&#xA;&lt;li&gt;The sharpest criticism (Surge AI&amp;rsquo;s &amp;ldquo;LMArena is a cancer on AI&amp;rdquo;, 246 points on HN in January 2026) argues the format &amp;ldquo;rewards superficiality over accuracy&amp;rdquo; because &amp;ldquo;the easiest way to climb the leaderboard isn&amp;rsquo;t to be smarter; it&amp;rsquo;s to hack human attention span&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;A preference rank is not a capability claim: verbosity and sycophancy win votes that lose tasks, so the number is routinely over-read by headlines.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free to use and to vote.&#xA;No public pricing page exists; the company is venture funded, and its &amp;ldquo;Try Arena&amp;rdquo; product surfaces are free at the time of verification.&#xA;No reader-facing price is stated, so no price history applies.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;: controlled first-party evals with published weights; choose it when you need price and speed, the arena when you need preference.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;: revealed preference (spend) versus stated preference (votes).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;: benchmark aggregation, which at least labels what it cannot verify.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended as the fastest read on which models feel better to people, and as a required filter on any headline of the form &amp;ldquo;X tops the arena&amp;rdquo;; not as the number you commit money against.&lt;/strong&gt;&#xA;My disagreeable claim: the leaderboard is the least valuable thing the arenas produce, because the vote stream is quietly one of the largest human-feedback datasets ever assembled, and whoever holds it holds a training asset, not just a ranking.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - the controlled-eval counterpart to crowd preference&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt; - what people pay for, against what they vote for&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt; - the research-nonprofit model of measurement this company left behind&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/you-cannot-out-review-a-machine-by-hand/&#34; &gt;You Cannot Out-Review a Machine by Hand&lt;/a&gt; - human judgment as a bottleneck, applied to review instead of ranking&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://lmarena.ai/&lt;/a&gt; - homepage meta: blind comparison, vote-driven leaderboards across text, image, and code (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://blog.lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=blog.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://blog.lmarena.ai/&lt;/a&gt; - Arena Intelligence Inc. identity, leaderboard families, 2026 post dates (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://blog.lmarena.ai/how-it-works/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=blog.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://blog.lmarena.ai/how-it-works/&lt;/a&gt; - the vote flow and identity reveal procedure (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2403.04132&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2403.04132&lt;/a&gt; - the founding paper: method and 240K+ votes (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2504.20879&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2504.20879&lt;/a&gt; - The Leaderboard Illusion: private testing, 27 Meta variants, sampling asymmetries (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=techcrunch.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&lt;/a&gt; - $100M seed, $600M valuation, investors, Berkeley origin (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.theverge.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming&lt;/a&gt; - the Maverick gaming episode and LMArena&amp;rsquo;s policy response (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://surgehq.ai/blog/lmarena-is-a-plague-on-ai&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=surgehq.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://surgehq.ai/blog/lmarena-is-a-plague-on-ai&lt;/a&gt; - the strongest critical essay on the format (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/lmarena&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/lmarena&lt;/a&gt; - methodology repositories including arena-rank, pushed 2026-08-04 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.lmarena.ai/&lt;/a&gt; - the WebDev arena surface (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
