<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>leaderboards on tomrochette.com</title>
    <link>https://tomrochette.com/tags/leaderboards/</link>
    <description>Recent content in leaderboards on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Thu, 24 Sep 2026 11:16:33 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/leaderboards/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>LMArena</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>evaluation</category><category>leaderboards</category><category>human-feedback</category>
      <description>&lt;p&gt;LMArena ranks AI models by blind human preference votes: two anonymous models answer the same prompt, you pick the winner, and Bradley-Terry-style statistics turn millions of those picks into leaderboards spanning text, image, video, vision, search, web development, and agents.&#xA;Facts below verified as of 2026-09-24; lmarena.ai is a client-rendered app, so vote and model counts beyond the founding paper&amp;rsquo;s figures could not be read from its HTML.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;It is the field&amp;rsquo;s mood ring: the most cited signal of which model people prefer, and the easiest leaderboard in existence to game.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A website (lmarena.ai), a family of leaderboards (Agent Overall, Text, WebDev, Image, Video, Vision, Document, Search), a WebDev arena (web.lmarena.ai), a blog, and open methodology repositories, run by Arena Intelligence Inc., the company that grew out of the UC Berkeley and LMSYS Chatbot Arena project.&#xA;The founding paper (arXiv:2403.04132, March 2024) describes the pairwise crowdsourcing method and 240K+ votes at the time; the leaderboard methodology source is published as the &lt;code&gt;arena-rank&lt;/code&gt; repository, pushed August 2026.&#xA;The company raised $100M at a $600M valuation in May 2025, led by Andreessen Horowitz and UC Investments, and labs including OpenAI, Google, and Anthropic partner with it to put flagship models in front of voters.&#xA;Recent product motion includes AutoEval scores added to the leaderboards (to complement slowly collected human votes), agent leaderboard categories with task costs (August 2026), and a HarnessTax research post (September 2026).&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;The most-cited leaderboard in the field: a critical paper describing Chatbot Arena as &amp;ldquo;the go-to leaderboard for ranking the most capable AI systems&amp;rdquo; is itself the best evidence of that status.&#xA;The blog posts within days of verification, the arenas run continuously, and HN threads routinely open with its rankings as the premise.&#xA;The company is well capitalized and has converted the academic project into a venture-scale business, which is also the source of its hardest questions.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Preference at scale is a measurement no benchmark suite replaces: it captures whatever makes people pick one answer over another, including style, format, and thoroughness.&lt;/li&gt;&#xA;&lt;li&gt;The methodology code and the voting procedure are public, so the statistics are checkable even when the data pipelines are not.&lt;/li&gt;&#xA;&lt;li&gt;The arena expansion (WebDev, agents, image, video) follows usage: it measures the surfaces people actually use models on.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The gaming record is documented: Meta&amp;rsquo;s Llama 4 Maverick episode (an &amp;ldquo;experimental chat version&amp;rdquo; tested on the arena that differed from the shipped model) forced a policy update, and LMArena&amp;rsquo;s own statement conceded &amp;ldquo;Meta&amp;rsquo;s interpretation of our policy did not match what we expect from model providers&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;The Leaderboard Illusion paper (arXiv:2504.20879) documents &amp;ldquo;undisclosed private testing practices&amp;rdquo; that &amp;ldquo;benefit a handful of providers&amp;rdquo;, counting 27 private Meta variants tested before the Llama 4 release and sampling-rate asymmetries favoring closed models.&lt;/li&gt;&#xA;&lt;li&gt;The sharpest criticism (Surge AI&amp;rsquo;s &amp;ldquo;LMArena is a cancer on AI&amp;rdquo;, 246 points on HN in January 2026) argues the format &amp;ldquo;rewards superficiality over accuracy&amp;rdquo; because &amp;ldquo;the easiest way to climb the leaderboard isn&amp;rsquo;t to be smarter; it&amp;rsquo;s to hack human attention span&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;A preference rank is not a capability claim: verbosity and sycophancy win votes that lose tasks, so the number is routinely over-read by headlines.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free to use and to vote.&#xA;No public pricing page exists; the company is venture funded, and its &amp;ldquo;Try Arena&amp;rdquo; product surfaces are free at the time of verification.&#xA;No reader-facing price is stated, so no price history applies.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;: controlled first-party evals with published weights; choose it when you need price and speed, the arena when you need preference.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;: revealed preference (spend) versus stated preference (votes).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;: benchmark aggregation, which at least labels what it cannot verify.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended as the fastest read on which models feel better to people, and as a required filter on any headline of the form &amp;ldquo;X tops the arena&amp;rdquo;; not as the number you commit money against.&lt;/strong&gt;&#xA;My disagreeable claim: the leaderboard is the least valuable thing the arenas produce, because the vote stream is quietly one of the largest human-feedback datasets ever assembled, and whoever holds it holds a training asset, not just a ranking.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - the controlled-eval counterpart to crowd preference&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt; - what people pay for, against what they vote for&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt; - the research-nonprofit model of measurement this company left behind&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/you-cannot-out-review-a-machine-by-hand/&#34; &gt;You Cannot Out-Review a Machine by Hand&lt;/a&gt; - human judgment as a bottleneck, applied to review instead of ranking&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://lmarena.ai/&lt;/a&gt; - homepage meta: blind comparison, vote-driven leaderboards across text, image, and code (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://blog.lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=blog.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://blog.lmarena.ai/&lt;/a&gt; - Arena Intelligence Inc. identity, leaderboard families, 2026 post dates (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://blog.lmarena.ai/how-it-works/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=blog.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://blog.lmarena.ai/how-it-works/&lt;/a&gt; - the vote flow and identity reveal procedure (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2403.04132&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2403.04132&lt;/a&gt; - the founding paper: method and 240K+ votes (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2504.20879&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2504.20879&lt;/a&gt; - The Leaderboard Illusion: private testing, 27 Meta variants, sampling asymmetries (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=techcrunch.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&lt;/a&gt; - $100M seed, $600M valuation, investors, Berkeley origin (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.theverge.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming&lt;/a&gt; - the Maverick gaming episode and LMArena&amp;rsquo;s policy response (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://surgehq.ai/blog/lmarena-is-a-plague-on-ai&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=surgehq.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://surgehq.ai/blog/lmarena-is-a-plague-on-ai&lt;/a&gt; - the strongest critical essay on the format (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/lmarena&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/lmarena&lt;/a&gt; - methodology repositories including arena-rank, pushed 2026-08-04 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.lmarena.ai/&lt;/a&gt; - the WebDev arena surface (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>Trackers and Leaderboards Feature Matrix</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/</guid>
      <category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>comparison</category><category>trackers-and-leaderboards</category><category>benchmarks</category><category>leaderboards</category><category>open-data</category>
      <description>&lt;p&gt;This matrix compares the six members of the Trackers and leaderboards category: sites whose product is a continuously refreshed number about the AI field itself.&#xA;The members split cleanly on what their number measures: what shipped (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt;), what the operator measured (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;), how the field moves (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt;), what public evidence aggregates to (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;), what people prefer (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt;), and what people pay for (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;).&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;No row of this matrix crowns a winner, because the rows are different questions; the failure mode is citing a site for a question it does not answer, usually LMArena ranks quoted as capability or OpenRouter tokens quoted as market share.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;Legend: ✓ supported, ✗ not supported, ~ partial or conditional, ? not verified.&#xA;Each column links to the full research note; every cell traces to a source cited there or in the references.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;The matrix&#xA;    &lt;div id=&#34;the-matrix&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#the-matrix&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;table&gt;&#xA;&#x9;&lt;thead&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/thead&gt;&#xA;&#x9;&lt;tbody&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Kind&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;release timeline plus flat-file corpus&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;independent benchmarking site and data business&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;research nonprofit with open datasets&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;composite aggregator with agent-facing API&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;blind preference arena&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;gateway usage rankings&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;The number measures&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;launch-day facts: what shipped, when, with which claimed scores&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;the operator&amp;rsquo;s own controlled evals, prices, and speed runs&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;long-run trends: compute, cost, capability over time&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;public benchmark evidence, normalized with uncertainty&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;blind human preference votes&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;tokens processed through one gateway&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Run by&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;To sider ApS, a Danish side project (one visible operator)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;venture-backed independent company (AI Grant seed)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;501(c)(3) nonprofit, itemized donors, about 50 people&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;ZeroEval Inc. (self-displayed YC backing)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Arena Intelligence Inc. ($100M seed at $600M)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;OpenRouter, a gateway acquired-by-Stripe (announced 2026-08-19)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Coverage&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;231 releases, 10 labs (as of 2026-08-26)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;671 models, 100+ providers, 1,000+ endpoints, chips&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;3,200+ models since 1950, data centers, chips, companies&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;398 canonical models, 50+ benchmarks claimed&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;text, image, video, vision, search, webdev, agent arenas&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;500+ models via the gateway, 80+ providers&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Update cadence&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;on each release; dataset last updated 2026-08-26&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;daily changelog; index versions iterate weekly&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;near-daily data updates, page-stamped&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;continuously; &amp;ldquo;within hours of release&amp;rdquo; claimed&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;continuous votes; product posts within days&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;daily UTC buckets, about one day of lag&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Methodology published&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ FAQ and data notes, no formal methodology&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ methodology hub with evaluation lists and weights&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ transparency page, papers, and data documentation&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ score construction published, details gated&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ methodology repo (arena-rank) plus founding paper&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ on-page caveats plus Data API docs&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Data access&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ /models.json and llms-full.txt, free with attribution&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ web charts free, data platform paid, no public API documented&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ CC BY datasets and a Python client&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ REST plus 11 MCP tools, free tier&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ arenas open, methodology open, raw votes not re-offered&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ CC BY 4.0 JSON via Data API, history to 2025-01-01&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Reader pricing&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free, no ads; Pro instant alerts, price unpublished&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free site; enterprise and data-platform tiers, prices unpublished&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free, donation funded&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free tier; Builder $99/month; Commercial contract&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free, attribution required&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Independence caveat&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;one operator, one company registration, no second pair of eyes&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;sells private benchmarking to the labs it ranks (disclosed policy)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;consulted for OpenAI and DeepMind, disclosed and audited&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;paid eval services and a consumer sibling in the same company&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;lab partnerships and a documented private-testing history&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;owned by the gateway it measures, being absorbed by Stripe&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Verification hooks&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ full corpus re-downloadable and diffable&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ changelog and methodology public, raw runs private&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ data, notebooks, and code public&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ API re-queryable, scoring internals partly gated&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ methodology code public, vote data not re-offered&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ daily snapshots re-fetchable under CC BY&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/tbody&gt;&#xA;&lt;/table&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Reading the matrix&#xA;    &lt;div id=&#34;reading-the-matrix&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#reading-the-matrix&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;The row that sorts the category is &amp;ldquo;the number measures&amp;rdquo;, and it doubles as a citation guide: release questions go to the tracker, present-tense quality to Artificial Analysis, trend claims to Epoch, quick composites to LLM Stats, preference to LMArena, spend to OpenRouter.&lt;/strong&gt;&#xA;The two most-quoted members are also the two most misquoted: LMArena ranks get read as capability, and gateway tokens get read as market share, when the matrix&amp;rsquo;s own rows show both measure something narrower.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Verification strength tracks institutional type, not popularity.&lt;/strong&gt;&#xA;The nonprofit (Epoch) publishes its funders, its consultations, and its data; the gateway (OpenRouter) publishes its raw daily numbers under CC BY; the side project (AI Release Tracker) lets you diff its whole corpus; while the three venture-scale companies publish methodology and products but keep raw runs, votes, or scoring internals private.&#xA;&lt;strong&gt;If you need to check the work rather than read the work, the checkable columns are the nonprofit&amp;rsquo;s, the gateway&amp;rsquo;s, and the side project&amp;rsquo;s, which is the opposite of what citation frequency would predict.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Every column carries a conflict row, and the conflicts are structural rather than scandals: the benchmark firm sells benchmarking to the benchmarked, the arena partners with the labs it ranks, the aggregator sells evaluation services, the gateway&amp;rsquo;s data is its own marketing, and the nonprofit consults for the labs.&lt;/strong&gt;&#xA;Epoch&amp;rsquo;s FrontierMath episode and LMArena&amp;rsquo;s Maverick episode are the two documented failures, and both produced their institution&amp;rsquo;s strongest disclosure artifacts, which is the pattern worth watching on every refresh of this page.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Choosing from the matrix&#xA;    &lt;div id=&#34;choosing-from-the-matrix&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#choosing-from-the-matrix&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Needing a dated record of what shipped and what the lab claimed that day: AI Release Tracker, via its flat files.&lt;/li&gt;&#xA;&lt;li&gt;Picking a model or provider this week and needing quality against price and speed: Artificial Analysis, with the Endpoint Accuracy Index for the cheap-provider question.&lt;/li&gt;&#xA;&lt;li&gt;Citing how fast the field moves, with downloadable data: Epoch AI.&lt;/li&gt;&#xA;&lt;li&gt;Wiring leaderboard data into an agent or a script: LLM Stats&amp;rsquo;s free API and MCP tools, accepting the gated internals.&lt;/li&gt;&#xA;&lt;li&gt;Sensing which model people prefer right now: LMArena, read as a mood ring and never as a capability claim.&lt;/li&gt;&#xA;&lt;li&gt;Asking what developers actually spend on: OpenRouter Rankings, quoted with its as-of date and its single-gateway caveat.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created with six columns (AI Release Tracker, Artificial Analysis, Epoch AI, LLM Stats, LMArena, OpenRouter Rankings) when the category was seeded at the owner&amp;rsquo;s request.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-selection-for-coding-tasks/&#34; &gt;Model Selection for Coding Tasks&lt;/a&gt; - the decision layer these measurements feed, and the leaderboard skepticism it argues for&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/people-and-publications/people-and-publications-feature-matrix/&#34; &gt;People and Publications Feature Matrix&lt;/a&gt; - the voices interpreting these numbers, compared on their own matrix&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/keeping-up-with-ai/&#34; &gt;Keeping Up With AI Is a Losing Strategy&lt;/a&gt; - the filtering argument for keeping this category small and pull-driven&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.archive.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/&lt;/a&gt; - AI Release Tracker column: scope, footprint, alerts, operator (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/&lt;/a&gt; - Artificial Analysis column: model count, indexes, provider coverage (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/methodology/intelligence-benchmarking&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/methodology/intelligence-benchmarking&lt;/a&gt; - Artificial Analysis column: published evaluation weights (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/about/transparency&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/about/transparency&lt;/a&gt; - Epoch AI column: nonprofit status, funders, consultations (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://llm-stats.com/developer&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/developer&lt;/a&gt; - LLM Stats column: API tiers, MCP tools, quotas (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=techcrunch.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&lt;/a&gt; - LMArena column: funding and company identity (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2504.20879&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2504.20879&lt;/a&gt; - LMArena column: the private-testing findings (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/rankings&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/rankings&lt;/a&gt; - OpenRouter Rankings column: methodology, caveats, licensing (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/&lt;/a&gt; - OpenRouter Rankings column: the ownership question (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://minimaxir.com/2026/05/openrouter-hy3/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=minimaxir.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://minimaxir.com/2026/05/openrouter-hy3/&lt;/a&gt; - OpenRouter Rankings column: the free-tier distortion evidence (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
