Eight loops automate research today, and the row that separates them is not capability but judging: everything with a Lean kernel or an official grader behind it produces checkable artifacts, everything without one produces prose or artifacts an adversarial critic and a human must accept, and the newest columns split between a bank account and a producer-critic pair. Every cell traces to its member note and that note’s fetched references.
The matrix #
| Row | Agon | AlphaProof | Anthropic Claude mathematical research | Harmonic Aristotle | Math Inc. Gauss | OpenAI Deep Research | OpenAI for Science | Pion |
|---|---|---|---|---|---|---|---|---|
| Operator | AutoResearch-Factory (University of Maryland and collaborators) | Google DeepMind | Anthropic | Harmonic | Math Inc. (DARPA expMath-supported) | OpenAI (product) | OpenAI (lab program) | Andon Labs (YC-backed) |
| The loop produces | Reviewed ideas, proposals, experiment workspaces, and paper drafts from a one-line topic | Competition-grade proofs and verified reasoning training | New theorems and formalized proofs | Machine-checked proofs of stated problems | Lean formalizations at record scale | Cited web-research reports | Benchmark firsts and claimed solutions, published with Lean artifacts | Real revenue-and-loss data from agents running actual businesses continuously |
| Human input in the loop | Topic and standards; built to run unattended for hours | 2024: manual Lean translation; 2025: none, end to end | One prompter, expert review after | A problem statement | Blueprints and scaffolding, review of key lemmas | A question | Case-study curation; disputed in the math claims | High-level direction only, through the Andonos managing agent |
| Who judges | Independent producer-critic agent loops on fresh contexts, plus the human for invisible failures | Lean kernel plus official IMO graders | Lean comparator plus named human experts | The Lean kernel | Lean comparator, specification-based | No machine judge; the human reads | Lean kernel on the published formalization; human acceptance and credit still in dispute | No machine judge; the bank account plus Andon’s own monitoring |
| Lean formal verification | No | 2024 yes, 2025 natural language | Yes (zeta and FLT artifacts) | Yes | Yes | No | Yes since 2026-09-08 (Lean 4 artifacts published; they cover Clay’s forced-blowup option, the weaker of the two formulations) | No |
| Surface today | MIT Claude Code plugin run from a separate artifacts workspace | Research system; Gemini 3 Deep Think on AI Ultra plus Gemini API early access | Unreleased models; artifacts on GitHub | Free web agent with login | OpenGauss open source; Gauss in beta | ChatGPT plans | Subscriptions and academic credits | Proprietary research preview with waitlist, no repo |
| Pricing as of 2026-09-18 | Free and MIT, no paid tier | Bundled in the Ultra subscription | Free artifacts, internal compute | Free; $1,000,000 grant program | OpenGauss free; about $25 per benchmark solve | Plan quotas; Pro at $200/month | Program-level; GPT-5 Pro at $200/month in case studies | ? none published; seed tokens funded, planned revenue share |
| Millennium-problem engagement | None claimed; mathematics is one of several domains, not the focus | None claimed; IMO as the public proxy | Attempted the Riemann hypothesis, failed productively (41.6 to 67.2 percent zero bound) | None public | Strong PNT as the gateway toward the Riemann hypothesis | None | Navier-Stokes claimed with a Lean certificate covering the forced-blowup option, credit and formulation both disputed | None; the eval lineage is Vending-Bench, not mathematics |
How to read it #
The judging row is the deciding one, and it repeats a pattern this section tracks in software tooling: outputs are exactly as trustworthy as the verifier behind them. AlphaProof’s 2024 result and Math Inc.’s formalizations carry kernel-level guarantees; the Anthropic results add named human reviewers on top of the kernel. Aristotle’s headline claims are real where Lean checked them and contested where only the vendor graded them. OpenAI’s two entries are prose-only loops: Deep Research cites, the science program claims, and neither has a machine judge. Pion is the first proprietary column and the only one whose output is neither artifact nor prose but money: its agents run real businesses, its judge is a bank account plus the operator’s own monitoring, and its launch thread’s contradiction (the operator calling autonomous resource acquisition the most troubling capability while releasing exactly that) is recorded in its note. Agon is the first column whose judge is a critic agent on a fresh context rather than a kernel, a grader, or a full-time operator, and its paper’s own taxonomy names the failure classes that oracle cannot see, which makes it the cheapest loop to run and the least verified. The Millennium column is uniformly no: nothing here has solved one, the closest engagement is a failed-but-productive Riemann attempt and a disputed Navier-Stokes claim, and FrontierMath’s problems are explicitly built below Millennium scale.
Changes #
- 2026-09-13 - Created in the same run the Automated research category was seeded, with six columns.
- 2026-09-16 - Extended from six to seven columns with Pion (Andon Labs), appended alphabetically after OpenAI for Science, with proprietary and no-repo cells marked as such and pricing marked unverified pending the revenue-share model.
- 2026-09-25 - Updated the OpenAI for Science column for the published Lean certificates (Navier-Stokes claim now machine-checkable, acceptance and credit still pending), removed the verification preamble, linked the header row, and normalized the separator row.
- 2026-09-27 - Extended from seven to eight columns with Agon, inserted first alphabetically, with the judging prose and the failure-taxonomy boundary updated.
- 2026-09-29 - The OpenAI for Science cells moved to the documented formulation fight (Scientific American, 2026-09-21): the certificate covers Clay option C, and a September 17 three-mathematician proof shows the method cannot extend to the unforced problem.
- 2026-10-02 - Moved the AlphaProof Surface today cell for the February 2026 Gemini 3 Deep Think update, which added the first Gemini API early-access path alongside the AI Ultra rollout.
See also #
- AlphaProof - the officially graded DeepMind lineage
- Anthropic Claude mathematical research - the subagent-fleet loop with comparator-checked artifacts
- Harmonic Aristotle - the hosted theorem prover
- Math Inc. Gauss - the autoformalization record and audited harness
- OpenAI for Science - the lab program behind the disputed claims
References #
https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/ - the AlphaProof column’s 2025 facts and the IMO grading caveat
https://www.anthropic.com/research/riemann-zeta - the Anthropic column’s zeta bound, subagent loop, and validation chain
https://aristotle.harmonic.fun/ - the Aristotle column’s surfaces, positioning, and grant program
https://math.inc/formalqualbench - the Gauss column’s audited benchmark numbers and comparator methodology
https://thenextweb.com/news/bubeck-navier-stokes-account-apology-altman - the OpenAI for Science column’s dispute facts
https://en.wikipedia.org/wiki/FrontierMath - the below-Millennium scope of the benchmark framing
https://andonlabs.com/blog/why-we-built-pion - the Pion column’s launch facts, deployment record, and the most-troubling admission