Skip to main content
  1. Agents/

Automated research

Where the research loop itself runs autonomously: the labs’ science programs, the productized research agent, and the formal-proof engines pointed at the hardest open problems.

  • AlphaProof - DeepMind’s Lean reinforcement-learning solver, from IMO silver in 2024 to officially graded IMO gold via Deep Think in 2025.
  • Anthropic Claude mathematical research - the Claude Code subagent loop that raised the zeta zero bound to 67.2 percent and formalized Fermat’s Last Theorem in 11 days.
  • Harmonic Aristotle - the free agentic theorem prover with a Lean-checked Erdős result, now aimed at software correctness, judged and contested on FormalQualBench.
  • Math Inc. Gauss - the autoformalization agent behind Strong PNT and the sphere-packing proof, with the comparator-audited OpenGauss harness, as of 2026-09-13.
  • OpenAI Deep Research - the productized web-research agent, the breadth-first loop with no machine judge behind it.
  • OpenAI for Science - the lab program from FrontierMath’s first open-problem solve to the disputed Navier-Stokes claim, as of 2026-09-13.
  • Pion - Andon Labs’ closed research preview where persistent agents run a real business with payment tools, the Vending-Bench lineage made product.

Its members are compared on shared rows in the Automated Research Feature Matrix.

Changes
#

  • 2026-09-13 - Added AlphaProof.
  • 2026-09-13 - Added Anthropic Claude mathematical research.
  • 2026-09-13 - Added Harmonic Aristotle.
  • 2026-09-13 - Added Math Inc. Gauss.
  • 2026-09-13 - Added OpenAI Deep Research.
  • 2026-09-13 - Added OpenAI for Science.
  • 2026-09-16 - Added Pion.