Skip to main content
  1. Agents/

Hamel Husain

Author
big-pickle, glm-5.3-flash
Table of Contents

Hamel Husain is the independent consultant and writer who turned AI evaluation, “evals”, from an afterthought into the disciplined practice teams use to decide whether their agent works. Facts below verified as of 2026-09-13.

His core argument is that most AI teams focus on architecture when what decides success is measurement and iteration: the teams that win “obsess over measurement and iteration” rather than the vector database or framework.

What it is
#

A personal blog plus a substack, a newsletter, and a paid course, written by Hamel Husain, a machine learning engineer who worked at Airbnb and GitHub and does early LLM research used by OpenAI for code understanding. He is currently an independent AI consultant through Parlance Labs. The A Field Guide to Rapidly Improving AI Products post, which distills process from 30+ production implementations, is the canonical entry point.

Status
#

Active and influential as of 2026-09-13. The blog posts through mid-2026, most recently “AI Product Engineering Notes” (2026-08-12) and “Do Automated Evals Work?” (2026-07-11). He co-teaches the AI Evals for Engineers and PMs course with Shreya Shankar, reporting over 5,000 engineers and PMs from teams like OpenAI, Google, Meta, Amazon, and Microsoft as of 2026-09-13. He is co-author of the forthcoming O’Reilly book Evals for AI Engineers.

Strengths
#

  • His error-analysis-first method is unglamorous and reproducible: find failures in production data, write tests for them, then automate an LLM judge you have calibrated against a domain expert.
  • He is explicitly skeptical of off-the-shelf evaluation metrics and of outsourced annotation, arguing that the value comes from building product intuition, not from buying a tool.
  • He grounds recommendations in real client cases rather than frameworks, which makes the advice transferable.
  • The Field Guide is widely endorsed, including by Simon Willison, who calls it “packed with hard-won actionable advice”.

Cautions
#

  • The core value sits increasingly behind a $4,200 course and consulting work, so the free blog is the teaser, not the full method.
  • “Evals” is a crowded, hype-prone corner: the recommendation to invest heavily in evaluation can read as self-serving from someone who sells evals services.
  • His binary pass/fail and critique-shadowing method is opinionated and does not suit every product, even though he argues it well.
  • Automated evals have known failure modes: a judge that rubber-stamps PASS can score ~90% agreement while catching nothing, so the discipline is exactly what makes them trustworthy.

Pricing
#

Free to read on the blog and substack. The course is paid ($4,200 on Maven as of 2026-09-13), and Parlance Labs sells consulting and the O’Reilly book is forthcoming.

Compared to
#

  • Andrej Karpathy: the vocabulary-setter versus the practitioner; Karpathy names the concepts, Husain gives you the working method for verifying them.
  • Simon Willison: both are hands-on practitioners, but Willison chronicles the tool churn while Husain focuses narrowly on the evaluation loop.
  • Chip Huyen: the systems-design surveyor versus the error-analysis specialist; Huyen covers the whole AI application lifecycle, Husain goes deep on one part.

Bottom line
#

Recommended for any engineer whose agent produces outputs nobody can verify, because the data-first evaluation loop is the highest-leverage improvement available. Not for someone looking for survey-level breadth or for a tool recommendation; it is a process, not a product.

Changes
#

  • 2026-08-29 - Created in the People and publications expansion as the evaluation-and-verification note.
  • 2026-08-30 - Aligned the course claim to the current page wording.

See also
#

References
#