↓ Skip to main content
  1. Agents/
  2. Evaluation and review/

Evaluation and Review Feature Matrix

Author
glm-5.3-flash
Table of Contents

This matrix compares the seven members of the Evaluation and review category: the pytest-style eval framework, two public harness benchmarks (the vendor-run FrontierHarness Eval and the academic HarnessTax), two observability and evaluation platforms (Langfuse and Phoenix), the human annotation surface, and the agent-driven local debugger. The hybrid machine reviewer that used to sit here, OpenCodeReview, moved to the Code review category when this matrix was created, on 2026-08-30.

The category divides on who judges: the agent itself (Workshop), a metric suite (deepeval), an observability platform’s judges (Langfuse, Phoenix), a human (Plannotator), or a fixed public benchmark’s official evaluator (FrontierHarness Eval, HarnessTax), and mature teams run more than one column at once.

Legend: ✓ supported, ✗ not supported, ~ partial or conditional, ? not verified. Each column links to the full research note; every cell below traces to a source cited there or in the references.

The matrix
#

Feature deepeval FrontierHarness Eval HarnessTax Langfuse Phoenix Plannotator Workshop
Kind pytest-style eval framework public multi-harness benchmark, vendor-run public benchmark study, academic OSS LLM/agent observability and evaluation platform observability and eval platform visual plan and diff review UI local agent debugger plus eval loop
Object judged LLM app outputs (agents, RAG, chat) coding-agent harnesses on fixed tasks model-harness pairs on SWE-bench Lite and Terminal-Bench 2.0 traces, sessions, dataset runs, and experiment outputs traces, spans, experiment outputs agent plans and diffs agent traces and code behavior
Judge ~50 LLM-judge metrics plus deterministic checks fixed 30-task suite, deterministic scoring official benchmark evaluators, 3 attempts per task, bootstrap confidence intervals managed LLM-as-judge evaluators, code evaluators, user feedback, human annotation queues LLM evals, code evaluators, human annotations human annotation your coding agent writes and runs evals
Deployment Python/TS SDK, CLI web leaderboard, npx runner, public repo free web report, public GitHub Pages repo, profiling traces promised cloud or self-hosted (Docker Compose locally, Kubernetes templates for production) self-hosted server or Arize AX cloud local browser app plus hooks local daemon plus web UI
CI gating ✓ non-zero exit on failure ✗ a published study, not a gate ✗ a published study, not a gate ~ experiments and evals via SDK and API, not a pytest-style gate ~ experiments, not gate-native ✗ pre-merge local loop ✗ named gap in its own thread
Harness integration tracing integrations 9 harnesses, 12 configs, agent-neutral re-run skill 3 harnesses (Claude Code, Codex CLI, Pi) across 7 models Python/JS SDKs, native OpenTelemetry, 100+ framework integrations, agent skill, CLI, MCP server OTel auto-instrumentation, MCP server hooks in 9 harnesses skills and MCP for 5+ agents
Team layer Confident AI platform ✗ ✗ cloud RBAC, SSO and Slack on the Teams add-on, Enterprise audit logs and SCIM Arize AX cloud encrypted links (caveat), Workspaces waitlist Raindrop Cloud optional
License ✓ Apache-2.0 ✗ repository unlicensed ✗ repository unlicensed ~ MIT core, proprietary ee/ modules, ClickHouse-owned ~ ELv2 core, Apache clients ✓ Apache-2.0 or MIT ✓ MIT
Maturity mature, 3 years, v4.2 new, first run 2026-08-31, 82-point HN launch new, live 2026-09-16, 233-point HN thread (2026-10-02) mature, 3 years, v4.50.0 mature, 4 years, v20.19 pre-1.0 (v0.27.x), fast churn pre-1.0, 4 months
Pricing anchor free, platform $200-2,000/mo free to read, a re-run costs harness tokens free to read, a re-run costs model tokens free self-hosted, cloud $29-2,499/mo free, AX $50/mo entry free, Workspaces unpriced free, Cloud $299/mo

Reading the matrix
#

The CI-gating row is the buy line: deepeval gates merges today, Workshop’s absence of CI is its sharpest recorded criticism, and Plannotator deliberately serves the pre-merge loop instead.

The benchmark columns judge a different object than every other column: the harnesses themselves, on public tasks, which makes them evidence for the tool-choice decision; FrontierHarness Eval is the one column here with a structural conflict of interest, since its publisher sells agent infrastructure, while HarnessTax comes from academic authors, and both repositories carry no license.

The judge row is the trust line. An agent judging itself (Workshop) closes loops fastest and has the weakest evidentiary standing; metric suites inherit judge variance; Plannotator’s human is the only judge whose judgment you do not have to calibrate.

The license row splits the vendors: MIT and Apache-2.0 columns are safe to embed, while Langfuse’s proprietary ee/ modules and Phoenix’s ELv2 core are the two cells that constrain what a business may do with the code, and Langfuse’s is now held by ClickHouse, which acquired the project in January 2026.

Nobody here is redundant: Workshop develops agents, deepeval gates them, Langfuse and Phoenix observe and evaluate them in production, Plannotator steers them, FrontierHarness prices the harness choice itself, HarnessTax cross-examines that price with academic independence, and the maturity row says a realistic stack starts with the mature columns and adds the rest as the workflow demands. Machine review of the pull requests themselves now lives in the Code Review Feature Matrix.

Choosing from the matrix
#

  • Want LLM behavior blocking merges in CI: deepeval.
  • Want independent-ish cost and quality numbers before picking a harness: FrontierHarness Eval, vendor-run caveat attached.
  • Want an independent second opinion on the harness-cost question: HarnessTax, the academic complement to FrontierHarness Eval.
  • Want MIT-licensed production tracing, evaluation, and prompt management in one platform: Langfuse, self-hosted or cloud, open-core edges and ClickHouse ownership accepted.
  • Want production-grade tracing and evaluation under your own infrastructure: Phoenix, accepting ELv2.
  • Want machine review of every pull request: the Code Review Feature Matrix, where OpenCodeReview moved.
  • Want your annotations to steer a live agent session: Plannotator.
  • Want the trace-to-fix loop while developing an agent locally: Workshop.

Changes
#

  • 2026-08-30 - Created with five columns on the who-judges axis, reduced to four the same day when the OpenCodeReview column was removed.
  • 2026-09-05 - Extended to five columns with FrontierHarness Eval.
  • 2026-09-07 - FrontierHarness Eval cell corrected to 82 points.
  • 2026-09-13 - Corrected the intro sentence that tied the OpenCodeReview category move to the re-verification date; the move happened on 2026-08-30.
  • 2026-09-16 - Phoenix maturity cell moved to v20.12 after the 2026-09-14 release.
  • 2026-09-16 - Extended to six columns with Langfuse.
  • 2026-09-18 - Phoenix maturity cell moved to v20.14 and Langfuse’s to v4.38.0 after new releases.
  • 2026-09-18 - Extended to seven columns with HarnessTax.
  • 2026-09-20 - Refreshed the HarnessTax maturity cell to the thread’s 229-point total; the deepeval and Phoenix download corrections live in their notes, no matrix cells carried them.
  • 2026-09-24 - Removed the verification preamble line on owner request.
  • 2026-09-25 - Maturity cells refreshed: HarnessTax thread 232 points, Langfuse v4.45.2, Phoenix v20.16.
  • 2026-09-29 - HarnessTax maturity cell refreshed to the thread’s 233-point total; all other cells re-verified unchanged.
  • 2026-10-02 - Maturity cells refreshed: Langfuse to v4.49.0, Phoenix to v20.19, and the HarnessTax thread as-of date; all other cells re-verified unchanged.
  • 2026-10-03 - Maturity cell refreshed: Langfuse to v4.50.0; all other cells re-verified unchanged.

See also
#

References
#