↓ Skip to main content
  1. Agents/

Agents log

Author
Tom Rochette
Table of Contents

Append-only log of every change agents make in this section. One entry per run: date, articles touched, one line per change, model id. Changes to this section that do not appear here were made by a human and must be preserved, not reverted.

2026-08-22
#

  • Scaffolded the section: AGENTS.md, _index.md, log.md [glm-5.3]

2026-08-22 (daily refresh)
#

  • agentic-coding-tools-landscape: created queue item 1 tracker (finished), a four-layer map of the agentic dev tool landscape as of 2026-08-22; all 17 external URLs fetched and verified during this run, 4 internal links resolve to committed articles [glm-5.3]
  • _index.md: listed the new article [glm-5.3]
  • AGENTS.md queue: appended ? for tom: because the companion article the-agentic-development-environment-landscape/ is untracked, so cross-linking it would fail CI [glm-5.3]
  • Queue items 2 (model selection) and 3 (context management) untouched, left for their own runs to keep this one at quality [glm-5.3]

2026-08-22 (owner edit)
#

  • Moved the article queue out of AGENTS.md into agents/queue.md; queue item 1 marked done [owner, glm-5.3]
  • Reworked the section on the rywalker.com/research model: added research notes as a content type (structured skeleton, citation standards) to AGENTS.md, added agents/methodology.md, rewrote queue.md around building a research index by category, restructured _index.md into essays + research index, made the daily refresh a stalest-category-first rolling re-verification [owner, glm-5.3]
  • Diagnosed why /agents/ did not render like sibling sections: the root AGENTS.md file was built by Hugo as a page claiming the /agents/ URL, displacing the section _index.md; fixed by adding build render/list never front matter to both AGENTS.md files (Hugo ignores them now, tools still read them), and added visible/menu to _index.md for consistency with work/ and processes/ [owner, glm-5.3]
  • Removed type: post from every file in the section and added an AGENTS.md rule never to set a type: articles here are section pages, not posts [owner, glm-5.3]

2026-08-22 (daily refresh, harnesses)
#

  • Research index: seeded the Harnesses category (queue item 1) with eight notes, each with 5 sources fetched and verified during this run: claude-code, codex, gemini-cli, opencode, aider, crush, amp, junie [glm-5.3]
  • gemini-cli: status recorded as superseded for individuals, Google stopped serving free and AI Pro/Ultra users on 2026-06-18 per the developers.googleblog.com transition post (Antigravity CLI is the successor); enterprise licenses and paid API keys continue [glm-5.3]
  • agentic-coding-tools-landscape: cross-linked all eight harness notes, corrected the Gemini CLI line that still advertised the dead consumer free tier, updated the subscription-versus-BYOK split sentence, set updated field [glm-5.3]
  • _index.md: added the research index section with the eight harness notes alphabetically [glm-5.3]
  • opencode: recorded CVE-2026-22812 (unauthenticated server, mitigated in v1.1.10) and the March 2026 Anthropic legal-request removal of Claude Pro/Max OAuth (PR 18186) [glm-5.3]
  • claude-code: recorded the June 2026 prompt-steganography finding and the July 2026 systima token-overhead measurements (~33k baseline tokens vs OpenCode ~7k) [glm-5.3]
  • amp: recorded the December 2025 Sourcegraph spinout (Amp Frontier Corporation) and current subscription tiers [glm-5.3]
  • Committed the owner’s uncommitted queue.md edit adding four seed categories (Retrieval, Memory, Executions, Hybrid execution), unchanged, attributed here [owner, glm-5.3]
  • Queue items 2 (model selection) and 3 (context management) remain blocked on their named dependencies (harness notes now exist for item 2; context engine notes still missing for item 3); left for their own runs [glm-5.3]
  • Verification: all internal link targets checked on disk, front matter parsed as YAML, one-sentence-per-line and no-em-dash checks passed across all new and edited files [glm-5.3]

2026-08-23 (style conformance pass, claude-code)
#

  • claude-code: bolded the key insight of each section (sessions portability, status verdict, strengths lead, cautions lead, pay-per-token path, bottom line) to match the per-section bolding convention used across the corpus; meaning unchanged [glm-5.3]
  • claude-code: split the two longest sentences (What it is surfaces sentence, Cautions proxy-study sentence) into one sentence per line, per the short-direct-sentences rule [glm-5.3]
  • claude-code: named the referent in the OpenCode comparison (“the study above” to “the July 2026 proxy study”) [glm-5.3]
  • claude-code: set updated to 2026-08-23 per the revise rule; status, tags, and all other front matter unchanged [glm-5.3]
  • Verification: all 4 internal link targets exist on disk, 4 of 5 reference URLs fetched 200 during this run, the fifth (thereallo.dev) returns a Cloudflare challenge to automated fetches and was verified in the 2026-08-22 run, front matter parses, no em-dashes, one sentence per line [glm-5.3]

2026-08-23 (style conformance pass, codex)
#

  • codex: bolded the key insight of each section (cloud runs the same agent, status verdict, strengths lead, cautions lead, Free-tier inclusion, bottom line) to match the per-section bolding convention used across the corpus; meaning unchanged [glm-5.3]
  • codex: named the referent in the April 2025 caution (“that thread” to “that launch thread”) [glm-5.3]
  • codex: set updated to 2026-08-23 per the revise rule; status, tags, and all other front matter unchanged [glm-5.3]
  • Verification: all 6 internal link targets exist on disk, all 5 reference URLs fetched 200 during this run, front matter parses, no em-dashes, one sentence per line [glm-5.3]

2026-08-23 (style conformance pass, crush)
#

  • crush: hyperlinked the in-body source mentions (FSL-1.1-MIT license, Catwalk, “Crush, come home”, the launch thread, the README) per the create-article link-all-sources rule [glm-5.3]
  • crush: bolded the key insight of each section (LSP-enhanced TUI identity, status verdict, strengths lead, cautions lead, BYOK pricing path, bottom line) to match the per-section bolding convention used across the corpus; meaning unchanged [glm-5.3]
  • crush: split the two longest sentences (opencode-ai continuation, Hyper pricing) into one sentence per line, per the short-direct-sentences rule [glm-5.3]
  • crush: replaced two showy words (“terminal aesthetes” to “engineers who want the best-looking terminal”, “zero spectacle” to “no frills”) per the plain-vocabulary rule, and wrapped crushrc in code ticks [glm-5.3]
  • crush: set updated to 2026-08-23 per the revise rule; status, tags, and all other front matter unchanged [glm-5.3]
  • Verification: all 6 internal link targets exist on disk, all 5 reference URLs fetched 200 during this run, front matter parses, no em-dashes, no banned terms, one sentence per line [glm-5.3]

2026-08-23 (style conformance pass, junie)
#

  • junie: bolded the key insight of each section (ACP skill sharing, status verdict, strengths lead, cautions lead, BYOK pricing path, bottom line) to match the per-section bolding convention used across the corpus; meaning unchanged [glm-5.3]
  • junie: split the What it is sentence combining ACP sharing and remote control into one sentence per line, and named the referent (“It plans” to “Junie plans”) [glm-5.3]
  • junie: set updated to 2026-08-23 per the revise rule; status, tags, and all other front matter unchanged [glm-5.3]
  • Verification: all 7 internal link targets exist on disk, all 5 reference URLs fetched 200 during this run, front matter parses, no em-dashes, one sentence per line [glm-5.3]

2026-08-23 (style conformance pass, opencode)
#

  • opencode: bolded the key insight of each section (one-codebase surfaces, status verdict, strengths lead, cautions lead, BYOK pricing path, bottom line) to match the per-section bolding convention used across the corpus; meaning unchanged [glm-5.3]
  • opencode: hyperlinked the in-body source mentions (July 2026 proxy study, CVE-2026-22812 disclosure, March 2026 Anthropic legal requests) per the create-article link-all-sources rule [glm-5.3]
  • opencode: named the referents (“this is the continuation” to “OpenCode is the continuation”, “will not drive it” to “will not drive OpenCode”, “its auto-started server” to “the auto-started server”, “it is off by default” to “the server is off by default”), moved the pin-your-version advisory to its own bullet per the one-sentence-per-line rule, and smoothed “Claude Code’s about 33k” to “about 33k for Claude Code” [glm-5.3]
  • opencode: set updated to 2026-08-23 per the revise rule; status, tags, and all other front matter unchanged [glm-5.3]
  • Verification: all 6 internal link targets exist on disk, all 5 reference URLs fetched 200 during this run, front matter parses, no em-dashes, no banned terms, one sentence per line [glm-5.3]

2026-08-23 (daily refresh, harnesses re-verification + surfaces seed)
#

  • Research index: re-verified the Harnesses category (the stalest and only category with notes) against live sources; 39 of 40 reference URLs returned 200, and the fortieth (thereallo.dev, cited by claude-code) fetched successfully via the browser path this run [glm-5.3]
  • aider: status downgraded from “active and mature, but slower” to “development stalled” on GitHub API evidence fetched this run (no default-branch commits since 2026-05-22, last tagged release v0.86.0 from 2025-08-09); added a maintenance-risk caution and extended the bottom line; volatile numbers re-dated to 2026-08-23 [glm-5.3]
  • codex: refreshed volatile numbers to about 115k stars, 17.5k forks, about 9.7k commits as of 2026-08-23 via the GitHub API [glm-5.3]
  • amp, claude-code, crush, gemini-cli, junie, opencode: re-verified with no material change (stars within stated bounds, Amp pricing tiers unchanged, all links live); left untouched, and the new-entrant scan found none worth a note, Antigravity CLI stays covered inside the gemini-cli note per the queue’s grouping [glm-5.3]
  • Research index: seeded the Surfaces category (queue item 1) with five notes, each citing 5+ sources all fetched and verified during this run: cursor, vscode-copilot, windsurf, jetbrains, zed [glm-5.3]
  • windsurf: recorded the 2025 collapse-and-rescue arc (OpenAI $3B agreement May, deal off with the CEO leaving for Google July 11, Cognition definitive agreement July 14 citing $82M ARR, employee equity dispute thread) and the June 2, 2026 Devin Desktop rebrand per Cognition’s posts and docs.windsurf.com; windsurf.com itself returns 429 to every automated fetch, stated in the note with pricing left unverified [glm-5.3]
  • agentic-coding-tools-landscape: replaced the stale “Windsurf still plays the AI-editor game” line with the Devin Desktop fact, cross-linked all five new surface notes, bumped the as-of date and updated field to 2026-08-23 [glm-5.3]
  • _index.md: added the Surfaces category (five notes, alphabetical) and updated the aider summary to record the stall [glm-5.3]
  • Queue item 1 remains open with Orchestration as the next seed category; items 2 and 3 untouched; the standing ? for tom: about the untracked the-agentic-development-environment-landscape/ remains unanswered, so it is still not linked [glm-5.3]
  • Verification: all internal link targets across the section’s 15 files exist on disk, front matter parses as YAML, no em-dashes, no banned terms, one sentence per line [glm-5.3]

2026-08-24 (surfaces expansion, owner-requested)
#

  • Owner request in chat: the Surfaces category was missing entries, “Google Antimatter” and “Open Chamber” named explicitly plus a scan for other popular/relevant surfaces; Antimatter does not exist (no DNS, no docs page, the Google surface is Antigravity), OpenChamber verified real, so both were covered under their actual names [glm-5.3]
  • Research index: added five Surfaces notes, each citing 5+ sources all fetched and verified during this run: antigravity, openchamber, kiro, trae, void [glm-5.3]
  • antigravity: recorded the four-surface platform (2.0 command center, IDE, CLI, SDK), the free tier’s model list and weekly limits, the enterprise Google Cloud path, and the incident record (Nov 2025 prompt-injection exfiltration finding, Dec 2025 drive-deletion report, Feb 2026 bans, May 2026 bait-and-switch thread) [glm-5.3]
  • openchamber: recorded the MIT session cockpit around the OpenCode SDK (session goals, five-model fusion, cron scheduling, private relay), v1.20.0 released 2026-08-22, 9.1k stars, and the friction record from the corpus retrospective; cross-linked six-months-with-openchamber [glm-5.3]
  • kiro: recorded the spec-first IDE, credit tiers (free 50 to Power $200 with $0.04 add-on credits), the ACP-compatible-IDE clause, and the Aug 2025 prompt-injection code execution writeup [glm-5.3]
  • trae: recorded the ByteDance tiers (Lite $3 through Ultra $100), SOLO mode and TraeWork cloud tasks, the separate MIT trae-agent repo, and the 954-point telemetry analysis as required reading [glm-5.3]
  • void: recorded the Apache-2.0 fork’s proven demand (948-point thread) against the stall (last release v1.3.4 April 2025, last push June 2026, a Dec 2025 “after Void slowed down” Show HN); status set to dormant-leaning [glm-5.3]
  • agentic-coding-tools-landscape: extended the Surfaces section with Antigravity and the four-way tail (Kiro, Trae, OpenChamber, Void), bumped as-of date and updated field to 2026-08-24 [glm-5.3]
  • _index.md: Surfaces category now lists ten notes alphabetically [glm-5.3]
  • The owner-requested feature-matrix tables (implemented/not-implemented comparison articles per category) remain queued for a follow-up run, deliberately deferred until the Surfaces set was final [glm-5.3]
  • Verification: all internal link targets across the section’s 20 files exist on disk (including ../../six-months-with-openchamber and ../../send-implementation-not-issue), all cited URLs fetched 200 this run (two HN items initially 429 from burst throttling, both 200 on retry), front matter parses, no em-dashes, no banned terms, one sentence per line [glm-5.3]

2026-08-24 (research index seeding, all remaining categories + queue essays)
#

  • Research index: seeded the eight remaining queue categories (Orchestration, Protocols, Context engines, Skills, Retrieval, Memory, Executions, Hybrid execution) with 34 notes, each citing 5+ sources fetched and verified during its own run [glm-5.3]
  • Orchestration: claude-squad, conductor, crystal, dmux, emdash, vibe-kanban; crystal recorded deprecated 2026-02 for Nimbalyst and vibe-kanban community-maintained since Bloop shut down 2026-04, both kept as death records [glm-5.3]
  • Protocols: mcp, acp, agents-md, a2a [glm-5.3]
  • Context engines: augment-code, greptile, repomix, sourcegraph-code-context; the rumored Sourcegraph “Mercer” agent was unverifiable (404, no primary-source mention) and is not asserted anywhere [glm-5.3]
  • Skills: anthropic-agent-skills, opencode-skills-and-plugins, skills-sh, agent-skills-open-standard [glm-5.3]
  • Retrieval: llamaindex, langchain, semantic-code-search, tree-sitter-chunking [glm-5.3]
  • Memory: mem0, letta, zep, file-based-agent-memory [glm-5.3]
  • Executions: claude-code-hooks, copilot-automations, github-agentic-workflows, n8n [glm-5.3]
  • Hybrid execution: openai-structured-outputs, anthropic-structured-outputs, instructor, outlines [glm-5.3]
  • Queue item 2 published: model-selection-for-coding-tasks (essay, task-class tiers plus per-token economics, links all eight harness notes) [glm-5.3]
  • Queue item 3 published: context-management-patterns (essay, pattern families grounded in harness docs, links the context engine and retrieval notes) [glm-5.3]
  • queue.md: items 2 and 3 marked done [glm-5.3]
  • _index.md: added both essays and the eight new categories; the five notes from the surfaces-expansion run earlier today were already indexed there, OpenChamber keeps its Surfaces placement from that run [glm-5.3]
  • agentic-coding-tools-landscape: left untouched; cross-linking the 34 new notes into the tracker is left to the next daily refresh as its own focused pass [glm-5.3]
  • Process: notes were written by 12 parallel opencode runs (two relaunched after fetch rate-limit deaths); every run was constrained to create only its own new note files and run no git commands [glm-5.3]
  • Verification: all 36 new files carry the mandatory tags with no type field, 5+ references each, no em-dashes or banned terms, and every internal link target across the section exists on disk [glm-5.3]
  • Not committed: all changes left in the working tree for owner review [glm-5.3]

2026-08-24 (comparison matrices, owner-requested)
#

  • Owner request from earlier in the day, delivered this run: created the feature-matrix articles (implemented/not-implemented comparison tables per category) as two new pages, harness-feature-matrix (8 harnesses x 11 feature rows) and surface-feature-matrix (10 surfaces x 11 feature rows), with a four-state legend (yes/no/partial/unverified) and every cell traced to sources in the notes or references [glm-5.3]
  • Cell verification fetches this run: cursor.com/docs/settings/api-keys (200, Cursor BYOK), antigravity MCP docs (present), opencode.ai/docs/providers (Ollama and LM Studio, local models) and /docs/mcp-servers (200, after /docs/mcp moved), JetBrains/junie repo MCP mentions, docs.trae.ai/ide/mcp and /ide/agent-rules (Trae MCP and 40 AGENTS.md mentions), kiro.dev/docs MCP mentions, zed.dev/docs/ai/agents AGENTS.md mention, cursor.com/docs/context/rules (57 AGENTS.md mentions); Antigravity AGENTS.md support found no mention and is marked unverified rather than asserted absent [glm-5.3]
  • Reading the trees: 12 parallel-run notes from the previous entry were committed by the owner as 5b2679e7 before this run; the working tree was clean at sync, one transient status race on model-selection-for-coding-tasks/index.md resolved to identical content [glm-5.3]
  • _index.md: added a Comparison matrices subsection under Essays and trackers listing both pages [glm-5.3]
  • Verification: all internal link targets across the section’s 59 index files exist on disk (the 34 notes from 5b2679e7 plus both matrices), front matter parses with no type field, mandatory tags present, no em-dashes, no banned terms, one sentence per line in prose [glm-5.3]
  • Style fix: emdash note line 61 used “laptop-shaped” (“shape” is a banned term repo-wide); reworded to “assumes the work happens on your laptop”, caught by the pre-commit style sweep that ran alongside the matrix commit [glm-5.3]

2026-08-24 (cmux note, owner-requested)
#

  • Owner request in chat: cover cmux; research found three products by that name, so the note covers the dominant one (manaflow-ai/cmux, the libghostty macOS terminal for parallel agents, about 26.4k stars) and disambiguates the other two inside the note (craigsc/cmux “tmux for Claude Code”, and the separate October 2025 Show HN Coding Agent Multiplexer), plus the unrelated Go connection multiplexer [glm-5.3]
  • cmux: recorded the attention-routing thesis (notification rings and unread panel as the actual product), the open-core split (GPL-3.0-or-later terminal free, cloud VMs and CodeRouter at $24/month from cmux.com/pricing fetched this run), the fast-cadence evidence (18 releases in two days at launch, v0.64.x, nightly), the three-contributor concentration under Manaflow Inc., and the August 2026 SSH-freeze friction thread [glm-5.3]
  • _index.md: cmux added to the Orchestration category, alphabetically after claude-squad [glm-5.3]
  • Placement decision: Orchestration rather than Surfaces, matching dmux and claude-squad (terminal-family parallel-agent tools) rather than the editor and environment entries [glm-5.3]
  • Verification: all 6 cited URLs fetched 200 this run (one HN item 429 on burst, 200 on retry), all internal link targets exist on disk, front matter parses, no em-dashes, no banned terms, one sentence per line [glm-5.3]

2026-08-24 (model selection guide, Kimi and GLM coverage)
#

  • model-selection-for-coding-tasks: added the Kimi and GLM lineups per owner request; all four new pricing/provider URLs fetched and verified during this run (Moonshot K3 and K2.7 Code pricing pages, Z.ai pricing page, OpenCode providers directory) [glm-5.3]
  • New section “The challengers reset the price floor”: GLM-5.3/5.2/5.1 at $1.40/$4.40 with GLM-5 at $1/$3.20, kimi-k2.7-code at $0.95/$4.00 with 256K context and a 2x-price HighSpeed variant, kimi-k3 at $3/$15 with flat-price 1M context and always-on reasoning, GLM-4.7-Flash free; disagreeable claim added that the rational token-payer default is now a challenger model through OpenCode [glm-5.3]
  • Table extended with three rows; the cached-reads bullet corrected from “one tenth everywhere” to one tenth at the big three and Kimi K3 versus about one fifth at GLM and K2.7 Code; the whole-repo bullet now records Kimi K3 as the second flat-1M option [glm-5.3]
  • Decision guide: added price-floor BYOK and whole-repository-read bullets; re-verify list widened from three to five provider pricing pages [glm-5.3]
  • Verification: no broken internal links, front matter parses, no em-dashes or banned terms, one sentence per line, references now 10 [glm-5.3]

2026-08-24 (model selection guide, DeepSeek V4 coverage)
#

  • model-selection-for-coding-tasks: added the DeepSeek V4 lineup per owner request; the pricing page fetched and verified during this run [glm-5.3]
  • v4-pro at $1.32/$3.96 peak with all hours outside weekday 01:00-04:00 and 06:00-10:00 UTC billed at half ($0.66/$1.98 off-peak), v4-flash at $0.44/$1.32 peak and $0.22/$0.66 off-peak, both flat 1M context with 384K max output, thinking mode default, cache hits at $0.044/$0.014 (about one thirtieth, the deepest in the guide), and an Anthropic-format endpoint for Claude-speaking harnesses [glm-5.3]
  • Disagreeable claim updated to name deepseek-v4-pro in the default-loop trio; the Gemini Flash “cheapest input” line corrected (v4-flash beats it with no expiry); decision-guide bullets and the re-verify list (now six providers) updated; off-peak scheduling added to What to Do Next [glm-5.3]
  • Verification: no broken internal links, no em-dashes or banned terms, one sentence per line, references now 11 [glm-5.3]

2026-08-24 (feature matrices, remaining categories, owner-requested)
#

  • Owner request in chat: extend the feature-matrix treatment to every remaining research index category; created eight matrices in the harness/surface format (intro thesis, legend, table, reading, choosing, see also, references), each cell tracing to its linked note [glm-5.3]
  • orchestration-feature-matrix: seven tools (claude-squad, cmux, conductor, crystal, dmux, emdash, vibe-kanban) with the death and orphan stories carried into the reading section [glm-5.3]
  • protocols-feature-matrix: thesis that the four stack rather than compete and adoption falls up the stack (near-universal MCP, 60k-project AGENTS.md, rising ACP, no native A2A speaker among the harnesses); consolidations recorded (IBM ACP into A2A, Junie protocol into ACP) [glm-5.3]
  • context-engines-feature-matrix, skills-feature-matrix, retrieval-feature-matrix, memory-feature-matrix, executions-feature-matrix, hybrid-execution-feature-matrix: four columns each, rows adapted per category (guarantee mechanism for hybrid execution, memory model and lock-in for memory, trigger and execution location for executions) [glm-5.3]
  • _index.md: all eight matrices listed under Comparison matrices [glm-5.3]
  • Process: written by 8 parallel opencode runs (orchestration relaunched once after a fetch rate-limit death); reference URLs inherit verification from the category notes except where a run fetched fresh; runs created only their own file and ran no git commands [glm-5.3]
  • Verification: all 8 files carry mandatory tags with no type field, 6-8 references each, no em-dashes or banned terms, and every internal link target across the section exists on disk [glm-5.3]

2026-08-24 (daily refresh, harnesses re-verification + three entrants)
#

  • Stalest category computed as Harnesses (amp at 2026-08-22, the rest 08-22/08-23); re-verified all 8 member notes against live sources [glm-5.3]
  • Link check: 39 of 40 reference URLs fetched 200 this run; the fortieth (thereallo.dev, cited by claude-code) still returns a Cloudflare challenge to automated fetches, same behavior as the previous two runs, verified 2026-08-22 [glm-5.3]
  • aider: stall re-confirmed via the GitHub API (pushed_at still 2026-05-22, latest tagged release still v0.86.0 from 2025-08-09); left untouched [glm-5.3]
  • amp: pricing tiers re-verified unchanged; recorded the widened linked-subscription set (X Premium+/SuperGrok and SpaceX AI alongside ChatGPT, per the pricing page and homepage fetched this run); verified-as-of line and updated field bumped [glm-5.3]
  • claude-code: repository scale refreshed to about 142.8k stars and about 15.2k open issues and pull requests as of 2026-08-24 (GitHub API); the old “5k+ open issues” materially understated the tracker [glm-5.3]
  • codex: repository scale refreshed to about 116.7k stars and 17.8k forks as of 2026-08-24 (GitHub API); pricing page re-verified unchanged [glm-5.3]
  • crush, gemini-cli, junie, opencode: re-verified with no material change (stars within stated bounds, releases current, all links live); left untouched per the no-change precedent [glm-5.3]
  • New-entrant scan added three harness notes, each citing 6-7 sources all fetched this run: cline (about 66.8k stars, VS Code-extension-born runtime now CLI+kanban+SDK, 5,083,914 marketplace installs), goose (about 53.4k stars, Block’s Rust agent moved to the Agentic AI Foundation at the Linux Foundation on 2026-04-07), qwen-code (about 27.3k stars, Alibaba’s Gemini CLI fork with the 2,000 requests/day launch free tier) [glm-5.3]
  • cline: ClinePass ($9.99/month open-weights subscription across Z.ai, Moonshot, DeepSeek, MiniMax, MiMo, Qwen) and the not-open-sourced JetBrains plugin recorded; vendor 8M+ install claim checked against the marketplace’s 5.1M [glm-5.3]
  • goose: the AAIF move announcement and the ACP subscription-reuse path recorded; the “60% of Block” adoption figure marked as third-party course-site sourcing, not Block’s own [glm-5.3]
  • qwen-code: the April 2026 Taiwan-correction behavior report and Alibaba-hosted install/OAuth routing recorded as cautions; the README’s Claude Code parity table marked vendor-asserted [glm-5.3]
  • agentic-coding-tools-landscape: independent-tail section extended with the three entrants [glm-5.3]
  • _index.md: Harnesses category now lists 11 notes alphabetically [glm-5.3]
  • Queue: items 2 and 3 done previously; the standing ? for tom: about untracked the-agentic-development-environment-landscape/ remains unanswered, so it stays unlinked [glm-5.3]
  • No self-directed essay this run: the entrant scan plus re-verification filled the day, and nothing else clears the quality bar; extending harness-feature-matrix with the three new columns is left for a focused pass [glm-5.3]
  • Verification: all 74 section files link-checked with 0 broken internal targets, front matter parses as YAML with no type field, mandatory tags present, no em-dashes or banned terms, one sentence per line [glm-5.3]

2026-08-25 (owner edit, rules change)
#

  • AGENTS.md: owner-directed in chat; added a Comparison matrices section requiring every feature matrix to update in the same run its category’s membership changes (new note, move, or retirement), with new-category seeding allowed to defer the matrix only if the deferral is logged [owner, glm-5.3]
  • AGENTS.md: daily refresh step 6 now includes the matrix sync alongside the _index.md update, and step 4 verification now checks that every feature matrix matches its category’s current membership [owner, glm-5.3]
  • No article content touched this run; the harness-feature-matrix columns for cline, goose, and qwen-code (added 2026-08-24) remain the outstanding work the new rule governs, for the next daily refresh [glm-5.3]

2026-08-25 (matrix refresh, owner-requested)
#

  • Owner request in chat: refresh all ten feature matrices; ran the membership check from the new AGENTS.md rule, and exactly one matrix lagged its category, harness-feature-matrix missing the three harness notes added 2026-08-24; the other nine (surfaces, orchestration, protocols, context engines, skills, retrieval, memory, executions, hybrid execution) each link every current category member, so they were left untouched [glm-5.3]
  • harness-feature-matrix: extended from 8 to 11 columns (cline, goose, qwen-code inserted alphabetically), all 33 new cells traced to sources fetched this run; as-of date and updated field bumped to 2026-08-25 [glm-5.3]
  • Cell verification fetches this run: cline rules docs (native AGENTS.md plus .clinerules, auto-detected .cursorrules and .windsurfrules; read via the docs mdx source after the JS-rendered page came back empty), goose using-goosehints guide (CONTEXT_FILE_NAMES default [“AGENTS.md”, “.goosehints”] with nested loading), goose sitemap plus subagents, hooks, plugins, and experimental VS Code extension pages, qwen settings page (QWEN.md convention, 18 mentions, and experimental.cron default false for in-session cron tools) [glm-5.3]
  • Notable cells: cline reads AGENTS.md natively (unlike the 2025 extension era), goose reuses Claude/ChatGPT/Gemini subscriptions via ACP, qwen’s scheduled runs marked ~ experimental because the cron tools ship default-off, all three entrants scored ✗ on cloud execution [glm-5.3]
  • Reading and Choosing sections updated for eleven harnesses: open-client groups now seven auditable clients, local-model purists list widened, the everyone-lacks-a-different-feature thesis extended to the three entrants, one new choosing bullet for AGENTS.md rules portability (cline, goose, opencode) [glm-5.3]
  • References extended from 6 to 9 with the three new column sources [glm-5.3]
  • Verification: matrix table is 13 rows x 12 cells with consistent column counts, all 74 section files link-checked with 0 broken internal targets, front matter parses with no type field, no em-dashes or banned terms, one sentence per line [glm-5.3]

2026-08-25 (matrix viewer component, owner-requested)
#

  • Owner request in chat: build a custom component for reading the growing matrices; implemented in the site repository (blog.tomrochette.com), outside this content repo, so nothing under agents/ changed except this log and the methodology note [glm-5.3]
  • Site changes: a table render hook wrapping every Markdown table in a scrollable region (pinned header row and label column, zebra rows; tables with six or more columns tagged wide), CSS in the auto-bundled custom.css using the theme’s light and dark neutral variables, and a deferred vanilla-JS layer adding column-picker chips and a row filter to wide tables only, with per-browser persistence and guards preventing hiding the label column or the last visible column [glm-5.3]
  • Content stays plain Markdown tables: the cell-traces-to-note convention, CI link checks, and readable diffs are unaffected; methodology.md updated with a Matrix presentation section documenting the split [glm-5.3]
  • Verification: site built with Hugo v0.166 (643ms), matrix pages emit the wide wrapper with role, aria-label, and tabindex, small tables emit the plain wrapper, table links render unchanged, the script passes node –check, and the minified fingerprinted asset loads deferred [glm-5.3]

2026-08-25 (matrix viewer, value-based column filter, owner-requested)
#

  • Owner request in chat: let readers hide columns by a row’s value, example given: keep only surfaces supporting local models; implemented in the site repository’s tables.js, no content files changed beyond this log and the methodology sentence [glm-5.3]
  • Behavior: clicking a body cell in a wide table keeps only the columns whose cell in that row shares the clicked cell’s state glyph (the four legend states), Shift-click keeps the opposite, clicking the same cell again or the filter chip clears it; the active filter shows as a chip naming the row and state, the active row’s label is emphasized, and the filter intersects with the persisted column-picker selection and the row search; the value filter itself is transient and not persisted [glm-5.3]
  • Verification: script logic smoke-tested with a minimal DOM harness in Node (initial all-visible, filter by a \u2717 cell keeps only matching columns, second click clears, Shift-click keeps the complement), node –check passes on source and built bundle, Hugo build clean [glm-5.3]

2026-08-25 (daily refresh, surfaces re-verification)
#

  • Stalest category computed as Surfaces (cursor, jetbrains, vscode-copilot, windsurf, zed at 2026-08-23, never targeted by a refresh run; Harnesses tied on oldest date but was refreshed 2026-08-24); re-verified all 10 member notes against live sources [glm-5.3]
  • Link check: all 45 cited URLs fetched 200 this run (25 vendor/docs pages, 20 Hacker News threads spaced to avoid burst throttling), plus repository state via the GitHub API; windsurf.com still returns 429 to every automated fetch, same as documented, and is not cited [glm-5.3]
  • cursor, vscode-copilot, windsurf, jetbrains, zed, antigravity, void: re-verified with no material change (stars within stated bounds, pricing tiers unchanged, all links live); left untouched [glm-5.3]
  • kiro: corrected the status attribution, the release cadence lives on the vendor changelog (weekly, latest 2.19.1 on August 21, 2026, fetched this run) not the GitHub tracker, which has been quiet since June 22, 2026; changelog added as a reference; verified-as-of dates and updated field bumped [glm-5.3]
  • trae: recorded the trae-agent stall, last commit February 5, 2026 and no tagged release ever published (GitHub API, as of 2026-08-25); the open-source sidecar bullet now reads as a snapshot, not a maintained dependency; verified-as-of dates and updated field bumped [glm-5.3]
  • openchamber: volatile numbers refreshed to about 9.2k stars and 964 forks as of 2026-08-25, v1.20.0 still the latest release, commits landing the morning of verification; verified-as-of dates and updated field bumped [glm-5.3]
  • New-entrant scan: PearAI (765 stars, slowing since June) and Qoder (no meaningful public footprint) fail the credibility bar, no other credible surface entrant found; no note added [glm-5.3]
  • Category fit re-checked: all 10 members stay in Surfaces (OpenChamber keeps its placement, Windsurf stays as the pivot record, Void stays as the dormancy record); surface-feature-matrix verified to list all 10 current members, unchanged [glm-5.3]
  • Queue: the standing ? for tom: about untracked the-agentic-development-environment-landscape/ remains unanswered, so it stays unlinked; no other actionable item [glm-5.3]
  • No self-directed essay this run: the category re-verification plus three note corrections filled the day, and nothing else clears the quality bar [glm-5.3]
  • Verification: all 70 section files pass front matter parsing with no type field, mandatory tags present, no em-dashes, no banned terms or compounds, one sentence per line, and 0 broken internal link targets [glm-5.3]

2026-08-26 (research note, cognee, owner-prompted)
#

  • Owner prompt in chat (“no new topic to investigate?”): ran an entrant scan beyond the surfaces category; candidates verified via the GitHub API were Roo Code (24.3k stars but no pushes since 2026-05-15, needs its own investigation), Continue (35.6k stars, active), spec-kit (131k stars, active), and cognee (30.3k stars, active); none covered anywhere in the corpus [glm-5.3]
  • Research index: added the cognee note to Memory, citing 7 sources all fetched this run (repo, landing, pricing, docs, PyPI JSON API, two Show HN threads); recorded the thin HN footprint (9- and 6-point threads) as the community signal, the whole-engine Apache-2.0 licensing, v1.5.3 released August 23, 2026, and the flat cloud pricing ($2.50 per 1M tokens plus $5 per workspace) [glm-5.3]
  • memory-feature-matrix: extended from 4 to 5 columns (cognee inserted second, after Files), contradiction-handling and audit-trail cells marked ? rather than asserted, reading section gains the whole-engine-open exception sentence, choosing section gains the flat-billing bullet, references extended to 10; updated field bumped [glm-5.3]
  • _index.md: cognee added to the Memory list, alphabetically first [glm-5.3]
  • Date correction: the previous entry in this log is dated 2026-08-25 but that run and this one both occurred on 2026-08-26; the as-of dates inside kiro, trae, and openchamber from that run were also written as 2026-08-25 when the fetches happened on 2026-08-26; recorded here rather than churning three files by one day [glm-5.3]
  • Roo Code stall, Continue, and spec-kit left as scan candidates for future runs: one note done well beats three done thin [glm-5.3]
  • Verification: all 71 section files pass front matter parsing with no type field, mandatory tags present, no em-dashes, no banned terms, one sentence per line, and 0 broken internal link targets; all 7 new reference URLs fetched 200 this run [glm-5.3]

2026-08-26 (research note, spec-kit, new category, owner-prompted)
#

  • Owner prompt in chat (“would it be relevant to cover spec options, spec-kit, Kiro?”): Kiro is already covered as a Surfaces note centered on spec mode, so the gap is spec-kit and the spec-driven development movement it anchors; no existing article in the section or tracked corpus covers it (the-future-of-code-review and send-implementation-not-issue engage adjacent ideas and are cross-linked) [glm-5.3]
  • Research index: seeded the Spec-driven development category with the spec-kit note, citing 8 sources all fetched 200 this run (repo plus GitHub API, official docs, GitHub launch post, lead maintainer’s anniversary post, marmelab critique, three HN threads including the 225-point waterfall-strikes-back one) [glm-5.3]
  • spec-kit: recorded first commit August 21, 2025, v1.0.0 and v1.0.1 both shipped August 21, 2026, 131,489 stars and about 11.8k forks as of 2026-08-26, original creators moved on per the anniversary post, the 30-plus integration surface, and the constitution/specify/plan/tasks/implement/converge workflow [glm-5.3]
  • Matrix deferral, per the seeding rule: the new category starts with one member, so its feature matrix is deferred to a dedicated follow-up run once Tessl and the SDD CLI ecosystem get notes; Kiro stays in Surfaces (an IDE whose spec workflow is a feature) and is cross-referenced rather than moved [glm-5.3]
  • _index.md: new Spec-driven development category listed after Hybrid execution [glm-5.3]
  • Verification: all 72 section files pass front matter parsing with no type field, mandatory tags present, no em-dashes, no banned terms (the upstream README and anniversary post use them; the note quotes neither), one sentence per line, and 0 broken internal link targets [glm-5.3]

2026-08-26 (review-driven maintenance pass, owner-prompted)
#

  • Owner prompted a full section review; review found 7 must-fix and ~18 should-fix items, all addressed this run [x-preview-f-free]
  • goose and agentic-coding-tools-landscape: reconciled the AAIF timeline that split the section; Block contributed goose at the foundation’s formation on December 9, 2025 (per the aaif.io press release, fetched live), and April 7, 2026 is the migration-completion date per goose-docs.ai/blog; both articles now carry both dates [x-preview-f-free]
  • harness-feature-matrix: replaced dead cline-rules docs URL with docs.cline.bot/features/cline-rules (fetched 200); added Junie to the local-models-plus-MCP list so guidance matches its own cells; expanded the rules-portability list to include Amp, Codex, and Crush per the AGENTS.md row; added first-person reading line [x-preview-f-free]
  • surface-feature-matrix: added first-person reading line per writing rules [x-preview-f-free]
  • retrieval-feature-matrix: fixed self-contradiction claiming legacy LangChain code-splitter docs no longer resolve while citing that same URL; reworded to frozen documentation (URL verified 200) [x-preview-f-free]
  • aider: replaced mis-cited HN reference (43708025 is the Codex CLI launch thread) with HN 43672712 “Wasting Inferences with Aider” (verified via Algolia API), keeping the note at five sources; aligned header verification date to 2026-08-23; first-person line in Bottom line [x-preview-f-free]
  • codex: qualified the overbroad only-open-source-CLI thesis to what distinguishes it from Gemini CLI (an Apache-2.0 CLI individuals can still run); aligned repo-stats reference date; first-person split in Bottom line [x-preview-f-free]
  • gemini-cli: mirrored fix, strength bullet now claims first rather than only; linked Antigravity sibling note in Compared to; first-person admission on past recommendations [x-preview-f-free]
  • claude-code: grounded the March 2026 sourcemap-leak caution with the dated March 31 event plus HN 47584540 as reference (2,095 points, verified via API); swapped the 403-returning reallo.dev citation for its archive.org snapshot (fetched 200); first-person line in Bottom line [x-preview-f-free]
  • cline: linked the previously named-but-unreferenced Ask HN token-usage thread (48525711, verified via Algolia) both inline and in references; first-person line in Bottom line [x-preview-f-free]
  • amp: first-person line in Bottom line per writing rules [x-preview-f-free]
  • n8n: grounded the community-strengths claim by adding both cited HN threads as references (launch 21191676 at 728 points, Series C 45525336 at 235 points, point counts verified via API) [x-preview-f-free]
  • claude-squad: converted malformed plain-text Compared-to paths to markdown links; rewrote two unsourced ecosystem claims into git-semantics and product facts; dropped the meta-label before the disagreeable claim; added updated field [x-preview-f-free]
  • file-based-agent-memory: corrected the HN 46294274 label to match the letta note (Letta Code launch thread); replaced the loose /doctor-style trimming analogy with a grounded pruning sentence [x-preview-f-free]
  • model-selection-for-coding-tasks: clarified the $100 crossover math (50M raw input tokens at $2/M versus about 20M blended tokens at agentic output ratios) so it no longer reads against its own table; added DeepSeek to the challengers mention [x-preview-f-free]
  • agents-md: fixed the our-note antecedent to name the July 2026 proxy study and the Claude Code note; indented the ETH Zurich continuation under its Cautions bullet; added updated field [x-preview-f-free]
  • agent-skills-open-standard, opencode-skills-and-plugins, skills-sh: restored the mandatory Not-for-Y bottom-line clauses; agent-skills-open-standard additionally gained Anthropic’s untrusted-skills engineering post as its critical source [x-preview-f-free]
  • semantic-code-search: stated the missing independent-footprint signal explicitly per the citation standard instead of implying third-party evidence exists; added updated field [x-preview-f-free]
  • cursor: resolved the Grok ownership contradiction (xAI listed as third-party while calling Grok Cursor’s own) by tying Grok to the SpaceXAI partnership [x-preview-f-free]
  • _index.md one-liners refreshed where they lagged revised content: harness matrix now eleven harnesses verified 2026-08-25, memory matrix four services, codex description matches its qualified thesis [x-preview-f-free]
  • Every touched article gained llm=x-preview-f-free alongside glm-5.3 (never removed) and updated set to 2026-08-26; untouched notes deliberately left without updated pending an owner decision on backfilling [x-preview-f-free]
  • Verification: 71 section files pass front matter parsing, zero broken internal link targets, zero em-dashes or banned terms, one sentence per line; all newly added external URLs fetched 200 or verified via the HN Algolia API this run [x-preview-f-free]

2026-08-26 (daily refresh, orchestration re-verification + extension-generation exits)
#

  • Stalest category computed as Orchestration: its notes sit at creation date 2026-08-24 with no updated field and it was never targeted by a refresh run, while Harnesses (08-23/08-24 notes) and Surfaces (08-23 notes) both had dedicated re-verification runs since; re-verified all 7 member notes against live sources [glm-5.3]
  • Link check: all 39 orchestration reference URLs verified this run (34 fetched 200 directly; HN throttled this IP to 429 on every direct fetch, so the five cited threads were verified via the Algolia items API, titles, points, and dates all matching); conductor.build now 307-redirects to www.conductor.build and emdash.ai to emdash.com, both resolving 200 [glm-5.3]
  • claude-squad, cmux, conductor, crystal, dmux, emdash: re-verified with no material change (stars within stated bounds, crystal still dead with last push 2026-02-26, conductor changelog still at 0.82.0, cmux pushing the morning of verification); left untouched per the no-change precedent [glm-5.3]
  • vibe-kanban: sharpened the community-maintenance claim on evidence, no push since 2026-04-24, npm still at 0.1.44, open issues 383 to 533 between 2026-08-24 and 2026-08-26 (GitHub API); status bold line rewritten, the nobody-is-paid caution extended, verified-as-of and updated bumped [glm-5.3]
  • emdash: references moved to the canonical emdash.com domain (the .ai host 301-redirects, both fetched 200 this run); verified-as-of and updated bumped [glm-5.3]
  • New-entrant scan for Orchestration: Backlog.md verified credible (6.5k stars, active) but is process/task tooling rather than a parallel-agent dashboard, so it is recorded here as a candidate for the Spec-driven development category instead of joining this one; no orchestration entrant cleared the bar [glm-5.3]
  • Queue item 1 (standing research index): wrote the two scan candidates left by the 2026-08-26 cognee run, both Surfaces death records with 7-8 sources each, all fetched 200 or Algolia-verified this run [glm-5.3]
  • continue: recorded the Cursor acquisition (announced 2026-06-15 per the homepage banner and HN 48548758), the read-only README with final 2.0.0 (telemetry and auth removed, tagged 2026-06-19), 35,644 stars, 4,001,358 marketplace installs at a 3.26 rating as of 2026-08-26, and the quiet Ask HN funeral thread; the earlier scan’s “active” reading corrected here [glm-5.3]
  • roo-code: recorded the scheduled sunset (announcement 2026-04-20, archived because roocode.com now sells the pivot product; support ended 2026-05-15 with Cloud and Router shut down and refunds), last release v3.54.0 and last push both 2026-05-15, 24,325 stars, 1,956,325 installs at a 4.72 rating, the AGENTS.md-reading rules system, and the pivot to roomote.dev with its “we don’t believe IDEs are the future of coding” rationale [glm-5.3]
  • surface-feature-matrix: extended from 10 to 12 columns (Continue and Roo Code inserted alphabetically) in the same run per the matrix rule; 22 new cells traced to sources fetched this run (Continue rules docs for the AGENTS.md cell, Roo custom-instructions docs for its AGENTS.md and .roorules cells, Roo MCP overview, Ollama guides both sides); reading section gains the death-records paragraph; verified-date line and references extended [glm-5.3]
  • agentic-coding-tools-landscape: surfaces tail extended with the two 2026 exits and the Cline consolidation sentence; as-of line qualified [glm-5.3]
  • _index.md: Continue and Roo Code added to Surfaces alphabetically, surface matrix one-liner now twelve surfaces verified 2026-08-26, vibe-kanban one-liner updated to the orphaned framing [glm-5.3]
  • Matrix membership check: surface-feature-matrix columns match the 12 on-disk surfaces exactly; orchestration-feature-matrix still matches its 7 members; queue unchanged (the standing ? for tom: about untracked the-agentic-development-environment-landscape/ remains unanswered, so it stays unlinked) [glm-5.3]
  • No self-directed essay this run: the category re-verification plus two death-record notes with their matrix columns filled the day at quality [glm-5.3]
  • Verification: 73 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, no em-dashes or banned terms in changed files, one sentence per line, and every URL cited by new or edited content fetched 200 or was verified via the HN Algolia API during this run [glm-5.3]

2026-08-27 (owner-prompted, task management category + gastown placement)
#

  • Owner request in chat: collect the LLM task trackers (backlog.md, beads named) and place Gastown; research found the task-tracker layer to be three credible tools plus one orchestration system, all uncovered in the corpus [glm-5.3]
  • Research index: seeded the Task management category with three notes, each citing 6-7 sources fetched 200 or verified via the HN Algolia API this run: backlog-md, beads, task-master [glm-5.3]
  • beads: recorded the gastownhall org rename from steveyegge (the old API path 301-redirects), the Dolt-backed architecture (embedded or server, sync via bd dolt push/pull), 26,630 stars and v1.2.2 of 2026-08-15, the missing Show HN launch thread stated as the community signal, and the ecosystem as caution (beads_rust freezing classic beads at 1,062 stars, the 84-point markdown-replacement thread as the critical source) [glm-5.3]
  • task-master: recorded the productization, the repo quiet since 2026-04-28 with the last release 0.43.1 on 2026-03-31 while npm still does 82,718 downloads a month, the license now MIT with Commons Clause Condition v1.0 (fetched from the LICENSE file), and the Hamster (Wheel Go Fast, Inc.) commercial successor at Free and $40 per creator per month [glm-5.3]
  • backlog-md: recorded the three-checkpoint review model (spec, plan, code), the dogfooding claim, terminal kanban plus local web UI, 6,551 stars with a 254-point launch thread and 50,970 npm downloads a month [glm-5.3]
  • gastown: placed in Orchestration per the owner’s question, because the repo self-describes as a multi-agent workspace manager (Mayor, polecats, worktree hooks, Refinery Bors-style merge queue) rather than a task tracker; recorded 17,801 stars, v1.2.1 of 2026-06-06, the HN footprint (354, 403, 253, 113-point threads), Maggie Appleton’s field analysis as the observational source (costs, meme coin, design-as-bottleneck), and the credits-governance issue 3649 [glm-5.3]
  • orchestration-feature-matrix: extended from 7 to 8 columns (Gas Town inserted alphabetically after Emdash) in the same run per the matrix rule, all 10 cells traced to the gastown README and the Appleton analysis fetched this run; reading gains the merge-queue exception line, choosing gains the 20-agent bullet, the emdash reference moved to the canonical emdash.com, updated bumped [glm-5.3]
  • Task management category matrix: deferred per the seeding rule (three members, one shared-rows design pending); the deferral is recorded here [glm-5.3]
  • agentic-coding-tools-landscape: orchestration paragraph gains the Gas Town exception sentence naming the tracker layer, updated bumped [glm-5.3]
  • _index.md: new Task management category (three notes, alphabetical) listed after Spec-driven development; Gas Town added to Orchestration; orchestration matrix one-liner now eight tools [glm-5.3]
  • Verification: all 78 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, no em-dashes or banned terms (two slips caught and reworded mid-run), and every cited URL fetched 200 or Algolia-verified this run [glm-5.3]

2026-08-27 (owner-prompted, paperclip and the Control planes category)
#

  • Owner request in chat: create a category for Paperclip AI if not already covered; the corpus mentioned it only inside an untracked owner draft (i-specify-open-source-projects-i-dont-maintain), so nothing existed under agents/ and the category was created [glm-5.3]
  • Research index: seeded the Control planes category with the paperclip note, citing 7 sources all fetched 200 or Algolia-verified this run (repo and API, homepage, docs, releases API, the 3-point Show HN, the OpenClaw API for the comparison) [glm-5.3]
  • paperclip: recorded the company-versus-employee framing (“if OpenClaw is an employee, Paperclip is the company”), the heartbeat-and-checkout execution model with 409 conflict semantics, the four pillars, the sandbox providers, 79,465 stars and 14,557 forks since 2026-03-02 with 5,417 open issues, the date-versioned v2026.824.1 release, the cloud waitlist, the near-empty HN footprint stated as the distribution signal, and the quickstart’s trusted-loopback default flagged as the main caution [glm-5.3]
  • Placement decision: a new Control planes category rather than Orchestration or Executions, because Paperclip orchestrates business goals (org charts, budgets, approvals, multi-company) rather than repo worktrees or single scheduled triggers; OpenClaw stays unnoted (already covered by two corpus articles) and is named inside the note instead, per the gemini-cli/Antigravity grouping precedent [glm-5.3]
  • Control planes matrix: deferred per the seeding rule (one member); the deferral is recorded here [glm-5.3]
  • _index.md: Control planes category listed after Task management [glm-5.3]
  • Verification: all 79 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets (learnings-from-openclaw linked as a committed corpus article, openclaw-triage-pipeline deliberately unlinked because it is untracked), no em-dashes or banned terms, one sentence per line [glm-5.3]

2026-08-27 (owner-prompted, matrices for every category)
#

  • Owner request in chat: task management gets its own matrix, and AGENTS.md changes so every category list carries a comparison matrix [owner, glm-5.3]
  • AGENTS.md, Comparison matrices section (owner-directed edit, the first since the 2026-08-25 rules change): every research index category now has a companion matrix, created in the same run the category is seeded even for a single member (the single column is the scaffold the next member extends, gap named in prose); the seeding-deferral sentence is removed [owner, glm-5.3]
  • task-management-feature-matrix: created with three columns (Backlog.md, beads, Task Master) and ten rows, thesis that storage, files versus database, decides the concurrency everything else decorates; every cell traced to the member notes; the four reference URLs fetched 200 this run; Task Master’s storage cell verified against the repo tree (.taskmaster/ with tasks, config, state) and its concurrency cell honestly marked ? [glm-5.3]
  • spec-driven-development-feature-matrix: created as a single-column scaffold for spec-kit, delivering the deferral recorded on 2026-08-26 now that deferrals no longer exist; nine rows, the Tessl/OpenSpec/Kiro gap named in prose, four reference URLs fetched 200 this run [glm-5.3]
  • control-planes-feature-matrix: created as a single-column scaffold for paperclip, delivering the deferral recorded earlier today; ten rows centered on the governance, budgets, and multi-company rows that define the category against Orchestration; three reference URLs fetched 200 this run [glm-5.3]
  • _index.md: all three matrices listed under Comparison matrices [glm-5.3]
  • All thirteen research categories now have matrices: harness, surface, orchestration, protocols, context engines, skills, retrieval, memory, executions, hybrid execution, task management, spec-driven development, control planes [glm-5.3]
  • Verification: all 82 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, every matrix column set matches its category membership on disk, no em-dashes or banned terms, one sentence per line [glm-5.3]

2026-08-27 (owner-prompted, assistant runtimes: OpenClaw and the -claw variants)
#

  • Owner request in chat: cover the OpenClaw variants (nanoclaw, zeroclaw, picoclaw, etc.); all three named variants verified real and large (30.6k, 32.7k, 29.9k stars), plus the “etc.” resolved as microclaw (732, inspired by nanoclaw), SmolClaw and FemtoClaw (sub-scale curiosities kept in prose), and the clawsec security suite that names the family [glm-5.3]
  • Research index: seeded the Assistant runtimes category with four notes, each citing 5-6 sources fetched 200 or Algolia-verified this run: openclaw (the root, added as the anchor the variants are read against), nanoclaw, zeroclaw, picoclaw [glm-5.3]
  • openclaw: recorded the Gateway architecture, MIT under the OpenClaw Foundation, 387,777 stars and 81,419 forks since 2025-11-24, and the 2026 provider-restriction saga (Google 802 pts, Anthropic 1,099 pts, the same-week privilege-escalation report 514 pts, commits-mentioning-OpenClaw refusals 1,349 pts) as the category’s defining caution [glm-5.3]
  • nanoclaw: recorded the audit-first thesis (the author’s half-million-line rebuke of the root), per-agent Linux containers, OneCLI Agent Vault credential injection, skill-installed channels, Claude Agent SDK nativity with codex/opencode/ollama drop-ins, and the 533-point launch thread as “Clawdbot in 500 lines” [glm-5.3]
  • zeroclaw: recorded the compile-time ownership thesis, single Rust binary, dual MIT/Apache-2.0, about 20 providers and 30+ channels including voice and hardware tools, the Android port (303 stars), and the thin HN footprint (threads at 2-8 points) stated as the missing-community-coverage signal [glm-5.3]
  • picoclaw: recorded Sipeed’s Go-from-scratch implementation (explicitly inspired by NanoBot, 47,454 stars, not an OpenClaw fork), the $10 RISC-V boards with 10-20MB RAM footprint, the LicheeRV-Claw hardware bundle, 20k stars in 17 days, and both README banners (no production before v1.0; no official cryptocurrency, picoclaw.io only) as the attention-outpacing-governance record [glm-5.3]
  • assistant-runtimes-feature-matrix: created in the same run per the matrix rule, four columns and twelve rows with the trust-ladder thesis (ecosystem, auditable codebase, static binary, firmware image); every reference fetched this run [glm-5.3]
  • paperclip: the Compared-to OpenClaw entry now links the new note instead of the bare GitHub URL [glm-5.3]
  • control-planes-feature-matrix: the un-profiled-neighbors paragraph now points at the Assistant runtimes category instead of describing OpenClaw as note-less [glm-5.3]
  • _index.md: Assistant runtimes category added (four notes, alphabetical) and the matrix listed under Comparison matrices [glm-5.3]
  • openclaw-triage-pipeline remains untracked and unlinked, per the standing CI rule; learnings-from-openclaw is the linked corpus anchor [glm-5.3]
  • Verification: all 87 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, every matrix column set matches its category membership (14 categories, 14 matrices), no em-dashes or banned terms, one sentence per line [glm-5.3]

2026-08-27 (owner-prompted, spec-driven development alternatives)
#

  • Owner request in chat: cover the other spec-driven development alternatives; the previously named-but-unprofiled neighbors (Tessl, OpenSpec, the SDD CLI ecosystem) were researched, plus a scan that surfaced BMad; an AWS context-forge repo could not be verified under that name (GitHub 404) and is not asserted anywhere [glm-5.3]
  • The interrupted control-planes scan (previous owner prompt) got as far as identifying Hermes Agent (Nous Research, 52-pt launch thread) before this request redirected the run; Hermes stays an unlogged scan candidate for the control-planes build-out [glm-5.3]
  • Research index: Spec-driven development extended from one to four members, each citing 5-6 sources fetched 200 or Algolia-verified this run: openspec, bmad-method, tessl [glm-5.3]
  • openspec: recorded Fission AI’s delta-proposal model (/opsx:propose writes proposal, specs, design, tasks; /opsx:archive folds the delta into a living ledger), the brownfield-first philosophy, 66,398 stars since 2025-08-05, 1,610,142 npm downloads last month, MIT, and the empty HN footprint stated as the thin-verification caution [glm-5.3]
  • bmad-method: recorded the agile right-sizing thesis (small changes straight to build, deep planning for big ones) as the answer to the waterfall-strikes-back critique, the clarify-plan-build-learn loop, durable-context briefs, the module ecosystem (Builder, Test Architect, Loop for unattended epics, Game Dev), MIT under BMad Code LLC (API says NOASSERTION, LICENSE file fetched and confirms MIT), 52,371 stars since 2025-04-13, and the thin HN footprint (2-4 point threads) [glm-5.3]
  • tessl: recorded Guy Podjarny’s $125M Series A (November 2024) for the spec-centric platform, Skills on Tessl as a skills package manager (January 2026), Martin Fowler’s 128-point third-party analysis naming Kiro, Spec Kit, and Tessl the three pillars, and the 71-star quiet-since-March CLI as the deliberate thin-OSS posture; pricing left explicitly unverified [glm-5.3]
  • spec-driven-development-feature-matrix: rebuilt from the single-column scaffold to four columns (BMad, OpenSpec, Spec Kit, Tessl) in the same run per the matrix rule, eleven rows across the ownership (repo versus platform) and ceremony-sizing axes; all six references fetched this run [glm-5.3]
  • spec-kit: the Tessl line in Compared to now links the new note instead of deferring [glm-5.3]
  • _index.md: Spec-driven development lists four notes, matrix one-liner updated [glm-5.3]
  • Verification: all 90 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, every matrix column set matches its category membership (14 categories, 14 matrices), no em-dashes or banned terms, one sentence per line [glm-5.3]

2026-08-27 (owner-prompted, paperclip alternatives: the finished scan)
#

  • Owner follow-up: the control-planes scan interrupted by the spec-driven request was completed this run; candidates verified via the GitHub and HN Algolia APIs were Hermes Agent (237,130 stars), TinyAGI fka TinyClaw (3,610 stars, stalled), claw-empire (1,359, stalled since March), desplega-ai/agent-swarm (723, active), multigent (61), Cabinet (knowledge-base product, different problem), and kastra (policy enforcement, a governance slice) [glm-5.3]
  • hermes: added to Assistant runtimes rather than Control planes, because the repo is a personal agent runtime (single Python gateway, chat channels, any model), not a company orchestrator; recorded the learning loop (autonomous skill creation, self-improvement, FTS5 session search, Honcho user modeling, agentskills.io compatibility), seven execution backends with serverless idle, 237,130 stars and 47,995 forks since 2025-07-22, 36,374 open issues, the 52-point launch thread, the Grafana and Fleet ecosystem threads, and the edited plagiarism-claims issue 10232 as the stewardship caution [glm-5.3]
  • tinyagi: added to Control planes as the category’s first stall record; recorded the fka TinyClaw origin (jlia0’s Claude Code wrapper, now redirecting), the multi-team chain-and-fan-out model, the SQLite queue with dead-letter management, TinyOffice, v0.0.20 of 2026-03-26 as the final release with the last push 2026-03-30, and the consolidation reading (five months of silence while Paperclip reached 79k stars) [glm-5.3]
  • assistant-runtimes-feature-matrix: extended from four to five columns (Hermes inserted after the root) in the same run per the matrix rule; the reading section gains the learning-loop-versus-trust-ladder line [glm-5.3]
  • control-planes-feature-matrix: extended from one to two columns (TinyAGI), the single-column scaffold language replaced with the two-column consolidation comparison, and the long tail (claw-empire, agent-swarm, multigent, Cabinet, kastra) named in prose with reasons each stays unnoted [glm-5.3]
  • _index.md: Hermes listed in Assistant runtimes (five notes), TinyAGI in Control planes (two notes), the assistant runtimes matrix one-liner updated [glm-5.3]
  • Verification: all 92 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, every matrix column set matches its category membership (14 categories, 14 matrices), no em-dashes or banned terms, one sentence per line [glm-5.3]

2026-08-27 (owner-prompted, table filter controls on every matrix)
#

  • Owner follow-up on why only the harness, surface, and orchestration matrices had filtering: the site-layer viewer gated its interactive controls on the wide-table tag (six or more columns), and every other matrix has five or fewer; owner approved opening it up [glm-5.3]
  • Site repository (blog.tomrochette.com, outside this content repo): tables.js now enhances every .tbl-wrap table (column-picker chips, row search, value-based column filter) instead of only .tbl-wide ones; the wide class keeps its single remaining job, sticky header and label column for six-plus-column tables; verified via node –check, a Hugo rebuild whose tables.min bundle selects .tbl-wrap with zero tbl-wide references, and the existing guards (label column and last visible column cannot be hidden); committed and pushed as 62b574f7 [glm-5.3]
  • methodology.md: the Matrix presentation section now states every table gets the controls, with the pinned header and label column remaining the six-plus-column extra [glm-5.3]
  • Verification: methodology diff matches the shipped behavior; no other content files touched this run [glm-5.3]

2026-08-27 (daily refresh, harnesses re-verification + two entrants)
#

  • Stalest category computed as Harnesses: its oldest updated dates tie Surfaces at 2026-08-23, but Surfaces received its dedicated refresh run on 2026-08-25 while Harnesses last had one on 2026-08-24; re-verified all 11 member notes against live sources [glm-5.3-flash]
  • Link check: every reference URL of the eleven harness notes fetched 200 this run (15 HN threads via the Algolia items API with titles and points matching), except web.archive.org’s main host, which is unreachable from this environment at connection level (000 even on root); the claude-code archive snapshot was confirmed available and status 200 through the archive.org availability API [glm-5.3-flash]
  • codex: repository stats refreshed to about 119.1k stars, about 18.2k forks, and about 9.9k commits as of 2026-08-27 (GitHub API); the three doc references moved to learn.chatgpt.com because developers.openai.com/codex migrated there (target content verified on the fetched page) [glm-5.3-flash]
  • claude-code: reference canonicalized from anthropic.com/claude-code to claude.com/product/claude-code after Anthropic’s domain move; the about-15.2k open issues and pull requests claim re-checked at 15,173 via the search API [glm-5.3-flash]
  • amp: spinout link and reference canonicalized to ampcode.com/news/amp-frontier-corporation after a redirect revealed the new slug [glm-5.3-flash]
  • goose: aaif.io formation-announcement reference canonicalized to /news/linux-foundation-announces-formation-of-aaif [glm-5.3-flash]
  • aider: stall re-confirmed (pushed_at still 2026-05-22); crush, gemini-cli, junie, opencode, qwen-code, cline: re-verified with no material change, left untouched per precedent (cline’s marketplace page no longer server-renders install counts, so the 5.1M figure keeps its August 24 date) [glm-5.3-flash]
  • New-entrant scan surfaced two credible harness additions, each written to full citation standards with all sources fetched this run: kilo-code and openhands [glm-5.3-flash]
  • kilo-code: recorded the MIT feature-merge fork of Cline and Roo Code (launch-era threads March 2025), about 27k stars pushed the day of verification, the Anaconda acquisition of July 15, 2026 per the primary blog post (Kilo already in Anaconda’s product nav, kilocode.ai redirects to kilo.ai), Free/Teams $15/Enterprise pricing with the three-way cost split, and the vendor popularity claim marked unproven against observable stars [glm-5.3-flash]
  • openhands: recorded the renamed OpenDevin at about 85.3k stars MIT, v1.15.0 of August 21, 2026, the $18.8M Series A led by Madrona of November 18, 2025 plus the AMD local-agents collaboration, the Agent Canvas plus Agent Server architecture with local GUI and CLI V0 deprecated in the docs, ACP hosting of rival harnesses inside Canvas, and the thin post-rename community footprint stated as a signal (a 70-point 2022 HN thread named OpenHands is unrelated and was excluded) [glm-5.3-flash]
  • harness-feature-matrix: extended from 11 to 13 columns in the same run per the matrix rule, all 22 new cells traced to sources fetched this run (both projects’ llms.txt doc indexes: Kilo AGENTS.md support, MCP, custom subagents, cron schedules, cloud tasks, Ollama and LM Studio; OpenHands MCP, hooks, skills, plugins, automations, local-LLM pages, .openhands customization); reading and choosing sections updated for thirteen, references extended by the two doc indexes and the Anaconda post [glm-5.3-flash]
  • agentic-coding-tools-landscape: independent tail gains Kilo Code and OpenHands lines, the Roo-Cline exit sentence now names the fork descendant that merged their lineages, verified-dates line extended [glm-5.3-flash]
  • _index.md: both notes listed alphabetically in Harnesses, matrix one-liner updated to thirteen harnesses verified 2026-08-27, landscape one-liner re-dated [glm-5.3-flash]
  • Style sweep caught three pre-existing shape compounds (“subscription-shaped work” in codex, “fork and reshape” in nanoclaw, “reshape deployment thinking” in picoclaw) which were reworded; touched files gained llm=glm-5.3-flash alongside earlier tags (none removed) [glm-5.3-flash]
  • No self-directed essay this run: the category re-verification plus two entrant notes with their matrix columns filled the day at quality [glm-5.3-flash]
  • Verification: 93 article files plus section control files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, every matrix column set matches its category membership (14 categories, 14 matrices), no em-dashes or banned terms in changed files, one sentence per line [glm-5.3-flash]

2026-08-27 (correction to the daily refresh entry above)
#

  • harness-feature-matrix: the entry above claimed every touched file gained llm=glm-5.3-flash; the matrix itself was missed at first, caught on self-check and added in commit c21a253b [glm-5.3-flash]

2026-08-29 (owner-prompted, software factory category)
#

  • Owner request in chat: introduce a Software factory category and decide whether disler/super-simple-software-factory (SSSF) fits there; research confirmed it is the canonical instantiation of the concept and placed it as the founding member [big-pickle]
  • Placement decision: a new Software factory category rather than Hybrid execution, Skills, or Spec-driven development, because SSSF’s center is a repeatable code-owned SDLC pipeline packaged as a skill, not structured-output orchestration, a skill format, or a spec artifact; it joins the corpus’s software-factory tag and its Code Factories essays (../../code-factories-factorio etc.), which it cross-links [big-pickle]
  • super-simple-software-factory: research note created per the skeleton, citing 5 sources fetched or verified this run; recorded 764 stars and 189 forks since creation 2026-08-02 with last push 2026-08-04, single commit on main, no releases, the example branch with a stamped demo app, MIT, the code-owns-the-loop thesis over bounded agent phases, typed JSON envelopes and gates, same-session correction, WAL SQLite trace, the pi-only coding agent (claude_code stubbed, README pi link 404), the no-sandbox/no-merge/no-approval gap, and the placeholder test gates [big-pickle]
  • software-factory-feature-matrix: created as a single-column scaffold for sssf in the same run per the matrix rule, ten rows centered on the loop-owner, agent-boundary, and sandbox rows, the gaps named in prose; three references verified this run [big-pickle]
  • _index.md: Software factory category added to the research index and the matrix listed under Comparison matrices [big-pickle]
  • Verification: all section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets (both code-factories corpus anchors and the spec-driven/hybrid-execution/executions links confirmed on disk), every matrix column set matches its category membership (15 categories, 15 matrices), no em-dashes or banned terms, one sentence per line [big-pickle]

2026-08-29 (owner-prompted, software factory alternatives)
#

  • Owner follow-up: find other software-factory-style projects and assess credibility; a web search plus GitHub API scan surfaced and triaged the long tail, then two credible members cleared the bar and were written to full citation standards with all sources fetched or verified this run [big-pickle]
  • Scorecard of scanned candidates: fluent and har admitted; smartcomputer-ai/lightspeed (91 stars, deterministic Temporal harness, general orchestration rather than a factory) kept out with the reasons in the matrix long-tail; determinagent (10 stars, library), agentic-software-factory (1 star, inflated agent counts), ai-factory (4 stars, NOASSERTION license), and Takk8IS/software-factory (0 stars, NOASSERTION) did not clear the citation or credibility bar [big-pickle]
  • fluent: research note added as the category’s self-improving member; recorded 84 stars and 1 fork since creation 2026-07-10, 1,450 commits, v0.2.0 of 2026-08-15, Apache-2.0 Rust binary plus skill for Codex/Claude Code/Pi (macOS only), the Observations-to-Merge-Candidate loop with EARS Behavior Specs, isolated worktrees with AWS Fargate remote execution, the deterministic final Tester bound to the reviewed commit, and the Learner writing reusable Expertise as the self-improvement lead [big-pickle]
  • har: research note added as the harness-style member; recorded 82 stars and 10 forks since creation 2026-06-28, 351 commits, v1.0.0 of 2026-08-28 with daily releases, Apache-2.0 TypeScript CLI plus MCP server (npm @osfactory/har), the .har contract replacing scattered run knowledge, per-agent worktree/port/database slots, deterministic verify with a per-commit evidence trail, Mission Control dashboard, and Kerno sponsorship [big-pickle]
  • software-factory-feature-matrix: rebuilt from the single-column scaffold to three columns (SSSF, Fluent, HAR) in the same run per the matrix rule, with the who-owns-the-loop thesis, the isolation row separating factories from single-agent calls, Fluent’s self-improvement lead, SSSF’s pi-only dependency as the reason to look past headline ease, and the rejected long tail named in prose so no future run re-adds it silently [big-pickle]
  • super-simple-software-factory: Compared-to rewritten to name the new peers (Fluent’s ceremony, HAR’s fleet harness) alongside the cross-category comparisons [big-pickle]
  • _index.md: Fluent and HAR listed alphabetically in Software factory, matrix one-liner updated to the three-way who-owns-the-loop framing [big-pickle]
  • Verification: all section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets (all four code-factories corpus anchors and the intra-category links confirmed on disk), every matrix column set matches its category membership (15 categories, 15 matrices), no em-dashes or banned terms, one sentence per line [big-pickle]

2026-08-29 (owner-prompted, people and publications category)
#

  • Owner request in chat: introduce a category that collects the websites and people shaping the domain as it evolves; selected the people-plus-outputs scope and a full seed plus companion matrix confirmed via two clarifying questions [big-pickle]
  • Placement decision: a new People and publications category, because the five members are cross-cutting carriers of signal who do not fit any tool category; the category spans hands-on practice, industry synthesis, and conceptual vocabulary rather than a single tool type [big-pickle]
  • simon-willison: research note created, the daily hands-on chronicler, citing six sources verified this run (blog, TIL, LLM CLI, predictions post, agentic-engineering guide, and a critical press take on vibe coding) [big-pickle]
  • latent-space: research note created, the AI engineering newsletter-podcast-conference of record by swyx and Alessio, recorded 198,000+ subscribers and the multi-event AI Engineer series, sources verified this run [big-pickle]
  • the-pragmatic-engineer: research note created, Gergely Orosz’s org-and-data view, noting the paywalled core and the annual AI tooling surveys, sources verified this run [big-pickle]
  • steve-yegge: research note created, the operative who builds what he predicts (Gas Town, brute squad), linked to the existing gastown and beads notes, sources verified this run [big-pickle]
  • andrej-karpathy: research note created, the vocabulary-setter (vibe coding, Software 3.0, agentic engineering), now at Anthropic pre-training, sources verified this run [big-pickle]
  • people-and-publications-feature-matrix: created with five columns in the same run per the matrix rule, segmented on the focus axis (hands-on, industry, vocabulary) with the missing enterprise-hands-on cell named as the scaffold for a future member, references verified this run [big-pickle]
  • _index.md: People and publications category added to the research index (five notes listed alphabetically) and the matrix listed under Comparison matrices [big-pickle]
  • queue.md: new category appended to the standing seed-categories list and marked done with the date and matrix slug [big-pickle]
  • Verification: all section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets (intra-category, gastown, beads, landscape, and corpus anchors all confirmed on disk), the new matrix’s column set matches its category membership (16 categories, 16 matrices), no em-dashes or banned terms, one sentence per line [big-pickle]

2026-08-29 (daily refresh, deep re-verification + fx entrant)
#

  • Scan: all 96 article files as they stand after this run pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, every matrix column set matches its category membership, no em-dashes or banned terms, one sentence per line [glm-5.3-flash]
  • Deep refresh of the stalest cluster (updated 2026-08-23): crush, jetbrains, junie, opencode, vscode-copilot, windsurf, zed re-verified against live sources; every cited URL fetched 200 this run (9 HN threads via the Algolia items API, titles and points matching), plus GitHub API checks for the five repo-citing notes [glm-5.3-flash]
  • crush, jetbrains, junie, opencode, windsurf, zed: re-verified with no material change (stars within stated bounds, junie pricing re-confirmed at $8.33/$25 with 10/35 credits, zed pricing page re-confirms all tiers and the planned-not-shipped SSO/SAML/SCIM line); left untouched per the no-change precedent [glm-5.3-flash]
  • vscode-copilot: the April 22, 2026 self-serve pause is now documented as covering Copilot Business and Copilot Enterprise with sign-ups reopening soon for card and PayPal payers; contraction sentence rewritten, stars re-dated to about 190k as of 2026-08-29, reference dates bumped, updated field set, llm=glm-5.3-flash added [glm-5.3-flash]
  • New-entrant scan (12-day HN window plus GitHub checks) surfaced one credible harness: fx (vercel-labs/fx) [glm-5.3-flash]
  • fx: research note added to Harnesses, citing 9 sources all fetched this run (repo plus GitHub API, the site, four docs pages, the 317-point launch thread via Algolia, vercel.com/oss); recorded the embed-first thesis (about 6.19 MiB Zig binary, 10 microsecond cold start, Wasm build), the three-credential model (Vercel AI Gateway, Codex OAuth, Grok OAuth) with locally stored tokens, native AGENTS.md, Claude-family skill directories, persistent-child subagents, the fx acp server, and the every-request-routes-through-AI-Gateway caution; 2,589 stars, v0.0.7 of 2026-08-29 [glm-5.3-flash]
  • harness-feature-matrix: extended from 13 to 14 columns in the same run per the matrix rule, all 11 fx cells traced to the docs pages fetched this run; reading section gains the embed-first line, choosing gains the embedded-agent bullet, references extended by 3, verified-date line names the fx column’s own date [glm-5.3-flash]
  • agentic-coding-tools-landscape: harness tail gains the fx line, verified-dates line extended, updated bumped to 2026-08-29 [glm-5.3-flash]
  • Attribution note: the vscode-copilot, landscape, harness-matrix, and _index.md changes above were left uncommitted in the working tree while this run proceeded, and the concurrent big-pickle software-factory run committed them in ba32d985 alongside its own work; this entry records the authorship [glm-5.3-flash]
  • The standing ? for tom: about untracked the-agentic-development-environment-landscape/ remains unanswered, so it stays unlinked [glm-5.3-flash]
  • No self-directed essay this run: the deep re-verification plus the fx note with its matrix column filled the day at quality [glm-5.3-flash]

2026-08-29 (correction to the daily refresh entry above)
#

  • The attribution note above was wrong: only the _index.md changes went in with ba32d985; the vscode-copilot, agentic-coding-tools-landscape, and harness-feature-matrix changes were still uncommitted and landed now in their own commit [glm-5.3-flash]

2026-08-29 (people and publications category expansion)
#

  • Owner request in chat and the standing category directive both call for building out the People and publications category; the five-member seed left the enterprise-hands-on cell and the model, evaluation, systems, and education bands open, and these six members fill them [big-pickle]
  • hamel-husain: research note added as the evaluation-and-verification band, the data-driven method for deciding whether an AI product works, citing the Field Guide, the evals post, the Maven course (4,500+ students from 500+ companies), and a critical Arize take on why evals fail, sources verified this run [big-pickle]
  • lilian-weng: research note added as the durable research-reference band, whose LLM Powered Autonomous Agents post is the canonical survey of planning, memory, and tool use, sources verified this run [big-pickle]
  • chip-huyen: research note added as the systems-survey band, the most-read O’Reilly AI book of 2025, sources verified this run [big-pickle]
  • nathan-lambert: research note added as the model-and-post-training band, the open-ecosystem voice from inside the labs, recording the Get Good at Agents and Claude Code pieces and the RLHF Book, sources verified this run [big-pickle]
  • deeplearning-ai-andrew-ng: research note added as the education-and-on-ramp band, the voice that popularized the four agentic design patterns and published the AI Engineering Skills Map, sources verified this run [big-pickle]
  • ai-jason: research note added as the video-practitioner band, the channel that shows agent workflows and context engineering working end to end, recording 230K subscribers and 99 videos, sources verified this run [big-pickle]
  • people-and-publications-feature-matrix: extended from five to eleven columns in the same run per the matrix rule, with the thesis re-segmented onto five focus bands (hands-on, evaluation, model-and-research, industry, systems-and-education), the been-empty enterprise-hands-on cell re-named as the remaining scaffold, and all eleven reader-slot rows filled from the member notes [big-pickle]
  • _index.md: the six new notes listed alphabetically in People and publications, and the matrix one-liner bumped from five to eleven voices [big-pickle]
  • queue.md: left untouched; the People and publications category item is already marked done, and expanding a done category is self-directed growth, so no new queue line applies (matching how the fluent and har additions did not touch the queue) [big-pickle]
  • Verification: all section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets (all intra-category links and corpus anchors confirmed on disk), the matrix column set matches its category membership (11 members, 11 columns), no em-dashes or banned terms, one sentence per line [big-pickle]

2026-08-29 (caleb writes code added to people and publications)
#

  • Owner request in chat names the channel directly (youtube.com/@calebwritescode); added as the twelfth member of People and publications, filling the same-week video release-explainer slot next to AI Jason’s workflow-builder one [big-pickle]
  • caleb-writes-code: research note added for the YouTube explainer channel of Caleb Eom; identity, description (“part editorial, part informational on AI”), join date 2025-03-16, and LinkedIn/X/Patreon links fetched from the channel about page, scale fetched from a live counter (105,664 subscribers, about 6.3M views, 114 videos, as of 2026-08-29), cadence and topics fetched from the channel RSS (uploads about twice a week, latest 2026-08-27) [big-pickle]
  • caleb-writes-code critical grounding: 11 of the last 15 upload descriptions carry a sponsor or affiliate read (computed from the RSS descriptions this run), no Hacker News threads and no substantive Reddit discussions were found, and one third-party analytics site’s stale inactivity prose contradicts its own FAQ and the live feed, all recorded as cautions [big-pickle]
  • people-and-publications-feature-matrix: extended from eleven to twelve columns in the same run per the matrix rule, Caleb Writes Code column placed after AI Jason in the hands-on cluster, hands-on reading line rewritten to name the two distinct video modes, choosing list gains the release-explainer bullet [big-pickle]
  • _index.md: Caleb Writes Code listed alphabetically between Andrew Ng and Chip Huyen, matrix one-liner bumped from eleven to twelve voices [big-pickle]
  • Verification: front matter parses with no type field and mandatory tags, all internal links resolve on disk, matrix columns match category membership (12 members, 12 columns), no em-dashes or banned terms, one sentence per line [big-pickle]

2026-08-30 (jcode added to harnesses)
#

  • Owner request in chat: cover jcode (github.com/1jehuang/jcode); placed in Harnesses, verified this run via repo page, GitHub API, jcode.sh (home, docs, about, pricing), the grigio.org independent comparison, and two HN threads via the Algolia items API [glm-5.3-flash]
  • jcode: research note created; recorded the resource-efficiency thesis (27.8 MB PSS with embedding off, ~10 MB per added session, 14 ms to first frame, all self-published), the embedding memory graph, native swarm, self-dev mode, confidence stepping and auto-poke, 671-token system prompt, session portability, 30-plus providers with subscription OAuth and multi-account switching, stdio-only MCP, $10/month hosted inference, YC S26 solo founder, 18,796 stars and 2,144 forks and v0.81.1 of 2026-08-26 as of verification, and the thin independent footprint (two HN threads at 3 and 5 points, zero comments) as a stated signal [glm-5.3-flash]
  • harness-feature-matrix: extended from fourteen to fifteen columns in the same run per the matrix rule, jcode placed between goose and Junie, all eleven cells traced to the note’s sources; verified-date line names the jcode column’s own date; open-client and local-model group lists updated; choosing list gains the parallel-agents bullet; references extended by 2 [glm-5.3-flash]
  • agentic-coding-tools-landscape: harness tail gains the jcode line, verified-dates sentence extended, updated bumped to 2026-08-30 [glm-5.3-flash]
  • _index.md: jcode listed alphabetically in Harnesses, matrix one-liner bumped to fifteen harnesses verified 2026-08-30 [glm-5.3-flash]
  • Verification: all 112 section files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets in the touched files, matrix columns match category membership (15 members, 15 columns, 15 cells per row), no em-dashes or banned terms, one sentence per line [glm-5.3-flash]

2026-08-30 (company row added to harness matrix)
#

  • Owner request in chat: add an associated-company row to the harness feature matrix; added as the first row under the header since it is maker identity, not a feature axis, with a linking cell per column [glm-5.3-flash]
  • harness-feature-matrix: Company row names each maker (aider none/solo, Amp Frontier, Anthropic, Cline Bot, OpenAI, Charm, Vercel, Google, AAIF ex-Block for goose, Solo Systems, JetBrains, Anaconda post-acquisition for Kilo Code, Anomaly, All Hands AI, Alibaba) with every attribution traced to the column’s own note; reading-the-matrix prose gains one line framing the row as context, not a feature axis [glm-5.3-flash]
  • External link verification: all 14 new company URLs fetched this run (openai.com returns 403 to curl as bot protection but is the canonical site), Anomaly’s site resolved to anoma.ly via the anomalyco GitHub org profile after anomaly.co/anomaly.com proved to be unrelated namesakes [glm-5.3-flash]
  • Verification: Company row has 16 cells matching the 16-column header, no em-dashes or banned terms, one sentence per line, updated already at 2026-08-30 [glm-5.3-flash]

2026-08-30 (owner-directed star sweep: twelve new notes, six matrices extended, one category seeded)
#

  • Owner request in chat: analyze all GitHub stars and write the top missing projects now; twelve stars were selected as high-priority gaps after diffing the star list against section membership [glm-5.3-flash]
  • deepseek-harness: created harnesses note (DeepSeek’s MIT everything-is-a-plugin harness, 203,755 stars, alpha, five references fetched) [glm-5.3-flash]
  • pi: created harnesses note (Earendil’s minimal extensible harness, 99,163 stars, no-MCP policy, oh-my-pi fork recorded, six references fetched) [glm-5.3-flash]
  • graphify: created context-engines note (local AST knowledge graph as skill and MCP server, 112,374 stars, self-benchmarked caveat recorded, six references fetched) [glm-5.3-flash]
  • claude-mem: created memory note (session capture, compression, reinjection, 92,602 stars, token-cost and privacy cautions, seven references fetched) [glm-5.3-flash]
  • rtk: created context-engines note (CLI output proxy, 77,844 stars, bytes-versus-bill caveat from its own docs, seven references fetched) [glm-5.3-flash]
  • qmd: created context-engines note (Tobias Lütke’s local hybrid search, 29,359 stars, security-fix history recorded, five references fetched) [glm-5.3-flash]
  • skillopt: created skills note (Microsoft Research skill optimizer, 16,477 stars, validation gate, unreplicated-results caution, seven references fetched) [glm-5.3-flash]
  • paseo: created orchestration note (cross-device daemon orchestration, 15,479 stars, solo-maintainer risk, six references fetched) [glm-5.3-flash]
  • superset: created orchestration note (source-available agentic IDE, 13,498 stars, ELv2, macOS-first caution, six references fetched) [glm-5.3-flash]
  • worktrunk: created orchestration note (worktree CLI, 6,752 stars, single-maintainer risk, six references fetched) [glm-5.3-flash]
  • ouroboros: created software-factory note (hidden-grading verification loop, 5,729 stars, thin independent validation recorded, six references fetched) [glm-5.3-flash]
  • agentsview: created session-analytics note (local session indexer, 5,589 stars, seeded the new Session analytics category) [glm-5.3-flash]
  • session-analytics-feature-matrix: created as the new category’s companion matrix, single-column scaffold with the live-observation gap named in prose [glm-5.3-flash]
  • harness-feature-matrix: DeepSeek Harness and Pi columns added, matrix at seventeen columns, prose and references extended [glm-5.3-flash]
  • context-engines-feature-matrix: Graphify, qmd, and rtk columns added, matrix at seven columns, kind-row prose and choosing list extended [glm-5.3-flash]
  • memory-feature-matrix: claude-mem column added, matrix at six columns, architecture-row prose updated to four vendors [glm-5.3-flash]
  • skills-feature-matrix: SkillOpt column added, matrix at five columns, stack prose names the new quality layer [glm-5.3-flash]
  • orchestration-feature-matrix: Paseo, Superset, and Worktrunk columns added, matrix at eleven columns, funding-spectrum sentence and choosing list added [glm-5.3-flash]
  • software-factory-feature-matrix: Ouroboros column added, matrix at four members, self-improvement and hidden-grading prose updated [glm-5.3-flash]
  • _index.md: twelve notes listed alphabetically in their categories, Session analytics section added, seven matrix one-liners updated [glm-5.3-flash]

2026-08-30 (owner-directed sweep two: two categories seeded, assistant runtimes expanded to nine)
#

  • Owner request in chat: cover the zero-note evaluation/review and sandboxing projects and expand assistant runtimes; fourteen notes written from research fetched this run [glm-5.3-flash]
  • workshop: created evaluation note (Raindrop’s local agent debugger, 1,064 stars, CI-disconnect criticism recorded, five references fetched) [glm-5.3-flash]
  • deepeval: created evaluation note (pytest-style eval framework, 17,960 stars, open-core split caution, six references fetched) [glm-5.3-flash]
  • phoenix: created evaluation note (Arize’s OTel observability platform, 11,245 stars, ELv2 license caution, six references fetched) [glm-5.3-flash]
  • open-code-review: created evaluation note (Alibaba’s hybrid reviewer, 21,636 stars, independent 12-percent-precision run recorded, six references fetched) [glm-5.3-flash]
  • plannotator: created evaluation note (nine-harness review surface, 8,235 stars, unencrypted-small-share discrepancy recorded, six references fetched) [glm-5.3-flash]
  • openshell: created sandboxing note (NVIDIA’s agent runtime, 8,420 stars, alpha and default-telemetry cautions, six references fetched) [glm-5.3-flash]
  • agent-sandbox: created sandboxing note (Kubernetes SIG Apps CRD, 3,680 stars, isolates-nothing-by-itself threat model recorded, six references fetched) [glm-5.3-flash]
  • aigate: created sandboxing note (14 stars, kept per the missing-footprint-is-a-signal rule, no-audit caution primary, six references fetched) [glm-5.3-flash]
  • flue: created sandboxing note (Astro team’s framework, 8,083 stars, three-tier sandbox taxonomy, single-author risk, six references fetched) [glm-5.3-flash]
  • artifact-fs: created sandboxing note (Cloudflare’s FUSE provisioning driver, 1,127 stars, vendor-benchmark caveat, six references fetched) [glm-5.3-flash]
  • nanobot: created assistant-runtimes note (HKUDS’ Python runtime, 47,529 stars, alpha and bus-factor cautions, six references fetched) [glm-5.3-flash]
  • qwenpaw: created assistant-runtimes note (AgentScope team assistant, 34,673 stars, telemetry-auto-accept caution, six references fetched) [glm-5.3-flash]
  • openwork: created assistant-runtimes note (OpenCode-based Cowork alternative, 23,202 stars, license-split and security-boundary cautions, six references fetched) [glm-5.3-flash]
  • eigent: created assistant-runtimes note (CAMEL workforce desktop, 15,160 stars, corrected-GAIA-claim history recorded, six references fetched) [glm-5.3-flash]
  • evaluation-review-feature-matrix: created, five columns on the who-judges axis [glm-5.3-flash]
  • sandboxing-feature-matrix: created, five columns with the kind row separating boundaries from plumbing [glm-5.3-flash]
  • assistant-runtimes-feature-matrix: Nanobot, OpenWork, Eigent, and QwenPaw columns added, matrix at nine columns, desktop-split prose added, one write failure mid-edit repaired [glm-5.3-flash]
  • _index.md: two category sections added, fourteen notes listed, three matrix one-liners updated [glm-5.3-flash]

2026-08-30 (owner edit: matrices must sort entries alphabetically)
#

  • Owner request in chat: added to the Comparison matrices rules in AGENTS.md that column entries are sorted alphabetically by member title, case-insensitive, with no privileged position for founders, roots, or baseline conventions, re-sorted in the same run as any membership change, and that an unsorted matrix is a defect fixed in the commit that surfaced it [owner, glm-5.3-flash]
  • Compliance scan of the sixteen matrices: eight currently deviate from the new rule and will be re-sorted by the next daily run: assistant-runtimes (OpenClaw root first), evaluation-review, memory (Files baseline first), people-and-publications, retrieval, sandboxing (kind-ordered), software-factory (SSSF founding member first), spec-driven-development [glm-5.3-flash]

2026-08-30 (matrices re-sorted per the new alphabetical rule)
#

  • Applied the same-day owner rule to the eight non-compliant matrices, reordering every row’s cells to match the sorted headers, ordering only, no cell contents or verification dates changed [glm-5.3-flash]
  • Re-sorted: assistant-runtimes (Eigent through ZeroClaw, dropping the root-first convention), evaluation-review (deepeval, OpenCodeReview, Phoenix, Plannotator, Workshop), memory (claude-mem, Cognee, Files, Letta, Mem0, Zep, dropping the Files-first baseline), people-and-publications (AI Jason through The Pragmatic Engineer, matching the index order), retrieval (LangChain first), sandboxing (Agent Sandbox, aigate, ArtifactFS, Flue, OpenShell, dropping kind-order), software-factory (Fluent first, dropping the founding-member-first convention), spec-driven-development (GitHub Spec Kit after BMad) [glm-5.3-flash]
  • retrieval, spec-driven-development, and people-and-publications gained the llm=glm-5.3-flash tag for this edit; updated bumped to 2026-08-30 on all eight [glm-5.3-flash]
  • harness-feature-matrix: corrected two column-ordinal references in prose (DeepSeek Harness is column seven, Pi column sixteen, the sentences said sixteenth and seventeenth) [glm-5.3-flash]
  • Verification: all sixteen matrices now sort case-insensitively by header title with no ragged rows [glm-5.3-flash]

2026-08-30 (repair: invalid YAML from the sort commit broke the site build)
#

  • Owner report: the deployed people-and-publications matrix still showed the pre-sort order [glm-5.3-flash]
  • Root cause found in this log’s own previous run: the tag-insertion step in the re-sort run appended , llm=glm-5.3-flash] before the original closing bracket, leaving ]] on the tags line of people-and-publications, retrieval, and spec-driven-development; the invalid front matter failed the site build, so the deployment kept serving the last good (pre-sort) build [glm-5.3-flash]
  • Repaired the three tags lines to single closing brackets, normalized the double-space table padding the sort script introduced across the eight re-sorted matrices, and re-verified: all section front matter parses as YAML, all sixteen matrices still sort case-insensitively with no ragged rows [glm-5.3-flash]
  • Process note: the pre-commit verification that run checked sort order and links but skipped YAML parsing, which is how the defect shipped; front matter parsing is part of the standing verification list and stays there [glm-5.3-flash]

2026-08-30 (daily refresh, context engines re-verification + semble entrant)
#

  • Section scan: all 142 article files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, every matrix column set matches its category membership (19 categories, 19 matrices), columns case-insensitively sorted, no ragged rows [glm-5.3-flash]
  • Banned-term sweep caught five “shape” slips the earlier owner-directed sweeps missed: “multi-layer shaping” in fluent, “shaping OSS agent expectations” in letta, “install shapes” in ouroboros, “board-shaped” twice in paseo, and “production-shaped” in phoenix, all reworded, meaning unchanged; fluent and letta gained updated 2026-08-30 and llm=glm-5.3-flash (fluent keeps llm=big-pickle) [glm-5.3-flash]
  • Stalest cluster computed as the 2026-08-24 notes never touched by a refresh run; Context engines carried the most members, so it took this run’s deep refresh: augment-code, greptile, repomix, sourcegraph-code-context re-verified against live sources [glm-5.3-flash]
  • Link check: all 27 reference URLs of the four notes fetched 200 this run (5 HN threads via the Algolia items API, titles and points matching), greptile and sourcegraph pricing pages re-confirmed byte-for-byte on their claims ($30/seat plus credits; $16K entry plus AI credits), both left untouched per the no-change precedent [glm-5.3-flash]
  • repomix: volatile numbers refreshed to 28,122 stars, 1,508 forks, 4,476 commits, 373,996 npm downloads for 2026-07-31 to 2026-08-29, v1.18.0 of 2026-08-08 (GitHub and npm APIs); verified-as-of and updated bumped, llm=glm-5.3-flash added [glm-5.3-flash]
  • augment-code: blog cadence extended from August 14 to August 28, 2026 on the fetched blog index; no-free-tier claim re-dated to the 2026-08-30 pricing-page fetch; verified-as-of and updated bumped, llm=glm-5.3-flash added [glm-5.3-flash]
  • New-entrant scan surfaced Semble (MinishLab/semble) as the credible context-engine addition; Mantic.sh scanned and kept out (554 stars, stalled since July 13); OneCLI and Bullet noted as harness candidates for a future run, not written this one [glm-5.3-flash]
  • semble: research note added to Context engines, citing 6 sources all fetched or Algolia-verified this run (repo plus GitHub API, the 445-point launch thread with the founders’ no-agent-level-claim admission and the grep-baseline criticism, PyPI plus PyPIstats 119,273 downloads for the month ending 2026-08-29, the potion-code-16M-v2 model page, the benchmarks README); recorded the sub-second CPU indexing thesis, the installer’s three delivery surfaces (MCP, AGENTS.md instructions, sub-agent), the RRF-plus-rerank stack, the self-measured semble savings counter, and the two-maintainer pre-1.0 risks [glm-5.3-flash]
  • context-engines-feature-matrix: extended from 7 to 8 columns (Semble inserted between rtk and Sourcegraph, alphabetically) in the same run per the matrix rule, all 11 cells traced to the note’s sources; reading, choosing, and references sections extended [glm-5.3-flash]
  • _index.md: Semble listed alphabetically in Context engines, matrix one-liner bumped from seven to eight tools [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the ? for tom: about untracked the-agentic-development-environment-landscape/ remains unanswered, so it stays unlinked [glm-5.3-flash]
  • No self-directed essay this run: the category re-verification plus the Semble note with its matrix column filled the day at quality, and five earlier runs already published today [glm-5.3-flash]
  • Verification: all 143 article files pass front matter parsing, 0 broken internal link targets, all 19 matrices match their categories, no em-dashes or banned terms, one sentence per line, and every URL cited by new or edited content fetched 200 or was Algolia-verified this run [glm-5.3-flash]

2026-08-30 (owner-directed: parallel refresh model + first full parallel refresh)
#

  • Owner directive in chat: the research index line “refreshed one category at a time, stalest first” becomes refresh-all-in-parallel; the rotation is retired [owner, glm-5.3-flash]
  • AGENTS.md daily refresh step 3 rewritten (owner-directed): every category refreshed in parallel each run, no stalest ranking, matrices re-verified in the same run; _index.md research-index line and methodology.md Refresh cadence section updated to match [owner, glm-5.3-flash]
  • Verification policy under the new model: every re-verified note gets its “Facts below verified as of” line bumped even with no material change (that line documents when facts were last checked); the updated field still moves only on material revision; recorded here so the owner can veto the bump policy [glm-5.3-flash]
  • First full parallel refresh executed this run: 15 sub-runs across 3 waves, one to four categories each, every run confined to its own note and matrix files with no git commands; Context engines sat out (fully re-verified earlier today, Semble note and matrix column included); 128 files changed [glm-5.3-flash]
  • Coverage: protocols, retrieval, memory, executions, hybrid execution, skills, surfaces, orchestration, harnesses (split A and B), task management, control planes, session analytics, spec-driven development, evaluation and review, sandboxing, software factory, assistant runtimes, people and publications; all 19 matrices re-dated to 2026-08-30 with cells updated only where facts moved [glm-5.3-flash]
  • Material corrections this run: claude-code-hooks “5k+ open issues” was materially wrong, now about 15k (search API); plannotator “about 30 contributors” rewritten to the verified 135-commit-contributor figure with 859 commits from the founder; eigent’s “200-plus built-in MCP tools” dropped, the cited README no longer states any count; task-master gained the missed Hamster Enterprise tier ($200/creator/month); picoclaw release line corrected to v0.3.1 of 2026-07-03; hamel-husain course claim aligned to the current page wording; deeplearning-ai lost an unverifiable jobpocalypse quote and a dead aiagentrank reference (replaced with the verified Coursera specialization page) [glm-5.3-flash]
  • URL fixes: openchamber release reference bumped v1.20.0 to v1.22.0, steve-yegge gastown repo URL canonicalized to gastownhall/gastown (org move, old path redirects), karpathy Medium and O’Reilly book pages confirmed bot-blocked and kept with their dates, skillregistry.io download count left client-rendered-unverifiable with its 2026-08-24 date [glm-5.3-flash]
  • Volatile refresh highlights: claude-code 143.5k stars and 15.5k open issues+PRs, codex 120.1k, cline 5,154,938 installs, openclaw 388,084, hermes 238,457, spec-kit 132,331, deepseek-harness dsh-v0.1.2-alpha.2 published today, HAR v1.4.0 released today, openchamber v1.22.0 released today, instructor 1.16.0, openhands v1.16.0, conductor 0.83.0, gh-aw v0.87.9 with the retired-release notice gone from the README, kiro 2.20.1 [glm-5.3-flash]
  • Entrant candidates recorded for future runs, no notes written: Warp Agent CLI (111 pts) and Zerostack (575 pts) as harness prospects, OneCLI (88 pts, also named in nanoclaw) and Tilde.run (205 pts) as sandboxing prospects, Tuneloop as the session-analytics gap column, microsoft/agent-host-protocol (286 stars) as a protocols watch; Bullet (121 pts) carried over as a known harness candidate [glm-5.3-flash]
  • _index.md one-liners refreshed where they lagged: surface and people matrix dates to 2026-08-30, Cline 5.1 to 5.2 million installs, Spec Kit 131k to 132k, Hermes 237k to 238k [glm-5.3-flash]
  • Verification: all 143 article files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 19 matrices match their category membership with columns sorted and no ragged rows, no em-dashes or banned terms, and every URL verified this run was fetched 200, Algolia-confirmed, archive-API-confirmed, or recorded as bot-blocked with its date kept [glm-5.3-flash]

2026-08-30 (owner-requested: Code review category)
#

  • Owner request in chat: do we cover AI-assisted code review tools; answer was only partially (Greptile sat in Context engines, OpenCodeReview in Evaluation and review, and no note for CodeRabbit, Qodo, Ellipsis, or the space itself), so the category was introduced [glm-5.3-flash]
  • Research index: seeded the Code review category with five members, three written fresh this run to full citation standards by parallel sub-runs (all sources fetched 200 or Algolia-verified): coderabbit, qodo, ellipsis, plus greptile and open-code-review moved in on category-fit grounds (both are review-first products; the context-engines matrix had already labeled Greptile “review service, API”) [glm-5.3-flash]
  • coderabbit: recorded the $16M/$60M/$143M funding ladder (August 2026 Series C at a $1.5B valuation), 17k customers and 2M+ weekly reviews as of 2026-08-30, the Kudelski RCE-and-write-access exploit chain on 1M repos with the vendor’s remediation timeline, the Pullflow 40.3M-PR analysis showing Copilot overtaking it on monthly volume while it keeps the cumulative-2025 lead, and the free-forever public-repo tier [glm-5.3-flash]
  • qodo: recorded the open-core split (donated, community-maintained MIT PR-Agent at 12,769 stars, versus paid Qodo Merge), the $11M plus $40M funding, the CodiumAI rename grounded via TechCrunch after the rebrand post itself proved unlocatable, the $0.012 credit pricing, and the second Kudelski chain (PR comment to AWS admin key, fixed October 2025) [glm-5.3-flash]
  • ellipsis: status set to pivoted, the July 2026 Agent Cloud launch (managed infrastructure for running nine coding agents, agents-as-code YAML, budget caps, BYOC) demotes review to a use case; recorded the $2M seed, the 121-point 2024 Show HN with split sentiment, and the usage-plus-10%-fee pricing [glm-5.3-flash]
  • code-review-feature-matrix: created with five columns (CodeRabbit, Ellipsis, Greptile, OpenCodeReview, Qodo, alphabetical) and ten rows; the thesis is that where your code runs decides the purchase, three vendor clouds versus VPC versus fully self-hosted, with both Kudelski exploit records and the Copilot-overtake market number in the reading section [glm-5.3-flash]
  • context-engines-feature-matrix: Greptile column removed (8 to 7 columns), all rows and prose updated, a line added explaining the move; evaluation-review-feature-matrix: OpenCodeReview column removed (5 to 4), who-judges thesis rewritten to three judges [glm-5.3-flash]
  • Note cross-links updated: greptile and open-code-review now link CodeRabbit (previously unlinked name-drops), qodo and ellipsis gained their sibling links, open-code-review’s see-also now joins the Code review matrix, greptile gained the flash tag and updated 2026-08-30 [glm-5.3-flash]
  • _index.md: Code review category added with five notes and the new matrix one-liner; Greptile and OpenCodeReview removed from their old sections; context-engines one-liner now seven tools, evaluation one-liner now four columns [glm-5.3-flash]
  • Scan candidates recorded for the category’s long tail, none note-worthy yet: Kodus (1,335 stars, open-source CodeRabbit alternative), Sourcery (1,857 stars), Cubic (thin HN), Graphite Diamond ($52M vendor, JS-only presence), plus the Vigilant-PR/Codra/hive-review tail; Copilot’s platform-native reviewer stays covered inside vscode-copilot per the grouping precedent [glm-5.3-flash]
  • Verification: all 147 article files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 matrices match their category membership with columns sorted and no ragged rows, no em-dashes or banned terms, one sentence per line, and every URL cited by new content was fetched 200 or Algolia-verified during this run [glm-5.3-flash]

2026-08-30 (owner-directed: same-run candidate resolution, full pile processed)
#

  • Owner directive in chat: collecting candidates without processing them in the same refresh has no justification; codified in AGENTS.md step 3 and methodology.md, every entrant candidate now resolves in the run that surfaces it, a note or an explicit logged rejection, and a candidate carried as “pending” is a defect [owner, glm-5.3-flash]
  • The existing candidate pile was processed to zero this run: 37 candidates resolved, 11 accepted as notes (written by parallel sub-runs to full citation standards), 26 explicitly rejected with fresh evidence fetched today [glm-5.3-flash]
  • Accepted into Harnesses (6): warp-agent-cli (Warp’s standalone agent binary, 64,668 stars, AGPL-3.0, credit-meter pricing), zerostack (solo GPL-3.0 Rust agent, 575-point launch, 1,609 stars, compile-flag features), onecli (YC S26 Apache-2.0 team harness, per-request credential injection, 3,430 stars), bullet (YC S26 closed-source latency bet, 121-point launch, self-reported 479/500 SWE-bench, npm early), ante (Antigma Labs’ offline-first Rust harness with embedded llama.cpp, 169-point launch, 1,919 stars), juggler (JUCE creator’s AGPL GUI agent with inspectable conversation trees, 280-point launch, 576 stars) [glm-5.3-flash]
  • Accepted into Orchestration (1): omnara (YC S25 Apache-2.0 Go control plane, 310-point 2025 Show HN plus 147-point 2026 launch, 2,778 stars, placed here on product reality, a supervised-agent control plane, not a harness) [glm-5.3-flash]
  • Accepted into Protocols (1): agent-host-protocol (Microsoft’s sessions-server spec, v0.9.0, six-language SDKs, 288 stars, 83k npm downloads, VS Code reference host in microsoft/vscode) [glm-5.3-flash]
  • Accepted into Code review (3): kodus (AGPL-3.0 dual-license BYOK reviewer, 1,337 stars, funding stated unverifiable), sourcery (MIT refactoring lineage versus proprietary reviewer, Pro $12/Team $24, funding unverifiable), graphite-diamond (recorded with the Diamond brand deprecated into Graphite Agent and the December 2025 Cursor acquisition) [glm-5.3-flash]
  • Rejected with current evidence (8 named candidates): hax (121 pts, 681 stars, niche already covered by stronger members), zot (107 pts, 330 stars, serial self-promotion pattern), hoplite (81 pts, no public repo, closed product), cubic (1-pt and 4-pt HN, client-rendered vendor pages cannot carry the citation bar), tilde.run (shut down, its domain now redirects to the lakefs shutdown announcement, 205-point traction now historical), tuneloop (5 pts, 47 stars, retrospective analytics duplicating agentsview at lower traction), llm-eval-kit (repo deleted), machine0 (unchanged since the prior below-bar assessment) [glm-5.3-flash]
  • Rejected with current evidence (20 weak-tail candidates re-verified today, none escalated): polign (3 stars), mnemosyne (renamed gomaa, 30 stars), knowl (27 stars), zuse (51), rove (114, no product site), operator (201, quiet, no site), lanegate (2 stars, two days old), singular (37, stale), concord (closest call: 294 stars and growing but 9-point HN and only 4 verifiable sources), proliferate (446 stars, single-source traction), huzzah (384-point thread but a self-declared solo experiment, 188 stars), biscuit (5-year-old repo, 2 pts), shikigami (7 pts, no locatable repo), context.dev (119 pts but an extraction layer, category mismatch), parsewise (56 pts, category mismatch, no movement), spec27 (13 pts, four months stale), nexa-gauge (40 stars, three self-posts), dokimos (56 stars, eight months, no pickup), plus near-misses flagged for re-review only on new independent coverage: concord and huzzah [glm-5.3-flash]
  • Matrices extended in the same run: harness 17 to 23 columns (Ante, Bullet, Juggler, OneCLI, Warp Agent CLI, Zerostack, alphabetical, 14 rows by 24 cells verified), orchestration 11 to 12 (Omnara), protocols 4 to 5 (Agent Host Protocol), code review 5 to 8 (Graphite Diamond, Kodus, Sourcery); untraceable cells marked ? not guessed [glm-5.3-flash]
  • _index.md and agentic-coding-tools-landscape updated for all eleven new notes, matrix one-liners re-dated; the AHP note’s stale see-also line fixed [glm-5.3-flash]
  • Style exceptions this run: “rate-shaping” in onecli reworded to “rate limits”; the remaining scan hit, “protocol-shape” in ante, is the proper name of a crate in Antigma’s repository and stays as fact, not prose word choice [glm-5.3-flash]
  • Verification: all 158 article files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 matrices match their category membership with columns sorted and no ragged rows, no em-dashes, and every URL cited by new content was fetched 200 or Algolia-verified during this run [glm-5.3-flash]

2026-08-30 (correction to the entry above)
#

  • The entry says 158 article files; the correct count is 157 (146 before this run plus 11 new notes) [glm-5.3-flash]

2026-09-02 (daily refresh, full parallel re-verification + addy-osmani entrant)
#

  • Full parallel refresh under the no-stalest-ranking rule: 12 note sub-runs re-verified all 134 research notes across the 20 categories against live sources; every “Facts below verified as of” line bumped to 2026-09-02 per the standing bump policy, notes with material revisions got updated plus the flash tag, bump-only notes got neither [glm-5.3-flash]
  • Link check: about 450 cited URLs fetched 200 or Algolia-verified across the sub-runs; blocked pages kept their dated claims (windsurf.com 429, karpathy Medium and O’Reilly and openai.com 403, pypistats 429, some HN items 429 to direct fetch, all verified via the Algolia API) [glm-5.3-flash]
  • Surfaces: cursor recorded the OpenAI wind-down (notice 2026-08-28, model access shutoff November 12, per devops.com) and windsurf the about-$1B raise at a about-$47B valuation (The Edge, 2026-09-02); kiro gained 2.20.2, Kiro Web GA, and the ISO 27001 certification; continue installs 4,048,536; roo-code repo confirmed archived-flagged [glm-5.3-flash]
  • Harnesses: deepseek-harness dsh-v0.1.2-alpha.4 and 208,672 stars, pi crossed 100k, opencode about 203k, amp’s pricing page dropped SpaceX AI from the linked list, ante v0.preview.92, bullet v1.4.16, jcode’s upstream benchmark re-framed (about 10 MB per added session versus Claude Code 212.7 MB), junie gained the Junie Local on-device Mac launch, kilo-code’s README dropped its popularity claim, onecli v2.4.0 with restructured hosted tiers [glm-5.3-flash]
  • Orchestration: cmux pricing packaging changed (CodeRouter gone, unlimited active Cloud VMs), paseo v0.7 stable, worktrunk v0.76.0 with a breaking -x change, conductor 0.83.2, superset v1.25.1, vibe-kanban orphan state re-confirmed; Assistant runtimes: openclaw 2.0 shipped 2026-08-30 (v2026.8.1, 933 contributors), hermes about 240k stars, openwork pricing overhauled (Free to 5 users, Team $20/seat, Enterprise $50) [glm-5.3-flash]
  • Code review: coderabbit pricing restructured (Essentials/Team renames, new $72/$90 Advanced tier replacing the $40 Security plan) and kodus pricing restructured (Community free, Teams BYOK $10/dev, Enterprise), qodo v0.44.0, open-code-review v1.11.2, ellipsis gained public SDK mirrors [glm-5.3-flash]
  • Elsewhere material: spec-kit v1.0.2 and v1.0.3, paperclip v2026.831.1, agentsview v0.42.0, claude-mem v13.23.1, rtk v0.47.0, graphify v0.9.53, phoenix 20.5.0, plannotator v0.27.11, fluent v0.3.0 (Slack planning and recovery), har v1.7.0, ouroboros v0.53.0, github-agentic-workflows v0.87.10, skills-sh leaderboard numbers moved, andrej-karpathy’s verifiability essay corrected to the 2025-11-17 “Verifiability” essay, nathan-lambert’s Interconnects now 80,000-plus subscribers [glm-5.3-flash]
  • Entrant resolution in-run: addy-osmani added to People and publications, the enterprise-hands-on gap the matrix named, 9 sources fetched 200 this run including the critical HN discussion; rejected with evidence: 49 IDE, Murmull, Supafork, Almanac YC S26 (out of scope), Openheim, Superagent, OpenCrew, Shed, Agentdock, KanVibe, Wasmer local sandboxes, lemmalog, ownmem, memcode, Memcode CR, Armin Ronacher (traction), kli, Fountain, VT Code, Wolfpack, Intent, Decispher, Mantic.sh, PearAI and Qoder (no new coverage) [glm-5.3-flash]
  • Matrices: all 20 re-verified in the same run, membership matched, columns sorted, no ragged rows; people matrix 12 to 13 columns (Addy Osmani, enterprise-hands-on scaffold now filled), harness matrix Junie local-models cell updated for Junie Local, orchestration cells for cmux/paseo/vibe-kanban/worktrunk, assistant-runtimes cells for hermes/openclaw, code-review pricing cells for coderabbit/kodus, spec-driven star cells, control-planes status cells [glm-5.3-flash]
  • Essays: landscape as-of bumped to 2026-09-02 with the cursor wind-down and windsurf funding folded in; model-selection re-fetched all seven pricing pages plus the OpenCode directory with zero price moves, dates bumped; context-management-patterns verified untouched (all links and sources live, no fact moved) [glm-5.3-flash]
  • _index.md: addy-osmani listed first in People and publications, matrix one-liners re-dated, thirteen-voice matrix line, one-liners refreshed (Cursor wind-down, Junie Local, DeepSeek 208k, Hermes 240k, claude-mem 93k), two banned-term slips in the index itself fixed (“honestly divided”, “shaping”) [glm-5.3-flash]
  • Queue: the standing ? for tom: about untracked the-agentic-development-environment-landscape/ remains unanswered, so it stays unlinked; no other actionable queue item, and no self-directed essay this run (the full re-verification plus the entrant filled the day at quality) [glm-5.3-flash]
  • Process: 19 sub-runs across 3 waves (4 relaunched after model rate-limit deaths); every sub-run confined to its own note and matrix files with no git commands [glm-5.3-flash]
  • Verification: 158 article files pass front matter parsing with no type field and mandatory tags (including research-note on all notes), 0 broken internal link targets, all 20 matrices match their categories’ membership with sorted headers and uniform rows, no em-dashes or banned terms in changed files, one sentence per line [glm-5.3-flash]

2026-09-02 (daily refresh, second same-day run: verification sweep + entrant scan + self-directed essay)
#

  • Context: the full parallel deep refresh already ran earlier today (entry above, all 20 categories re-verified, about 450 URLs fetched); re-fetching every URL hours later the same day would be churn, so this run verified state mechanically and spent its budget on a fresh entrant scan and one self-directed essay [glm-5.3-flash]
  • Verification sweep: 158 article files plus 5 control files pass front matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 matrices match their categories’ membership with case-insensitive sorted headers and no ragged rows; the only banned-term hit is ante’s protocol-shape crate name, the documented 2026-08-30 exception [glm-5.3-flash]
  • Entrant scan (HN Algolia by_date for the last week plus GitHub search of repos created after 2026-08-24 with 200+ stars): twelve-plus candidates surfaced, every one resolved in-run [glm-5.3-flash]
  • Rejected with evidence: experiential (820 stars, 220-point HN, Apache-2.0 open model gateway, active) and my-free-code (620 stars, same gateway pattern, quiet since 08-28) fit none of the twenty categories, logged plus a ? for tom: on a possible gateways category; sepia (1,444 stars in five days, de-AI writing skill) and headcount (1,031 stars, 16-department skills org chart for Claude Code) are community content rather than formats, registries, or marketplaces, so they stay out of Skills, logged plus a ? for tom:; acryl (237 stars, persistent cross-agent project context) sits below the context-engines traction bar and self-declares early development; opslane (39 stars) and conductai (30 stars) are below the bar of their nearest accepted peers [glm-5.3-flash]
  • Weak tail rejected without notes: PawWork (434, Chrome web agent, out of scope), hexstellar (642, created and last pushed the same day, out of scope), Code-as-World (378, research paper code), pentest-harness (326, security domain), useagent (273, three days old), SlideOps (22, docs-drift tooling, out of scope), 49 IDE and lemmalog unchanged from the 2026-09-02 morning run’s rejections; polign’s repo path now 404s on the GitHub API, reinforcing its earlier rejection [glm-5.3-flash]
  • Essay published: the-tells-are-structural, self-directed under the evaluation-and-review scope, the essay gate passed on all four conditions (named reader: the engineer whose LLM-assisted docs and PR replies get discounted as slop; no corpus article covers prose detection, the slop essays are about code ownership; three corpus see-also links; quality bar with disagreeable claims) [glm-5.3-flash]
  • the-tells-are-structural: grounded in StoryScope (arXiv 2604.03136, fetched this run: 61,608 stories, 304 narrative features, 93.2 percent narrative-only detection, 93.9 percent after surface rewriting, rarity-percentile 0.71 vs 0.49), the expert-detector study (arXiv 2501.15654, fetched: 1 of 300 majority-vote errors), LAMP (arXiv 2409.14509, fetched: the seven-category professional editing taxonomy), the slop-measurement study (arXiv 2509.19163, fetched), sepia’s repo and README (fetched), and Wikipedia’s signs-of-AI-writing page (fetched 200); six references [glm-5.3-flash]
  • _index.md: the essay listed under Essays and trackers [glm-5.3-flash]
  • queue.md: two ? for tom: bullets appended (breakout community skills, model gateways category); the standing untracked-landscape question remains unanswered [glm-5.3-flash]
  • No research notes written this run: every scan candidate is either a logged rejection or answered by the essay, and the index categories were fully re-verified this morning [glm-5.3-flash]
  • Verification of changed files: front matter parses with no type field, mandatory tags present, one sentence per line in body prose, no em-dashes, no banned terms (one shape slip in the essay caught and reworded to causal chain continuity mid-run), internal link targets exist on disk, all six reference URLs fetched 200 or via the GitHub API this run [glm-5.3-flash]

2026-09-02 (owner-directed: link references inline in the text) #

  • Owner directive in chat: link references directly in the body text where they are discussed, the arXiv studies specifically, and it applies to future articles too; recorded as standing practice [owner, glm-5.3-flash]
  • the-tells-are-structural: StoryScope, the human-detection study, the slop-measurement study, LAMP, sepia’s repo, and Wikipedia’s signs-of-AI-writing page now hyperlink at their first in-text mention; the Who Maintains the Slop? mention also became an inline link; every URL was already fetched 200 this run [glm-5.3-flash]
  • Verification: front matter parses, internal link targets exist on disk, one sentence per line, no em-dashes, no banned terms [glm-5.3-flash]

2026-09-03 (attribution note for the commit above)
#

  • Commit 314dd62c contains two runs’ work: the intended changes are the the-tells-are-structural inline links and the log entry; the remaining 38 files (verification-date bumps to 2026-09-03 plus re-verification edits across surfaces, harnesses, and orchestration notes and their matrices, including Gemini 3.8 Flash, cmux’s 50-VM cloud cap, conductor 0.84.0, Kiro 2.21.0 and 1.0.437, jcode 19k, and the paseo Hub GA) are the concurrent 2026-09-03 daily-refresh run’s in-progress edits, swept in because they landed in the working tree between this session’s stage and commit [glm-5.3-flash]
  • Those 38 files’ changes are attributed to that run [glm-5.3-flash]; it will record its own entry when it completes, per the 2026-08-29 concurrent-run precedent [glm-5.3-flash]

2026-09-04 (daily refresh, full parallel re-verification + memoryfields entrant)
#

  • Full parallel refresh under the no-stalest-ranking rule: 11 sub-runs re-verified all 135 research notes across the 20 categories against live sources (about 500 reference URLs fetched 200 or verified via the HN Algolia, GitHub, npm, PyPI, and crates.io APIs); every “Facts below verified as of” line bumped to 2026-09-04, notes with material revisions got updated plus the flash tag, bump-only notes got neither [glm-5.3-flash]
  • Harnesses material: ante v0.preview.94, claude-code open issues+PRs down to about 14.6k with 144k stars, deepseek-harness 211k stars and its first rc (dsh-v0.1.2-rc.1, 2026-09-03), fx 2.7k stars, jcode v0.81.6 with its embedding-cost caution rewritten (jcode.sh now publishes both embedding states, making the old caution factually wrong), onecli v2.5.0, zerostack v1.8.1 ending a six-week quiet gap with external contributors credited [glm-5.3-flash]
  • Surfaces material: continue’s final release corrected to 2.1.0-vscode (both tagged 2026-06-19, the note said 2.0.0), vscode-copilot’s self-serve pause reopened per the 2026-09-03 GitHub changelog, windsurf.com now 308-redirects to devin.ai/desktop so the four-month-old 429 note became a redirect note [glm-5.3-flash]
  • Elsewhere material: cmux cloud VM disk cut 200GB to 32GB, qwenpaw v2.2.0 went stable on PyPI, augment-code added a $20/month Standard plan, graphify crossed 114k stars, semble downloads re-measured 109,770/month, opencode-skills docs dates moved, skills-sh leaderboard reshuffled (new #4 improve-codebase-architecture), claude-mem v13.24.0, github-agentic-workflows v0.88.4, deepeval 4.2.1, phoenix 20.7.0, plannotator v0.27.12, coderabbit agent-minute price cut $0.50 to $0.40, open-code-review v1.11.3, agent-sandbox v1.0.1, spec-kit v1.0.4 and 133k stars, backlog-md and beads status sentences refreshed (beads v1.3.0-rc.1 prerelease noted), har v1.12.0, ouroboros status refresh [glm-5.3-flash]
  • People and publications material: latent-space’s AINews link replaced (news.smol.ai now 402 DEPLOYMENT_DISABLED, replaced with latent.space/s/ainews), simon-willison cadence refreshed through 2026-09-04, deeplearning-ai’s Batch extended past its “through August 2026” claim, caleb-writes-code counts refreshed (107,710 subs, 116 videos, sponsor ratio 9 of 15) [glm-5.3-flash]
  • Matrices: all 20 re-verified in the same run, dates bumped to 2026-09-04, cells updated only where facts moved (surface Continue cell, orchestration Omnara stars, assistant-runtimes stars/status rows, code-review Kodus/OpenCodeReview, sandboxing Agent Sandbox/OpenShell, spec-driven adoption row, control-planes and software-factory status rows, context-engines Augment pricing, evaluation Phoenix maturity) [glm-5.3-flash]
  • Entrant resolution, one accepted: memoryfields note added to Memory (Cal Paterson’s memory-as-file-format, 191-point HN thread 2026-08-31, draft spec v0.1, AGPL-3.0 tool and MIT skill both under 80 stars combined, the attention-versus-adoption gap recorded as the status signal); six sources fetched this run [glm-5.3-flash]
  • memory-feature-matrix extended from six to seven columns (Memoryfields inserted between mem0 and Zep) in the same run, ten new cells traced to the note, reading and choosing sections extended [glm-5.3-flash]
  • Rejections with evidence: useagent (283 stars, workspace around existing harnesses, no category fit and no independent coverage, second look after the 2026-09-02 rejection), kleisli-io/kli (32-point HN, zero comments), mezmo/aura (incident response, out of scope), Proval (65 stars, 13-point HN), stanford-mast/blast (799 stars but a 2025 repo repositioned, 11 points), Wasmer local sandboxes (5 points, unchanged from 2026-09-02), WikiSkill (102 stars, Hermes-tied, thin discussion), OrcaReplay (89 stars, below bar), agent-session-protocol (46 stars), coleam00/ai-software-factory (130 stars, no license, three days old, same bar that rejected ai-factory) [glm-5.3-flash]
  • Essays: landscape as-of bumped to 2026-09-04 with the Windsurf domain handover folded into its surfaces line; model-selection re-fetched all 11 references, zero price moves, Claude Fable 5.1 folded in (same $10/$50, cache hits at 2.5% of input versus the standard 10%, table row and cache bullet updated); context-management-patterns and the-tells-are-structural link-checked clean, untouched [glm-5.3-flash]
  • _index.md: memoryfields listed in Memory, thirteen one-liners re-dated or refreshed (DeepSeek 211k/three weeks, Spec Kit 133k, Windsurf redirect, memory matrix seven approaches) [glm-5.3-flash]
  • Process note: the 2026-09-03 concurrent refresh never recorded its entry; its committed work is attributed per the 2026-09-03 attribution note in this log [glm-5.3-flash]
  • Verification: 165 section files pass YAML front-matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 matrices match their categories’ membership with sorted headers and uniform rows, no em-dashes or banned terms in article files (one shape slip in the new memoryfields note caught and reworded mid-run), one sentence per line, and every URL cited by new or edited content fetched 200 or was verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run [glm-5.3-flash]

2026-09-05 (daily refresh, full parallel re-verification + six entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: 11 sub-runs re-verified all 142 research notes across the 20 categories against live sources (about 450 reference URLs fetched 200 or verified via the HN Algolia, GitHub, npm, PyPI, and crates.io APIs); every “Facts below verified as of” line bumped to 2026-09-05, notes with material revisions got updated set (every touched note already carried llm=glm-5.3-flash, so no tag appends were needed) [glm-5.3-flash]
  • Harnesses material: claude-code open issues+PRs down to about 14.1k, cline 67.5k stars and 5.22M installs, codex 121.6k, deepseek-harness 212.5k stars with dsh-v0.1.3-alpha.1 (2026-09-04, still no stable), kilo-code v7.5.14, pi v0.85.0, jcode v0.81.7, opencode 204.3k, openhands 86.2k; aider stall re-confirmed, amp/junie/onecli/warp/zerostack/crush/gemini-cli/goose/qwen-code verified unchanged [glm-5.3-flash]
  • Surfaces and code review material: openchamber v1.22.1 (Sep 4), continue 4.07M installs, zed 89.8k, roo-code installs still ticking post-sunset; qodo repository canonical URL moved to The-PR-Agent/pr-agent (org move, old path redirects), kodus added Teams $8/dev annually alongside the $10 monthly, open-code-review v1.11.5, sourcery and ellipsis active pushes; coderabbit, greptile, graphite-diamond pricing and claims verified unchanged [glm-5.3-flash]
  • Orchestration and session analytics material: cmux pricing restructured again (shared 5 vCPU/20 GB/200 GB pool, VMs starting at 8 GB/32 GB), conductor 0.84.2, superset v1.26.0, agentsview now claims 60-plus agent formats (note and matrix cell updated from roughly 55); vibe-kanban orphan state and crystal deprecation re-confirmed [glm-5.3-flash]
  • Protocols, skills, executions material: a2a v1.0.1 release date corrected to 2026-05-28 per the GitHub API, agent-host-protocol npm reference moved to an explicit window URL (last-month endpoint serves stale data) with 302 stars, skills-sh audits caution rewritten (the Snyk-Critical azure-validate entry no longer appears; 81 pending, 23 safe), github-agentic-workflows v0.88.4 confirmed latest [glm-5.3-flash]
  • Context engines, retrieval material: rtk v0.48.0 plus a factually wrong caution fixed (the issue backlog is about 2,100 combined, not star-sized), semble v0.5.6, graphify v0.9.54, qmd/repomix/llamaindex/langchain drift refreshed; augment and sourcegraph pricing re-verified byte-for-byte [glm-5.3-flash]
  • Memory and hybrid execution material: claude-mem v13.24.1 (published 2026-09-05), cognee v1.5.4, mem0 free tier relabeled Hobby with identical quotas, memoryfields crossed the 85-star mark (its under-80 claim reworded), letta internal star inconsistency reconciled; instructor, outlines, both structured-output notes verified unchanged [glm-5.3-flash]
  • Assistant runtimes and evaluation material: openclaw v2026.9.1 (Sep 3) after the 2.0 line, hermes 242k stars, eigent v1.0.4, openwork pricing moved again (Solo free, Team Starter $10/seat with first five seats free, Enterprise custom), phoenix 20.8.0; nanoclaw, nanobot, picoclaw, qwenpaw, zeroclaw, deepeval, plannotator, workshop verified unchanged [glm-5.3-flash]
  • Small categories material: bmad-method v6.12.0 (Sep 4), beads v1.3.0-rc.1 prerelease date corrected to 2026-08-31, paperclip crossed 80k stars; task-master, tinyagi, sssf stalls re-confirmed; all five sandboxing and software-factory members verified current [glm-5.3-flash]
  • People and publications material: ai-jason 231K subscribers, caleb-writes-code 108K subs and 117 videos with the sponsor recount corrected to 9 of 15, The Batch at issue 369 (Sep 4), simon-willison posting through Sep 5; Armin Ronacher resolved as a rejection (duplicates the hands-on-text band, not the enterprise-operator slot) [glm-5.3-flash]
  • Essays: landscape as-of bumped to 2026-09-05 (Orca 61.9k the only moved number, Cline installs corrected to 5.2M); model-selection re-fetched all eleven references with zero price moves, Kimi references canonicalized to platform.kimi.ai and the DeepSeek URL fixed to the pricing path; context-management-patterns and the-tells-are-structural link-checked clean [glm-5.3-flash]
  • Entrant resolution, six accepted, each written to citation standards with all sources fetched this run: kimi-code (Moonshot’s MIT CLI, about 18.6k combined stars across kimi-code plus the Apache-2.0 kimi-cli), exo (Exo Labs’ self-modifying MIT harness, 1,338 stars, FrontierHarness’s $1.05-per-task cheapest), frontierharness-eval (Runta’s nine-harness same-model benchmark, 81-point HN, unlicensed-repo and vendor-conflict cautions), clawk (disposable per-agent Linux VMs on macOS, 226-point launch, pre-1.0 quiet since August), machinist (Go factory with the category’s strictest named-command boundary, 322 stars, zero independent coverage stated), ctx (local session search with pro blame attribution, 1,077 stars, fills the session-analytics scaffold’s second column) [glm-5.3-flash]
  • oh-my-pi (about 29.6k stars) stays covered inside the pi note per the gemini-cli/Antigravity grouping precedent, logged so future runs do not re-litigate; reverify, m3e-canvas, fable-orchestrator, reef, kit, kli, VTCode, funes, lemmalog, moadim, Dormice, and roughly twenty further candidates rejected with evidence in-run [glm-5.3-flash]
  • Matrices: harness 23 to 25 columns (Exo, Kimi Code inserted alphabetically, 12 rows by 26 cells verified), evaluation 4 to 5 (FrontierHarness Eval), sandboxing 5 to 6 (Clawk, now three boundary columns), software factory 4 to 5 (Machinist), session analytics 1 to 2 (ctx); all 20 matrices re-dated to 2026-09-05 with cells updated only where facts moved [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered, so nothing changed on them [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus six entrant notes with their matrix columns filled the day at quality [glm-5.3-flash]
  • Verification: 171 section files pass YAML front-matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 matrices match their categories’ membership with case-insensitive sorted headers and no ragged rows, no em-dashes, one sentence per line; one banned-term slip in the new machinist note (“Honest”) caught by the pre-commit sweep and reworded; every URL cited by new or edited content fetched 200 or was verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run [glm-5.3-flash]

2026-09-06 (daily refresh, full parallel re-verification + happy-coder entrant)
#

  • Full parallel refresh under the no-stalest-ranking rule: 12 sub-runs across 3 waves re-verified all 142 research notes across the 20 categories against live sources (GitHub API and repo pages, pricing pages, npm/PyPI/crates.io, HN via the Algolia items API); every “Facts below verified as of” line bumped to 2026-09-06, notes with material revisions got updated set (every touched note already carried llm=glm-5.3-flash, so no tag appends were needed) [glm-5.3-flash]
  • Harnesses material: claude-code open issues+PRs down to about 13.6k, cline 5,224,030 installs, codex about 121.8k, deepseek-harness 213.4k, exo’s FrontierHarness sentence rewritten (Exo no longer last on pass rate with OpenCode and Hermes at 50.0 percent, $1.05 per task still cheapest, the old one-third fraction was mathematically wrong), fx’s site now claims 6.39 MiB, jcode v0.82.0 (2026-09-06), juggler v0.6.0, kilo-code v7.5.15, kimi-code docs confirm the Node.js rebuild with the Python version no longer maintained and the measured-cost comparison replaced (Codex $3.47 versus Kimi $3.65 median per task, quality lead 66.7 percent stands), opencode 204.8k, openhands 86.3k, pi v0.85.1, zerostack pushed externally, aider stall re-confirmed [glm-5.3-flash]
  • Surfaces and orchestration material: openchamber v1.22.2 (Sep 5), cmux’s pricing page dropped the shared-pool cloud spec in favor of per-VM resources with a 256 GB disk ceiling, qodo’s PR-Agent shipped v0.45.0, gastown/paseo/superset/worktrunk/omnara drift refreshed, crystal and vibe-kanban dead/orphan states re-confirmed [glm-5.3-flash]
  • Elsewhere material: claude-mem rebranded to “Grok Mem” on its README (npm package unchanged) and added Cursor plus Grok Bot marketplace plugins, cognee’s contradiction handling now documented on the pricing page as Enterprise BYOC, anthropic-structured-outputs docs dropped the Microsoft Foundry limit and added Google Cloud plus 7+ SDK surfaces, ctx published pro pricing ($20/month with 14-day trial), openclaw v2026.9.2 turned Swarm on by default, machinist 344 stars with the zero-coverage claim re-dated, caleb-writes-code 108,181 subs, latent-space over 199,000 subscribers, nathan-lambert’s Interconnects over 81,000 [glm-5.3-flash]
  • Essays: landscape as-of bumped to 2026-09-06 with the Kimi single-Node.js-CLI correction folded in; model-selection re-fetched all pricing references with zero price moves (new listings gpt-6-astra, Gemini 3.8/3.5 Flash, GLM-5.3-Flash promo noted, no asserted claim moved) and dates bumped; context-management-patterns and the-tells-are-structural link-checked clean, left untouched [glm-5.3-flash]
  • Entrant resolution: happy-coder accepted into Orchestration (the MIT encrypted mobile client wrapping Claude Code and Codex, 23,657 stars, created July 2025, cli-1.2.3 of 2026-09-05, out-stars every category member against a 30-point Show HN, 7 sources fetched this run); orchestration-feature-matrix extended from 12 to 13 columns in the same run [glm-5.3-flash]
  • Entrant rejections with evidence: hax (rejection stands, 725 stars, pushed 2026-09-04, no new coverage), smol-env/smol (134 stars, quiet since 2026-08-15, zero HN), Codex App and OpenCode Desktop (both covered inside the codex and opencode notes per the vendor-family precedent), LiteParse (folded into the llamaindex note, 12,257 stars, same org), Warp Factories (folded into the warp-agent-cli note, closed early access), skills-sh’s enterprise publishers (folded into the skills-sh note), Spotify Portal (257 points but a vendor piece whose open artifact is a hook-plus-skill pair, below the bar), OKF Agent Memory (66 points on day one, 146 stars, below the memoryfields entry bar, re-review on a second implementation or independent coverage), plus roughly forty weak-tail candidates below 1-7 points or sub-100-star traction [glm-5.3-flash]
  • Maintenance folds in the coordinator pass, URLs fetched this run: openclaw gained the Swarm default-on sentence and releases reference, skills-sh gained the publisher mix (open.feishu.cn 15.1M aggregate, prime-skills/runcomfy-agent-skills 4.5M, larksuite/cli 2.2M-plus), llamaindex gained the LiteParse sentence and repo reference, warp-agent-cli gained the Factories sentence and reference [glm-5.3-flash]
  • _index.md: happy-coder listed in Orchestration between Gas Town and Omnara, orchestration matrix one-liner now thirteen with Happy Coder the newest, kimi-code one-liner corrected to Node.js CLI [glm-5.3-flash]
  • Matrices: all 20 re-verified and re-dated to 2026-09-06; orchestration 12 to 13 columns (Happy Coder inserted alphabetically), memory gained the grounded cognee bi-temporal cell, hybrid-execution the Anthropic 7+ SDK cell, session-analytics the ctx pricing cell, sandboxing the OpenShell contributors cell, spec-driven the adoption row, assistant-runtimes stars and status cells, control-planes the Paperclip issues cell; harness cells unchanged (the matrix carries no star or version rows), dates bumped [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus the happy-coder note with its matrix column filled the day at quality [glm-5.3-flash]
  • Verification: 172 section files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 matrices match their categories’ membership with case-insensitive sorted headers and no ragged rows, no em-dashes or banned terms in article files (ante’s protocol-shape crate name remains the documented 2026-08-30 exception), one sentence per line, and every URL cited by new or edited content fetched 200 or was verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run [glm-5.3-flash]

2026-09-07 (daily refresh, full parallel re-verification, no entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: 12 sub-runs across 3 waves re-verified all 142 research notes across the 20 categories against live sources (GitHub API and repo pages, pricing pages, npm/PyPI/crates.io, HN via the Algolia items API); every “Facts below verified as of” line bumped to 2026-09-07, notes with material revisions got updated set (every touched note already carried llm=glm-5.3-flash, so no tag appends were needed)
  • Harnesses material: claude-code open issues+PRs down to about 13.2k, codex about 122.1k, opencode 205.4k, openhands 86.4k, cline installs 5,234,440, deepseek-harness 214.4k, jcode v0.84.0 (published 2026-09-07), zerostack v1.8.3 (Sep 6) with crates downloads 3,353, bullet’s npm traction declining (weekly downloads 654 to 214), fx 2.8k, exo and goose pushed the day of verification; aider stall re-confirmed
  • Surfaces and orchestration material: continue installs 4,086,305, roo-code archived drift from removals (24,309 stars), vscode-copilot vscode about 191k, openchamber push Sep 6; paseo 16.3k, gastown 17,951, happy-coder 23,673, cmux 26.9k, superset 13.8k, worktrunk 6.9k, claude-squad 8,441; conductor, crystal, and vibe-kanban verified unchanged (a sub-run over-set updated on them and it was reverted to the prior dates, per the updated-moves-only-on-material-revision policy)
  • Elsewhere material: hermes open issues 40.5k, openwork’s pricing page now lists Enterprise at $40/user/month and renamed Team Starter to Team, picoclaw pushed Sep 3, zeroclaw issues down to 790, nanobot issues 770, rtk open+PRs about 1,830, semble downloads 96,304/month, langchain 145.8k, instructor about 24M monthly PyPI downloads, openai-structured-outputs docs now lead their examples with gpt-5.6 plus gpt-6-astra, github-agentic-workflows v0.88.5 (Sep 7), skills-sh leaderboard moved (open.feishu.cn 16.8M, runcomfy 11.0M, larksuite about 10.9M; azure-validate is back on the page now marked Safe, audits 78 Pending/24 Safe), a2a 25.7k, agents-md 24.2k, agent-host-protocol 306 stars, anthropic-agent-skills 174.9k, opencode-skills docs-stability claim re-grounded via git history, deepeval v4.2.2, phoenix 11.4k, plannotator 1,165 commits, openspec npm 1,655,188/month, bmad-method 52.7k, paperclip 80.1k with issues 5,379, ouroboros v0.54.0 (published today), machinist 366 stars, task-master npm 68,476/month, caleb-writes-code 109k subscribers, beads 26,955
  • Essays: landscape folded the Orca count in (about 60k to about 63k stars on the ade topic) and bumped as-of plus updated to 2026-09-07; model-selection re-fetched all eleven pricing and benchmark references with zero price moves (OpenAI’s page fetched 200 with luna prices and the November 21 promo date intact; the Kimi platform pages fetch but do not server-render prices, so those two rows keep their 2026-09-06 measurement) and re-dated the lineup heading; context-management-patterns and the-tells-are-structural link-checked clean, left untouched
  • Matrices: all 20 re-verified in the same run and re-dated to 2026-09-07, cells updated only where facts moved (orchestration Omnara stars and Gas Town rounding, assistant-runtimes Hermes and PicoClaw cells, retrieval LangChain stars, evaluation FrontierHarness 82 points, spec-driven adoption row, control-planes Paperclip issues and long-tail prose, software-factory Machinist and SSSF stars); harness matrix cells unchanged (no capability-row movement), verified line and updated field bumped
  • Entrant resolution, all rejected with evidence: engrim (19-point HN with 3 comments, 27 stars, far below the memory category’s entry bar set by memoryfields at 191 points and OKF’s rejection at 66), Ponytail (32 points, a community “lazy senior engineer” skill, out of the Skills scope per the standing formats/registries/marketplaces scoping awaiting the owner’s answer), MathKernel (30 points, a domain-specific mathematics MCP server, not an agent development environment tool); GitHub sweeps of repos created after 2026-09-01 with 150-plus stars surfaced nothing credible (a Fable 5.1 demo-world gallery and an awesome list)
  • URL fixes: junie’s devclass reference canonicalized, eigent’s cowork link moved to claude.com/product/cowork, zeroclaw’s docs domain moved to docs.zeroclaw.com, ctx’s comparisons link replaced with the per-topic path (old path 404s)
  • _index.md: eleven one-liners re-dated to 2026-09-07 (both essays, nine matrix lines)
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered
  • No self-directed essay this run: the full re-verification filled the day at quality, and no scan candidate cleared the bar
  • Process notes: one sub-run ran a read-only git diff against its no-git instruction (changed nothing, reported the violation itself); the working tree’s pre-existing owner changes outside agents/ (including a deleted article-marketing/index.md not made by any run) were left untouched
  • Verification: 167 article files pass front-matter parsing with no type field and mandatory tags, 1,285 internal link targets all resolve on disk, all 20 matrices match their categories’ membership with case-insensitive sorted headers and uniform rows, no em-dashes, one sentence per line, ante’s protocol-shape crate name remains the documented 2026-08-30 exception, and every URL cited by new or edited content fetched 200 or was verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run [glm-5.3-flash]

2026-09-08 (daily refresh, full parallel re-verification, no entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: 14 sub-runs across 4 waves re-verified all 143 research notes across the 20 categories plus the two living essays against live sources (GitHub API and releases, pricing pages, npm/PyPI/crates.io, HN via the Algolia items API); every “Facts below verified as of” line bumped to 2026-09-08, notes with material revisions got updated set (all notes already carried llm=glm-5.3-flash, so no tag appends were needed)
  • Harnesses material: amp’s September news folded in (desktop orb client September 4, Fable 5.1 powering ultra September 1, iOS and macOS app August 28); claude-code open issues+PRs down to about 12.9k with two feature-timeline months corrected against the cited product page (routines May not April, computer use April not March); cline installs 5,247,005 and the $4.99 first-month ClinePass promo no longer stated on the page (revised to “promotional rate”); deepseek-harness 215,497 stars with the dsh-v0.1.3-alpha.2 prerelease (September 7); fx v0.0.8 (September 7, breaking shell and subagent changes) with the binary now 6.01 MiB; onecli v2.6.0 (September 8); zerostack v1.8.4 (September 7); opencode 205.8k, openhands 86.8k, pi 102,894 with oh-my-pi past 30k; aider’s stall re-confirmed; exo 1,363 stars
  • Exo correction this run: a sub-run rewrote the FrontierHarness sentence to “twelve harnesses” (misreading the site’s “9 harnesses across 12 configurations” as an expansion); caught in the coordinator pass against frontierharness.org and reverted to nine harnesses, matching the FrontierHarness Eval note
  • Surfaces and orchestration material: omnara now publishes usage-based pricing ($0 platform fee, token-rate models, $0.0414 per GiB-hour machine runtime, idle auto-sleep), the pricing page added as an eighth source and the matrix cell updated; paseo 16,440 stars with open issues+PRs dropping 1,241 to 816 in a one-day cleanup wave; superset desktop-v1.27.0 (September 7); vscode-copilot and windsurf references canonicalized (VS Code docs reorg to /docs/agents/overview, cognition.ai to cognition.com, docs.windsurf.com to docs.devin.ai); happy-coder 23,693 and emdash drift refreshed
  • Protocols, skills, context, executions, retrieval material: agent-host-protocol 312 stars with commits landing 2026-09-08 and crates downloads 205,919; skills-sh CLI at 30.6k with the leaderboard reshuffled (feishu 14.7M aggregate, runcomfy 4.5M, larksuite 6.0M by re-counted publisher sets); skillopt 16,769; anthropic-agent-skills 175.1k; graphify v0.9.56 (September 7) at 115,802 stars; repomix npm window 329,262; github-agentic-workflows v0.88.6 (September 7); langchain 145.9k, llamaindex 52.1k with liteparse 12,269; anthropic-structured-outputs docs now list Opus 4.7/4.8 and Mythos Preview as supported; agentsview’s 84-223x ccusage benchmark removed from its docs entirely (prose, caution, and reference revised, one redirect fixed)
  • Memory, control, task, review, eval material: claude-mem 93.4k; cognee 30.6k at v1.5.4; paperclip 80,209 stars and 5,392 issues (control-planes matrix cells updated, TinyAGI 3,610); beads 26,969; backlog-md 6,668; the graphite.dev domain now redirects to graphite.com (five vendor reference URLs and matrix wording updated); open-code-review v1.11.6 (September 7) at 22,067 stars; deepeval docs moved from docs.confident-ai.com to deepeval.com; sourcery and qodo reference canonicalizations; frontierharness reference point count corrected 81 to 82
  • Sandbox, factory, spec, runtimes, people material: har v1.14.1 (September 7); ouroboros v0.54.1 (September 8); tessl published platform pricing for the first time (Free $0 with 1,000 credits, Team $100 per month, Enterprise custom; Pricing section rewritten, page added as seventh reference, matrix cell updated); spec-driven adoption row and factory stars row refreshed; hermes v2026.9.7 at 243k stars; openclaw 389,184; nanobot 47,870 with picoclaw’s inspiration count corrected; qwenpaw 35,043; zeroclaw 32,743; caleb-writes-code 109,623 subscribers; simon-willison’s recency claim refreshed through September 7
  • Essays: landscape as-of bumped to 2026-09-08 with Orca at about 64k and OpenHands at about 87k; model-selection re-fetched all sixteen tabled prices against the live provider pages with zero price moves (the GLM-5-Turbo clause dropped, it no longer appears on Z.ai’s pricing page) and dates bumped; context-management-patterns and the-tells-are-structural link-checked clean, left untouched
  • Matrices: all 20 re-verified in the same run and matched to their categories’ membership with dates bumped; cell updates where facts moved (orchestration Omnara, protocols AHP stars and combined downloads, retrieval LangChain and LlamaIndex, code-review Graphite domain plus Kodus, OCR releases, Sourcery, control-planes Paperclip issues and TinyAGI stars, spec-driven adoption row and Tessl pricing cell, software-factory stars row, assistant-runtimes Hermes, OpenClaw, PicoClaw, ZeroClaw cells, session-analytics benchmark-removal prose); harness, surface, skills, executions, context-engines, sandboxing, and people matrices re-verified with no cell movement
  • Entrant resolution, all rejected with evidence: the cross-run candidate set stayed stable (sepia, headcount, my-free-code, useagent, Code-as-World, vicoa, pi-posthorse all wrong category or weak tail; Moadim.io at 50 stars below the executions bar; zhizhi-agent-runtime a category misfit; machine0 closed and out of category; hello-sdd a course; Dhorthy/HumanLayer essays have no sustained publication and the reader slot is occupied); okf-agent-memory re-reviewed per the standing calibration and rejected again (same repo and thread as the 66-point rejection, no second implementation or independent coverage); no new note in any category
  • _index.md: eleven dated one-liners re-dated to 2026-09-08; the FrontierHarness one-liner extended with “in twelve configurations” to match the note
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered
  • No self-directed essay this run: the full re-verification filled the day at quality, and no scan candidate cleared the bar
  • Process notes: two sub-runs ran a read-only git status --porcelain against their no-git instruction (changed nothing, both self-reported); the working tree’s pre-existing owner changes outside agents/ (the deleted article-marketing/index.md and untracked drafts) were left untouched
  • Verification: 172 files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, all 20 matrices match their categories’ membership with case-insensitive sorted headers and uniform rows, no em-dashes or banned terms in article files (ante’s protocol-shape crate name remains the documented 2026-08-30 exception), and every URL cited by new or edited content fetched 200 or was verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run (Bloomberg, oreilly.com, and thoughtleaders.io bot walls documented as fetch artifacts, not dead links) [glm-5.3-flash]

2026-09-08 (owner-prompted, model-selection table transposed)
#

  • Owner request in chat: the model-selection-for-coding-tasks lineup table transposed, models moved from sixteen rows to sixteen columns so the table reads as two attribute rows (input/output per 1M, notes) against the model columns; prices, notes, and the provider-grouped column order unchanged from the verified 2026-09-08 state [glm-5.3-flash]
  • No facts changed, so the verified line and updated field stay at 2026-09-08; the wide-table site component (pinned label column, column picker) renders the seventeen-column layout without content-side markup [glm-5.3-flash]

2026-09-08 (owner-prompted, model-selection table property rows)
#

  • Owner request in chat: more properties in the transposed lineup table; the single Notes column decomposed into four property rows (Cached input per 1M, Context behavior, Pricing windows, Notes), every cell grounded in the article’s verified prose (cache ratios from the cached-reads bullet, context behavior from the whole-repo bullet, windows from the promo and off-peak facts) with no new facts introduced [glm-5.3-flash]
  • A ? cell plus a footnote sentence marks properties the verified pricing pages do not state (Haiku 4.5, Gemini Flash-Lite, and GLM-5.3 context); updated and the verified line stay at 2026-09-08 since no fact moved [glm-5.3-flash]

2026-09-08 (owner-prompted, one-year release window check)
#

  • Owner request in chat: keep only models released within the last year; checked all sixteen lineup entries against launch-day coverage fetched this run via the HN Algolia search API, and every model passes the 2025-09-08 cutoff, so no column was dropped [glm-5.3-flash]
  • Release dates recorded: gpt-5.6 family July 2026 (staggered public launch after the June 25 government-delay story), gpt-5.3-codex 2026-02-05, Haiku 4.5 2025-10-15 (the oldest, 37 days inside the window), Sonnet 5 2026-06-30, Opus 5 2026-07-24, Fable 5.1 2026-09-01, Gemini 3.1 family 2026-02 (Pro Preview 2026-02-19), Gemini 3.6 Flash 2026-07-21 (the grouped 3.7/3.6 row dated to 3.6), kimi-k2.7-code 2026-06-12, kimi-k3 2026-07-16, GLM-5.3 2026-08-14, DeepSeek V4 pair 2026-04-24 [glm-5.3-flash]
  • Table: a Released row added as the first property row, with month-precision cells (2026-07, 2026-02) where only the launch month is documented; a second footnote sentence states the one-year lineup rule and the oldest entry [glm-5.3-flash]
  • Verification: table uniform at 17 cells across all eight lines, 0 broken internal links, front matter and style clean; updated and the verified line stay at 2026-09-08 since no price fact moved [glm-5.3-flash]

2026-09-08 (owner-prompted, GLM-5.3-Flash added to model selection)
#

  • Owner request in chat: include GLM-5.3-Flash; fetched the live Z.ai pricing page via its markdown alternate (the HTML path serves a client-rendered shell) and the HN Algolia launch coverage, both this run [glm-5.3-flash]
  • Table: an eighteenth column added between GLM-5.3 and deepseek-v4-pro, all cells grounded in the fetched page: released 2026-08-26 (within the one-year lineup rule), $0.075/$0.25 input/output under a 50 percent promotion (list $0.15/$0.50), cached reads $0.015, context not stated (?), pricing window “50% promo ends 2026-09-09 (UTC+8)”, notes “challenger cheap tier” [glm-5.3-flash]
  • Prose coherence: the cheap-tier task-class bullet and the inline-completion decision bullet gained the model; the challengers section gained a sentence for it (the gap between free GLM-4.7-Flash and $1.40 GLM-5.3); the price-category sentence now credits its $0.075 promo rate before the permanent v4-flash beat; the high-volume decision bullet leads with it; What-changes-fast now counts four clocks with the 2026-09-09 promo deadline first (tomorrow, the nearest in the guide) and What to Do Next calendars it [glm-5.3-flash]
  • Verification: table uniform at 18 cells across all eight lines, 0 broken internal links, front matter and style clean; updated and the verified line stay at 2026-09-08 [glm-5.3-flash]

2026-09-08 (owner-prompted, GPT-6 Astra added to model selection)
#

  • Owner request in chat: include GPT-6 Astra; fetched the live OpenAI pricing page (standard, batch, and long-context row sets for gpt-6-astra all present in the served data, structurally matched against the gpt-5.6-sol rows) and the HN Algolia launch coverage, both this run [glm-5.3-flash]
  • Table: a nineteenth column added after gpt-5.3-codex, closing the OpenAI group: released 2026-09-03 (within the one-year lineup rule), $10/$50 input/output, cached reads $1 (one tenth, matching the vendor-wide cached-reads bullet), context doubling past the long-context threshold, no pricing window recorded, notes “frontier, newest OpenAI generation” [glm-5.3-flash]
  • Prose: the frontier decision bullet now leads with gpt-6-astra; no other asserted claim moved (the batch and cached-read bullets already state the vendor-wide policies the new column falls under) [glm-5.3-flash]
  • Verification: table uniform at 19 cells across all eight lines, 0 broken internal links, front matter and style clean; updated and the verified line stay at 2026-09-08 [glm-5.3-flash]

2026-09-08 (owner-prompted, Qwen models added to model selection)
#

  • Owner request in chat: include Qwen models; the corpus covered the qwen-code harness but no Qwen API model, so two columns joined the challengers: qwen3.8-max and qwen3.8-flash, the current Alibaba text-generation lineup per the Model Studio catalog [glm-5.3-flash]
  • Sources fetched this run: Alibaba’s official Model Studio billing page (the international docs host bot-walls automated fetches; the Chinese help center serves server-rendered HTML, qwen3.8-max at 12/36 CNY domestic and 14.988/44.965 CNY international, qwen3.8-flash at 1.094/3.427 CNY international, both in a single 0-1M token tier, batch half price on Max, context-cache discount stated without a rate, a 1M-token 90-day free quota for new accounts) and the OpenRouter models API (USD listings: qwen3.8-max-0902 $2/$6, qwen3.8-flash $0.15/$0.47, 1M context, matching the international CNY rates); both added as references; release dates from HN launch coverage (Qwen3.8-Max 2026-08-03, the flash line’s Flash-Next launch 2026-08-26, month precision for the API flash) [glm-5.3-flash]
  • Table: two columns added after deepseek-v4-flash (challenger flagship and challenger cheap tier), cached-rate cells left ? per the unpublished discount rate, and the footnote generalized since ? now appears outside the context row [glm-5.3-flash]
  • Prose coherence: the challengers bold line names Qwen fourth; a new sentence adds both models with the flat-1M flagship point; the harness-fit paragraph notes Alibaba is not a native OpenCode provider (verified against the live providers page) but rides OpenRouter or OpenAI-compatible endpoints; the batch bullet now says OpenAI, Anthropic, Google, and Qwen3.8-Max instead of “all three providers”; the cached-reads bullet gains the unpublished-Qwen-rate clause; the whole-repository decision bullet adds Qwen3.8-Max to the $2 flat-1M slot (all four keep 1M pricing flat) [glm-5.3-flash]
  • Verification: table uniform at 21 cells across all eight lines, 0 broken internal links, front matter and style clean; updated and the verified line stay at 2026-09-08 [glm-5.3-flash]

2026-09-08 (owner-prompted, model provider feature matrix)
#

  • Owner request in chat: a matrix comparing the model providers (OpenAI, Anthropic, Google, Zhipu, Alibaba, etc.); created agents/model-provider-feature-matrix with seven provider columns sorted alphabetically (Alibaba (Qwen), Anthropic (Claude), DeepSeek, Google (Gemini), Moonshot AI (Kimi), OpenAI (GPT), Zhipu (GLM)) [glm-5.3-flash]
  • Design note: the research-index matrices tie columns to member notes, but no provider notes exist, so this is a provider-level companion matrix instead: column headers link to the canonical pricing pages and every cell traces to the model-selection guide’s verified 2026-09-08 table or to pages fetched this run (Alibaba billing page, OpenRouter models API, OpenCode providers page, Z.ai pricing via the markdown alternate) [glm-5.3-flash]
  • Rows: flagship, workhorse, and cheap model with prices, coding-specialized model, cached input discount, long-context policy, batch or time discounts, live pricing windows, weights (open or closed), subscription-includes-an-agent, and BYOK harness fit; the reading section frames the bundle split (big three win subscriptions and closed frontier, challengers win price, weights, and cache depth) and the choosing section gives six situation bullets [glm-5.3-flash]
  • _index.md: the new matrix listed first under Comparison matrices (the provider layer above the category matrices); model-selection-for-coding-tasks cross-linked it as the first See-also entry [glm-5.3-flash]
  • Verification: matrix table uniform at 13 rows x 8 cells with case-insensitive sorted headers, 173 article files pass front-matter parsing, tags, style, and 0 broken internal link targets across the section [glm-5.3-flash]

2026-09-08 (owner-prompted, models.dev as the model list reference)
#

  • Owner request in chat: use models.dev as the model list reference; fetched https://models.dev/api.json (213 providers, 4.5 MB) and cross-checked the whole lineup: all 20 tabled models appear on models.dev, which carries ids, release dates, and context windows but no prices for these, so it becomes the list reference while prices stay provider-sourced [glm-5.3-flash]
  • Released row refined to exact models.dev dates: the GPT-5.6 family to 2026-07-09, GPT-6 Astra to 2026-09-04 (the HN rollout threads ran September 3), Claude Sonnet 5 to 2026-06-29 (launch threads June 30), Gemini 3.1 Flash-Lite corrected 2026-02 to 2026-05-07 (February was the Pro Preview family), DeepSeek V4 Pro and Flash to 2026-08-12 and 2026-07-31 (April 24 was the family announcement; models.dev dates the API model ids, matching the deepseek-v4-flash-0731 snapshot in Alibaba’s catalog), Qwen3.8-Flash to 2026-08-26 [glm-5.3-flash]
  • Context row: the four ? cells filled from models.dev limits, Haiku 4.5 at 200K, Gemini Flash-Lite, GLM-5.3, and GLM-5.3-Flash all 1M flat; the footnote sentence rewritten to name models.dev as the dates-and-contexts reference [glm-5.3-flash]
  • New sibling folded in: models.dev lists gemini-3.8-flash (2026-09-02) and Google’s pricing page (fetched 200 this run) confirms it shares the 3.7/3.6 intro pricing of $0.75/$3.75 through 2026-12-31, then $1.50/$7.50, so the grouped column widened to Gemini 3.8 / 3.7 / 3.6 Flash with per-model release dates and the high-volume decision bullet updated to match [glm-5.3-flash]
  • model-provider-feature-matrix: Zhipu’s long-context cell moved from ? to 1M flat on the models.dev limit plus Z.ai’s single-tier pricing, the flat-1M club in the reading and choosing sections widened from four vendors to five (Anthropic, DeepSeek, Alibaba, Moonshot K3, Zhipu), the intro now names the models.dev reference, and the models.dev URL joined the references of both articles [glm-5.3-flash]
  • Verification: guide table uniform at 21 cells across all eight lines (20 models), provider matrix 13 rows x 8 cells, 0 broken internal links, front matter and style clean; updated and verified lines stay at 2026-09-08 [glm-5.3-flash]

2026-09-09 (daily refresh, full parallel re-verification + kimi price move)
#

  • Full parallel refresh under the no-stalest-ranking rule: 15 sub-runs across 3 waves re-verified all 143 research notes across the 20 categories, the four essays, and all 21 matrices against live sources (GitHub, npm, PyPI, crates.io, HN Algolia); every “Facts below verified as of” line bumped to 2026-09-09, notes with material revisions got updated set [glm-5.3-flash]
  • Harnesses material: deepseek-harness 216,663 stars with dsh-v0.1.5-alpha.1 (September 8) and its desktop-port repo moved to the dsh-tauri-desk org (301 fixed); claude-code open issues+PRs down to about 12.6k; codex 122.6k stars and about 10.5k commits; cline installs 5,259,433; fx’s home page now advertises local models, gateways, and direct provider keys while the docs still describe only the three-credential model (caution rewritten as announced-but-undocumented); qwen-code’s Qwen OAuth free tier (2,000 requests/day) discontinued 2026-04-15, pricing section rewritten and thesis set to past tense; kilo-code v7.5.16, ante v0.preview.95, pi 103.3k, opencode 206k, openhands 87k, crush 28.0k, zerostack crates 3,399, goose pushed September 9; aider stall re-confirmed [glm-5.3-flash]
  • Surfaces and orchestration material: continue installs 4,103,000 with rating 3.2, kiro 4.3k, openchamber 9.7k with commits landing September 9, roo-code installs 1,992,472, zed about 90k; omnara’s pricing page gained a machine-retention line item ($0.20016 per GiB per 30 days); paseo 16,592 stars with open issues+PRs down to 607 and a v0.8.0-beta.1 prerelease (September 8); worktrunk v0.77.0 (September 8); gastown 17,976, emdash about 5.7k [glm-5.3-flash]
  • Assistant runtimes material: hermes gained paid tiers (Free/Plus/Super/Ultra via Nous Portal, prices behind a bot wall), retiring the no-paid-tier claim; openclaw v2026.9.3 (September 8, Node 24.16+ now required); openwork v0.18.44 with pricing restructured again (Free up to 5 users including the self-hosted control plane, Team $10 per seat up to 100 users, Enterprise $40 per user per month) [glm-5.3-flash]
  • Executions, protocols material: github-agentic-workflows v0.88.7 (September 8); n8n’s license reference moved to docs.n8n.io/n8n-community-license/sustainable-use-license.md after the old path 404ed; agent-host-protocol downloads refreshed (about 209k crates plus 76.9k npm in the explicit window) [glm-5.3-flash]
  • Context, retrieval, memory, hybrid material: cognee’s Standard cloud price cut from $2.50 to $1.00 per 1M tokens (note, matrix cell, and choosing bullet updated); mem0’s Hobby relabel confirmed at 65k stars; claude-mem 93.5k, letta 24.7k, memoryfields 92 stars across repos, instructor 1.17.0 (uploaded today); anthropic-structured-outputs docs now list Mythos 5, 5.1, and Preview; graphify 116,154, rtk 79,575 with about 1,857 combined open, langchain 146k, semble pushed September 8; augment and sourcegraph pricing re-verified unchanged [glm-5.3-flash]
  • Code review, evaluation, sandbox, skills, session material: qodo v0.45.0 (September 5) with install badges drifted, kodus 1,363, open-code-review 22,132, sourcery pushed September 9, deepeval 18,185, phoenix 20.9.0 (September 8), plannotator contributor recount (141 contributors, 929 of 1,168 commits reworded); skills.sh CLI 30.8k with the leaderboard drifted and the installs-into count now 79 (note plus matrix cells); skillregistry.io now returns 402 DEPLOYMENT_DISABLED, its status reworded to went-dark; clawk quiet window now 27 days, openshell forks decreased (1,252) and recorded as such; ctx’s comparison numbers moved from its comparisons page to its home page (caution updated); agentsview pushed September 9 with the ccusage benchmark still gone [glm-5.3-flash]
  • Small categories material: spec-kit v1.0.5 (September 8) at 134,311 stars; fluent v0.4.0 (September 8) with its status rewritten; beads 26,987 with open issues up to 1,098; har references corrected to v1.14.1; machinist 386 stars with its releases line corrected to v0.1.0 through v0.4.0; ouroboros at 26 PyPI releases, paperclip 80,287 with 5,420 open issues; task-master’s stall reworded to “since late April”; tessl pricing re-verified unchanged [glm-5.3-flash]
  • Essays: landscape as-of bumped to 2026-09-09 folding in Cline about 5.3 million installs, the qwen-code free-tier death, and Orca about 65k; model-selection caught the kimi-k2.7-code price drop $0.95/$4.00 to $0.71/$3.50 with cached reads $0.19 to $0.15 via OpenRouter (Moonshot’s platform page still lists the old price, lag noted in prose), the HighSpeed 2x-claim replaced with dollar figures ($1.90/$8.00), and the GLM-5.3-Flash promo verified still live today (ends 24:00 September 9 UTC+8), the four-clocks paragraph rewritten to three clocks once tonight’s deadline resolves and What to Do Next now calendars the 2026-09-10 price check; context-management-patterns and the-tells-are-structural link-checked clean, untouched [glm-5.3-flash]
  • Matrices: all 20 category matrices plus the provider matrix re-verified in the same run and re-dated to 2026-09-09; cell moves: omnara price-model and status, hermes stars/issues/status, openclaw release, picoclaw and zeroclaw issues, fx local-models cell moved from no to partial with the announced-but-undocumented caveat, qwen-code subscription cell moved to the Coding Plan weekly quota, cognee pricing, phoenix and kodus maturity, retrieval LangChain stars, spec-driven adoption row, task/control/factory status and stars rows, provider matrix Moonshot workhorse and windows cells plus the Zhipu cache-prose contradiction fixed (Z.ai does publish cached prices) [glm-5.3-flash]
  • Entrant resolution, all rejected with evidence: Meta Muse (491 points, a closed consumer personal assistant with a client-rendered page and no developer surface, mismatched with the category’s self-hostable runtimes); LongRun-Harness (101 stars two days after creation, zero independent coverage, below the bar exo set at entry with its 169-point launch); the community-skill wave (holo-card-studio 1,189, dream-loop 398, cyber-resume-reviewer-skill 126, claude-style-patch 106) stays out of Skills per the standing scope question; bankmcp (168) is a vertical MCP server, glm-flash-offline-client (106) a vertical chat client, Awesome-OKF (102) an awesome list, not a tool [glm-5.3-flash]
  • _index.md: matrix verified-dates and both essay dates re-dated to 2026-09-09; one-liners refreshed (Cline 5.3M, DeepSeek Harness prerelease-only, Qwen Code free tier, Spec Kit 134k, Paperclip 80k, Hermes 244k plus paid tiers, OpenClaw 389k) [glm-5.3-flash]
  • Style sweep: split two pre-existing two-sentence Bottom-line lines in ellipsis and qodo; no other violations [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus the Kimi price move and promo-clock management filled the day at quality [glm-5.3-flash]
  • Verification: 169 article files plus control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, all 20 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (plus the provider matrix at 13 rows by 8 cells), no em-dashes, no banned terms in article files (ante’s protocol-shape crate name remains the documented 2026-08-30 exception), one sentence per line, and every URL cited by new or edited content fetched 200 or was verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run [glm-5.3-flash]

2026-09-10 (daily refresh, full parallel re-verification + engrim entrant)
#

  • Full parallel refresh under the no-stalest-ranking rule: 16 sub-runs across 4 waves re-verified all 144 research notes across the 20 categories, the four essays, and all 21 matrices against live sources (GitHub, npm, PyPI, crates.io, HN Algolia); every “Facts below verified as of” line bumped to 2026-09-10, notes with material revisions got updated set [glm-5.3-flash]
  • Tag note: chip-huyen, lilian-weng, and the-pragmatic-engineer received their first edits under this model and gained llm=glm-5.3-flash alongside llm=big-pickle (retained); all other touched notes already carried the flash tag [glm-5.3-flash]
  • Harnesses material: ante v0.preview.96 (Sep 9), bullet v1.4.17 (Sep 10), deepseek-harness 217,979 stars with its first rc dsh-v0.1.5-rc.1 (Sep 10, still no stable), juggler v0.6.1 (Sep 9), openhands v1.17.0 (Sep 9), claude-code open issues+PRs flat at 12,583, codex 123k, cline installs 5,272,654 with the unverifiable ratings count dropped, opencode 206.2k, pi 103.6k with open issues+PRs 152 to 190, kimi-code 18.7k combined, zerostack crates 3,419; aider stall re-confirmed [glm-5.3-flash]
  • Surfaces and orchestration material: openchamber v1.23.0 (Sep 9), continue installs 4,111,234, roo-code installs 1,995,310 still ticking post-sunset; cmux pricing page reverted to shared-pool phrasing (24 GB RAM, 6 vCPU shared), conductor 0.85.0 (Sep 9), superset desktop-v1.28.0 (Sep 9), paseo 16,675 with v0.8.0 still beta [glm-5.3-flash]
  • Assistant runtimes material: openclaw patched the June line on 2026-09-10 (v2026.6.35) while no v2026.9.4 exists, openwork v0.18.46 (Sep 10), qwenpaw stars down about 350 to 34.7k with no source for the drop, recorded without editorializing [glm-5.3-flash]
  • Elsewhere material: graphify v0.9.57 (Sep 9) at 116.5k, semble PyPI 98,508/month, qmd pushed Sep 9 after v2.8.3, claude-mem v13.24.5, github-agentic-workflows v0.89.1 (Sep 10), n8n 204k, agent-host-protocol crates 211,871 with the npm window rebased, deepeval 18,199, plannotator v0.27.13 (Sep 10), spec-kit 134,479, beads 27,016 with issues 1,123, paperclip 80,364 with 5,438 issues, agentsview’s “more than 60 agent formats” claim adopted from its homepage, ouroboros v0.54.2 (Sep 9), skills-sh leaderboard reshuffled (tdd displaces frontend-design at #5) with JS-rendered publisher aggregates kept at their 09-09 dates [glm-5.3-flash]
  • Entrant resolution: engrim accepted into Memory, the 2026-09-07 rejection’s calibration met (Show HN 49594008 at 92 points and 64 comments including the dang AI-content flag and the in-repo-docs counterpoint kept as the critical source, 220 stars, v1.3.2 released today, 15 PyPI releases since June); memory-feature-matrix extended from 7 to 8 columns in the same run (Engrim inserted third, case-insensitive) [glm-5.3-flash]
  • Rejections with evidence: okf-agent-memory re-rejected (527 stars flat, no second implementation or independent coverage), Maxxwell, Hazzel, OtoDock (44 points but 111 stars and a non-OSI license, named in the control-planes long tail), opusfived.dev (1,100 points, a one-off gag site), Sebastian Raschka (model-layer band occupied), Mastra Agent Factory, agentflow, ChronoVec, Adios.dev, Proval, SureForge, plus the community-skill wave excluded per the standing scope question [glm-5.3-flash]
  • Essays: model-selection caught the day’s real price move, DeepSeek released deepseek-flash (V4.1-Flash) on 2026-09-10 at $0.30/$1.20 peak ($0.15/$0.60 off-peak, cache hits $0.006, 1M/384K, native vision) and retired V4-Flash, with v4-pro routing to V4.1-Flash rates from 2026-09-14 (new clock added to the clocks paragraph and What to Do Next); the GLM-5.3-Flash promo resolved exactly as calendared (list $0.15/$0.50, cached $0.03); Moonshot’s platform page still lags OpenRouter at $0.95/$4.00; all other provider pages zero moves; landscape as-of bumped with Kimi Code 18.7k the only moved number; context-management-patterns re-dated, the-tells-are-structural link-checked clean and untouched [glm-5.3-flash]
  • URL canonicalizations: aigate’s sandboxing reference to learn.chatgpt.com/docs/sandboxing, openshell docs to /about/overview, claude-code sandboxing docs to code.claude.com, e2b docs to docs.e2b.dev, tessl www to non-www, agentsview.io to www.agentsview.io, session-analytics matrix token-usage path [glm-5.3-flash]
  • Matrices: all 20 category matrices plus the provider matrix re-verified in the same run, membership matched, columns sorted, uniform rows; cell moves: harness OpenHands subscription and IDE rows, memory Engrim column, orchestration Omnara status, assistant-runtimes qwenpaw/hermes/openclaw/picoclaw/zeroclaw, code-review Kodus/OCR/Sourcery, sandboxing OpenShell contributors, spec-driven adoption row, control-planes status and long-tail prose, provider matrix DeepSeek and Zhipu cells [glm-5.3-flash]
  • _index.md: engrim listed in Memory, memory matrix one-liner now eight columns, DeepSeek Harness one-liner moved to first-rc phrasing, twelve dated one-liners re-dated to 2026-09-10 [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus the engrim note and the DeepSeek price move filled the day at quality [glm-5.3-flash]
  • Verification: 169 article files plus control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (provider matrix at 13 rows by 8 cells), no em-dashes, no banned terms in article files (ante’s protocol-shape crate name remains the documented 2026-08-30 exception), one sentence per line, and every URL cited by new or edited content fetched 200 or was verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run [glm-5.3-flash]

2026-09-12 (daily refresh, full parallel re-verification + three entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: 13 sub-runs across 2 waves re-verified all research notes across the 20 categories, the four essays, and all 21 matrices against live sources (GitHub, npm, PyPI, crates.io, HN Algolia); every “Facts below verified as of” line bumped to 2026-09-12, notes with material revisions got updated set [glm-5.3-flash]
  • Harnesses material: amp pricing restructured (new free Hobby tier, Individual $20 with 45,000 orb minutes and a Gigawatt toggle, Teams seat-free with pooled credits and SSO, the old no-free-tier caution rewritten); ante v0.preview.98 with the core-harness-in-private-repo posture folded into its licensing caution; claude-code about 145k stars and issues+PRs down to about 12.5k with two feature-timeline months corrected against the live product page; cline 5,303,954 installs; codex about 124k; fx v0.0.9 with reasoning-steering and subagent delegation plus the vendor-routing caution rewritten (docs now say Codex and Grok subscription requests go direct); goose 54.2k; jcode 19,553; juggler v0.6.3; junie 434 stars (its pricing page 404s, tiers keep their 09-09 date); kilo-code v7.6.2; kimi-code caught kimi-k3 at $2.65/$13.28 on OpenRouter; onecli 3,469; opencode about 207k; openhands v1.18.0 with the Local GUI deprecated into Legacy docs; pi 104,384; warp about 65k with the any-harness cloud beta noted; zerostack crates drift [glm-5.3-flash]
  • Surfaces material: continue installs 4,129,674; kiro CLI 2.21.3 patches and Crew 0.6.0 session-harness selection; openchamber v1.23.1; roo-code 2,001,981 installs with ZooCode named alongside Cline as the community fork; vscode-copilot 192k; antigravity, cursor, jetbrains, trae, void, windsurf, zed bump-only [glm-5.3-flash]
  • Orchestration material: paseo shipped v0.8.0 stable (September 10) with plugin providers and Hub follow-ups; happy-coder 23,762; superset’s README now lists 21 agent presets (up from the 12 cited); vibe-kanban orphan state re-confirmed via the commits feed; worktrunk about 7.2k [glm-5.3-flash]
  • Assistant runtimes material: hermes’ paid tiers went public on the official site (Free, Plus $20, Super $100, Ultra $200), retiring the bot-wall claim; openclaw v2026.9.4 (September 11); openwork’s pricing page dropped the $40/user and 100-user caps for custom Enterprise; qwenpaw v2.2.1; zeroclaw’s Android port was archived read-only by its owner in March (“too unsafe”), so the port is no longer an active strength [glm-5.3-flash]
  • Code review material: ellipsis Advanced support $7,500 to $10,000; open-code-review v1.12.0 (September 12) at 22,687; kodus, sourcery, qodo drift refreshed [glm-5.3-flash]
  • Evaluation and sandboxing material: frontierharness quality spread corrected from about 13 to about 17 points with the publisher’s own attribution folded in and Claude Code now second on quality; phoenix 20.11.0; plannotator v0.27.14 with Workspaces on public waitlist; agent-sandbox v1.0.2; clawk quiet window 30 days; flue and openshell drift [glm-5.3-flash]
  • Context, retrieval, memory, executions, hybrid material: graphify published four plans (Free, Pro $10, Teams $20/seat, Enterprise), its matrix cells moved from early-access phrasing; rtk v0.49.0; semble v0.6.0 at 92,479 downloads; repomix window re-based; claude-mem v13.24.23; engrim v1.4.2 with native Codex hook integration; github-agentic-workflows’ release line corrected (v0.89.x all prereleases, v0.88.7 the last stable); openai-structured-outputs examples now lead with gpt-6-astra [glm-5.3-flash]
  • Small categories material: spec-kit v1.0.6 at about 136k stars; har’s site now advertises HAR HQ, a hosted team layer with no published pricing, replacing the no-cloud claim; ouroboros 14 runtimes and v0.54.3; beads 27,104; openspec npm 1,768,196 for the month; ctx added a paid referral program; tessl launched Tessl Code Review (tiers unchanged); tinyagi and task-master stalls re-confirmed [glm-5.3-flash]
  • People material: addy-osmani’s bio states he departed Google in 2026 after 14 years, the note and the people-matrix enterprise-vantage cell reframed; latent-space passed 200,000 subscribers; hamel-husain’s course cohort 5,000-plus; deeplearning-ai Batch at 370; ai-jason’s cadence corrected to two-three uploads a month; caleb-writes-code’s sponsor recount 12 of 15; simon-willison posting through September 12 [glm-5.3-flash]
  • Essays: model-selection caught kimi-k3’s OpenRouter drop to $2.65/$13.28 (Moonshot’s platform page still lists $3/$15, lag recorded) and the DeepSeek V4 Pro Flash-routing cancellation (the row is permanent again, the September 14 clock removed); every other provider page zero moves and the GLM-5.3-Flash promo resolution stands as calendared; landscape as-of bumped with OpenHands about 88k the only moved number; context-management-patterns and the-tells-are-structural link-checked clean, untouched [glm-5.3-flash]
  • Entrant resolution, three accepted, each written to citation standards with all sources fetched this run: grok-build (SpaceXAI’s Apache-2.0 Rust TUI agent, 26,704 stars in nine weeks, 590-point open-source thread after the 539-point wire-level analysis caught whole-repo uploads, no GitHub releases, v1.0.25 via the x.ai installer) into Harnesses; Graft (NanoNets/Trail’s MIT context-graph CLI, 7,276 stars in ten weeks against a 3-point Show HN, every benchmark vendor-run, Trail Brain from $20k/year as the upsell) into Context engines; JetBrains Air (the standalone ADE running Codex, Claude Agent, Gemini CLI, and Junie in parallel with worktree/Docker/cloud isolation, bundled with AI Pro/Ultimate, thin HN footprint stated) into Orchestration, placed there rather than Surfaces per the Conductor/Superset precedent [glm-5.3-flash]
  • Entrant rejections with evidence: SNARC (7 stars, quiet since August 15, the wave’s “clearest entrant” claim did not survive verification); litelm (169-point HN but 187 stars and a LiteLLM-style gateway, the standing gateways ? for tom: still blocks a category); graphify-csharp folded into the graphify note as ecosystem evidence (41-point Show HN for the third-party C# port, the largest Graphify-linked thread); Proliferate re-checked and still below bar; the OpenAI-agents RubyGems incident (908 points) logged as incident news with no tool to profile, the sandboxing category already carrying that lesson [glm-5.3-flash]
  • Matrices: harness 25 to 26 columns (Grok Build), context engines 7 to 8 (Graft), orchestration 13 to 14 (JetBrains Air, re-sorted one position when the coordinator’s insertion instruction proved mis-sorted against Happy Coder); all 20 category matrices plus the provider matrix re-dated to 2026-09-12 with cells updated only where facts moved; provider matrix DeepSeek cells updated for the routing cancellation and Moonshot’s flagship cell carries the OpenRouter price with the platform lag [glm-5.3-flash]
  • Correction this run: happy-coder’s “out-stars every tool in the category” claim was factually wrong (cmux at 27k), fixed in the note, the matrix prose, and the index one-liner [glm-5.3-flash]
  • _index.md: three new notes listed alphabetically, matrix one-liners re-dated or re-counted, five one-liners refreshed (DeepSeek 221.5k, Hermes tiers, Spec Kit 136k, Osmani departure, Happy Coder recast) [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus three entrant notes with their matrix columns filled the day at quality [glm-5.3-flash]
  • Verification: 172 article files plus control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 20 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (provider matrix 13 rows by 8 cells), no em-dashes, no banned terms in article files (ante’s protocol-shape crate name remains the documented 2026-08-30 exception), one sentence per line, and every URL cited by new or edited content was fetched 200 or verified via the Algolia, GitHub, npm, PyPI, or crates.io APIs this run (x.ai 403s bots, its changelog fetched via a text-extraction proxy; junie.jetbrains.com/pricing 404 with tiers keeping their 09-09 date; ampcode.com pricing JS-hydrated, claimed only from served content) [glm-5.3-flash]

2026-09-13 (owner-directed: Changes sections)
#

  • Owner request in chat: every article carries a ## Changes section before ## See also documenting how it changed over time, backfilled from this log; codified in AGENTS.md (article format list, note skeleton, writing rules, daily refresh step 5) [owner, glm-5.3-flash]
  • AGENTS.md rule: ## Changes is an append-only, dated changelog of material revisions, creation bullet first, sourced from this log; verification-date bumps and routine volatile-number refreshes are not changes; the bullet is appended in the same run as the change it records [owner, glm-5.3-flash]
  • Backfill: all 172 articles gained the section, 397 bullets total, the creation bullet from each article’s front-matter created date plus curated material changes from this log (corrections, status and pricing moves, category moves, matrix membership changes); bump-only refresh runs and passing mentions left out per the new rule [glm-5.3-flash]
  • Every touched article confirmed to carry llm=glm-5.3-flash (appended to the seven missing it); updated fields left untouched since no fact moved [glm-5.3-flash]
  • Verification: all 172 articles carry ## Changes immediately before ## See also, bullets dated ascending with the first equal to created, no em-dashes or banned terms in the new sections, no links inside bullets, mandatory tags present, front matter parses as YAML, and no path outside agents/ was touched [glm-5.3-flash]

2026-09-13 (daily refresh, full parallel re-verification, maintenance-leaning)
#

  • Full parallel refresh under the no-stalest-ranking rule: 12 note sub-runs across two waves re-verified all 147 research notes across the 20 categories against live sources (GitHub, npm, PyPI, crates.io, HN Algolia), then 6 matrix sub-runs re-verified all 20 category matrices plus the provider matrix; every “Facts below verified as of” line bumped to 2026-09-13 [glm-5.3-flash]
  • Harnesses material: amp’s Sep 13 “Free Agent” announcement recorded (free with your own compute, subscriptions, or keys, BYOK token fees dropped outside Enterprise, nine more BYOK providers in early access) in the note, the harness matrix BYOK cell, the landscape tracker, and the index one-liner; cline installs 5,307,119; deepseek-harness 221,878 with dsh-v0.1.5-rc.2 still latest and no stable; claude-code 144,880 stars with issues+PRs 12,534; grok-build changelog v1.0.30 (Sep 11) still release-less on GitHub; juggler stars decreased 593 to 592, recorded without editorializing; exo’s FrontierHarness claims re-verified on the live page [glm-5.3-flash]
  • Surfaces material: windsurf’s docs domain now 404s with Devin Desktop’s docs the only survivor (note Changes bullet, surface-matrix reference swapped to docs.devin.ai, index one-liner recast); kiro CLI 2.21.4; continue installs 4,130,133 with no revival signal; roo-code install counter static post-sunset; cursor’s November 12 OpenAI shutoff verified unrevised [glm-5.3-flash]
  • Orchestration defect found and fixed: the Happy Coder and JetBrains Air body cells were transposed across all ten rows of the orchestration matrix since the Air column landed 09-12 (headers right, data columns swapped), restored against the notes with a Changes bullet; gastown’s status reframed to “active but cooling” (main branch quiet since 2026-07-23, no release since v1.2.1); emdash gained scheduled automations; omnara cli-v1.0.11 (Sep 12); vibe-kanban orphan wording corrected (repo pushed_at moved without any committed change) [glm-5.3-flash]
  • Code review and evaluation material: ellipsis added a free-for-individuals tier (own ChatGPT/Claude subscription) with the Advanced $10,000 package re-verified live; qodo launched the Agentic Toolbox (qodo-review/rules skills inside Claude Code, Codex, Kiro), both recorded in notes and matrix cells; open-code-review 23,041; plannotator gained a terminal sibling, Herdr Annotate (452 stars), noted in its note; the evaluation matrix intro’s “same day” claim corrected to the recorded 2026-08-30 date [glm-5.3-flash]
  • Memory, session analytics, small categories: engrim v1.4.3 (Sep 13, MCP outputSchema); agentsview’s docs now carry usage JSON schema version 6 (caution corrected, session-matrix prose updated); har’s HAR HQ published its first pricing (Team $400/mo up to 50 users with $100 monthly cloud credits, Enterprise custom), a new Pricing row added to the software-factory matrix; ouroboros v0.54.4 and its note’s one lingering “13 runtimes” instance aligned to the recorded 14; beads prerelease v1.3.0-rc.2; task-master stall re-confirmed [glm-5.3-flash]
  • Skills, executions, hybrid material: skills.sh leaderboard reshuffled (grill-with-docs 962.1K, new #8 setup-matt-pocock-skills) and skillregistry.io is live again with 61 skills after the 09-09 dark episode, recorded in the note; github-agentic-workflows v0.89.10 prerelease with v0.88.7 confirmed last stable via the releases API; skillopt 17.0k with the issues count corrected to 40 issues and PRs; instructor downloads refreshed to 22.3M/month [glm-5.3-flash]
  • People material: steve-yegge’s writing venue corrected to yegge.ai (165 essays, about 699k words, RSS) after his Substack proved an empty shell with zero posts (note plus people-matrix platform cell); latent-space 200k+ subscribers re-confirmed with the Sep 12 FDE issue; the-pragmatic-engineer 1.1M+ subscribers added with its September agent-adoption coverage; lilian-weng’s 2026 cadence documented (two posts, latest June 24); the latent-space note’s broken “Interconnects” link fixed to point at nathan-lambert [glm-5.3-flash]
  • Protocols: acp’s agent list sharpened to exactly 40 entries with its restructured docs reference updated; agent-host-protocol 219,376 lifetime crate downloads with the npm window rebased; mcp spec unchanged at 2026-07-28; the remaining /overview/clients reference in the acp note re-fetched and confirmed still live [glm-5.3-flash]
  • Essays: model-selection re-checked every provider price against live pages (OpenRouter, Moonshot, DeepSeek, Z.ai, OpenAI, Anthropic, Google, Alibaba) with zero moves, bump-only, the Moonshot platform lag and the calendared November 21 / December 31 clocks standing; landscape as-of bumped with the Amp line updated and the Orca count moved to about 68k (67,641); context-management-patterns and the-tells-are-structural link-checked clean, untouched [glm-5.3-flash]
  • Entrant resolution, all rejected with evidence across 12 scans of the 2026-09-10..13 Show HN window: Hazzel (9 points, 4 stars), gpty (93 points but a Godot PTY multiplexer, skepticism-heavy thread), Maestro (2 stars, two days old, no license), Rocky Surf (3 points, 5 stars), Charter (11 stars, team platform), no_human (305 stars but a 2-point thread and no independent coverage), hcode (no published source, self-described non-OSS MVP), Weftgate (1 point, 2 stars), Beta9 (2023 repo re-shared, execution infra not isolation), learnlance (2 points), ClaudeStatsBar (a single-harness widget), RagLeap (third dead self-posted Show HN), Benzi (stale self-run benchmark), Personal Context MCP (no public repo), plus out-of-category candidates (Copperhead EDA, Toast IDE, Graphify C#, Stroq) named in the group reports; standing exclusions held (community-skill wave, litelm, okf-agent-memory, SNARC, Proliferate) [glm-5.3-flash]
  • Matrices: all 20 category matrices plus the provider matrix re-verified in the same run, membership matched, columns sorted, re-dated to 2026-09-13; cell moves: harness Amp BYOK, surface Windsurf reference, orchestration Happy Coder/JetBrains Air un-transpose plus Gas Town and Vibe Kanban status, assistant ZeroClaw Android and issue counts, code-review Ellipsis and Qodo, evaluation intro correction, software-factory new Pricing row (HAR), people Yegge platform and AI Jason cadence [glm-5.3-flash]
  • _index.md: all dated one-liners re-dated to 2026-09-13; one-liners refreshed (Amp free on own compute, Windsurf redirect-and-docs, Ouroboros 14 runtimes) and the stale Osmani trailing clause dropped from the people-matrix line [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus the amp restructure and the orchestration matrix defect filled the day at quality [glm-5.3-flash]
  • Verification: 172 article files plus control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, all 20 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (provider matrix 13 rows by 8 cells), no em-dashes, no banned terms in article files (ante’s protocol-shape crate name remains the documented 2026-08-30 exception), one sentence per line, and every URL cited by new or edited content was fetched 200 this run (junie.jetbrains.com/pricing still 404 with tiers keeping their 09-09 date; devin.ai/pricing 429’d and no pricing claim was made from it; the MSR Skillopt blog 403s bots, fetched with a browser user agent) [glm-5.3-flash]

2026-09-13 (owner-prompted, agent-native note)
#

  • Owner request in chat: place BuilderIO/skills and www.agent-native.com; owner chose in chat the Skills-category route, one note covering the pack with the framework as context, over seeding a new agent-native-apps category [glm-5.3-flash]
  • Research index: added the agent-native note to Skills, six sources all fetched this run (both GitHub repos, the landing page, the npm registry and downloads APIs, and the HN Algolia zero-hit search); recorded the missing-community signal (no HN threads or independent coverage found as of 2026-09-13) and the hosted-endpoint, managed-instruction-block, and Clips-capture cautions [glm-5.3-flash]
  • skills-feature-matrix: extended from five to six columns (Agent-Native inserted first alphabetically) in the same run per the matrix rule, 12 cells traced to sources fetched this run; Kind prose extended to the pack, the stale “new column” phrasing on SkillOpt fixed, reading gains the pack paragraph, choosing gains the starter-set bullet, references extended to 10 [glm-5.3-flash]
  • _index.md: agent-native added to Skills alphabetically, skills matrix one-liner extended with the curated pack [glm-5.3-flash]
  • Verification: all internal link targets exist on disk (agents/agent-native/index.md new; skills-sh, agent-skills-open-standard, anthropic-agent-skills, mcp, agentic-coding-tools-landscape verified), all cited external URLs fetched 200 this run, front matter parses with no type field, mandatory tags present, no em-dashes, no banned terms, one sentence per line [glm-5.3-flash]

2026-09-13 (owner-directed: category directory reorganization)
#

  • Owner request in chat: reorganize agents/ by category; moved 168 article directories (148 research notes plus the 20 category matrices) into 20 category directories named for their index sections: harnesses/, surfaces/, orchestration/, protocols/, context-engines/, skills/, retrieval/, memory/, executions/, hybrid-execution/, code-review/, evaluation-review/, sandboxing/, spec-driven-development/, task-management/, control-planes/, assistant-runtimes/, software-factory/, session-analytics/, people-and-publications/ [glm-5.3-flash]
  • Left at the section root: the four essays (agentic-coding-tools-landscape, model-selection-for-coding-tasks, context-management-patterns, the-tells-are-structural), model-provider-feature-matrix (no category; it accompanies model-selection-for-coding-tasks), and the five control files; the agent-native note moved with its Skills peers [glm-5.3-flash]
  • Links: rewrote the internal links of every article, matrix, and control file (1,272 links changed in a first pass, 1,237 corrected in a second pass after the link check caught an off-by-one-level rewrite); final shape: same-category siblings keep ..//, cross-category links are ..///, essays and control files are ../../, corpus articles are ../../..//, and the section index is ../../_index.md from inside a category [glm-5.3-flash]
  • One inbound corpus link updated out of necessity: speeding-up-llm-work-on-a-single-codebase now points at agents/session-analytics/agentsview/index.md; root AGENTS.md’s agents/AGENTS.md link was unaffected; no other path outside agents/ was touched [glm-5.3-flash]
  • _index.md category lists now link through the category directories; article content and front matter are byte-identical apart from link paths, so no per-article Changes bullets; in the append-only files (log, queue) only link targets were retargeted mechanically, no entry text altered [glm-5.3-flash]
  • Verification: php test-links.php reports zero broken links involving agents/ (the 24 remaining broken targets are pre-existing corpus sections outside this section’s scope and were left alone); the section verifier passes 173 article front matters with mandatory tags, section order, and no type fields; every matrix still matches its category’s membership [glm-5.3-flash]
  • AGENTS.md is agent-untouchable and still documents the flat layout’s link conventions, so flagged via ? for tom: in queue.md [glm-5.3-flash]

2026-09-13 (owner-prompted, automated research category)
#

  • Owner request in chat: categorize autoresearch-style projects and cover what Anthropic and OpenAI are doing toward the Millennium problems; resolved in chat to a new Automated research category with six notes plus the same-run matrix [glm-5.3-flash]
  • Research index: added agents/automated-research/ with six notes (alphaproof, anthropic-claude-math, harmonic-aristotle, math-inc-gauss, openai-deep-research, openai-for-science), written under the in-flight per-category directory layout; no pre-existing path touched [glm-5.3-flash]
  • Facts recorded: Anthropic zeta zero bound 41.6 to 67.2 percent (Aug 10 2026) with the 60-subagent Claude Code loop, FLT formalized in 11 days at 13M lines (Sep 4 2026), Fable 5 credited with the Jacobian conjecture disproof; Math Inc. Strong PNT (25k lines, three weeks) and sphere packing (about 200k lines, Feb 2026), OpenGauss 8/23 on FormalQualBench with the caught Codex axiom-injection exploit; Harmonic Aristotle’s Erdős #124 solve (Lean-checked, a version of the problem) and its comparator-unverified 6/23 column; DeepMind IMO 2024 silver and IMO 2025 official gold 35/42 with the IMO validation disclaimer; OpenAI Deep Research lineage (o3 to GPT-5.2, MCP connectors, quotas); OpenAI for Science’s FrontierMath open-problem solve (GPT-5.4, March 2026), the deleted-posts episode (October 2025), and the disputed Navier-Stokes claim (Bubeck and Altman versus Buckmaster and Alpöge, September 2026) [glm-5.3-flash]
  • automated-research-feature-matrix: created in the same run per the matrix rule, six columns sorted alphabetically, eight rows, every cell traced to a member note’s fetched sources; the who-judges row named the deciding one in prose [glm-5.3-flash]
  • _index.md: added the Automated research category section (six notes, alphabetical) and the matrix line after Model Provider [glm-5.3-flash]
  • Sources: every URL cited by the new content was fetched 200 this run except openai.com and help.openai.com (bot-blocked 403), so no claim cites them; pcmag also 403’d and was not cited [glm-5.3-flash]
  • Note: a separate restructure of the section into per-category directories was in flight in the working tree during this run; new files follow the restructured layout [glm-5.3-flash]
  • No self-directed essay this run: the owner-directed category build filled the run at quality [glm-5.3-flash]

2026-09-13 (owner-directed: AGENTS.md conventions for category directories)
#

  • Owner request in chat: update agents/AGENTS.md to describe the new category-directory layout; the “never modify this file” guardrail binds agents acting on their own, so this edit was made only on the owner’s explicit instruction [owner, glm-5.3-flash]
  • AGENTS.md changes: the Article format now places articles at agents///index.md, names the 20 category directories, notes each category’s feature matrix lives alongside its notes, and keeps essays, model-provider-feature-matrix, and control files at the section root; the Writing rules internal-links bullet now documents same-category ..//, cross-category ..///, section-root ../../, control files ../../, and corpus ../../../ depths; the note skeleton’s See also example now points at ../../agentic-coding-tools-landscape/index.md [owner, glm-5.3-flash]
  • No article content changed; the reorganization’s link check still stands (zero broken links involving agents/) [glm-5.3-flash]
  • Correction, same run: the skeleton’s See also example link, read literally by CI from agents/AGENTS.md’s own directory, resolved to a broken path, so it became a non-link placeholder pointing at the Writing rules depths (consistent with the References placeholder style) [glm-5.3-flash]
  • Second correction, same run: the checker’s regex matches any bracket-parenthesis sequence, so the placeholder was rewritten without link syntax at all, angle-bracket text only like the References example [glm-5.3-flash]

2026-09-16 (daily refresh, full parallel re-verification + nine entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: 16 sub-runs across two waves re-verified all 154 research notes across the 21 categories, the four essays, and all 22 matrices against live sources (GitHub via gh api after anonymous rate limits, npm, PyPI, crates.io, HN Algolia, vendor and pricing pages); every “Facts below verified as of” line bumped to 2026-09-16, notes with material revisions got updated set [glm-5.3-flash]
  • Harnesses material: deepseek-harness 225,829 stars with dsh-v0.1.6-alpha.1 (Sep 15, still prerelease-only); fx v0.0.10 (Sep 14) and the docs added a custom-model-connections preview for Ollama, vLLM, and OpenRouter (harness-matrix local-models cell updated with the preview-build caveat); cline installs 5,341,071 with the vendor claim now 11M-plus; kilo-code v7.7.2, bullet v1.4.20, ante v0.preview.99 (eval best 83.9 percent on DeepSeek V4.1 Flash), juggler v0.6.4; jcode’s release pause recorded (v0.84.0 still latest); grok-build v1.0.34 via npm with the dist-tag lag noted; kimi-code caught kimi-k2.7-code’s OpenRouter output cut to $3.21; codex about 124.6k, opencode 207.7k, openhands 88.1k, pi 106k, claude-code 145,227 stars with issues down to 12,433; aider stall re-confirmed [glm-5.3-flash]
  • Surfaces material: kiro IDE 1.1 shipped Sep 14 with the GPT-5.6 1M-context raise; openchamber v1.23.2 (Sep 14); windsurf’s docs domain returned as a 308 chain into docs.devin.ai (the 404 note corrected); continue and roo-code counters tick on dead repos; antigravity gained the ToS account-suspension caution from the 337-point Sep 3 thread, resolving the flag the re-verification run surfaced in-run [glm-5.3-flash]
  • Orchestration material: happy-coder cli-1.2.4 with 1.2.5 betas; superset v1.29.0; jetbrains-air 262.834.41 (git-repository requirement removed); vibe-kanban recorded its first post-shutdown default-branch commit (Sep 15, a GPG-verified former Bloop maintainer, still no release); worktrunk jumped to about 7.9k stars; paseo tracker-cleanup caution extended; cmux’s pricing page stable this cycle [glm-5.3-flash]
  • Assistant runtimes, memory material: openwork’s pricing churned again (Solo free tier, Team Starter $10/seat, Enterprise custom); nanobot shipped v0.3.5 (Sep 15) after seven quiet weeks, retiring the stale release-cadence claim; hermes 246k with tiers live; engrim shipped v1.4.5/v1.4.6 and added OpenCode as a sixth CLI; claude-mem’s npm channel (13.25.x) runs ahead of its GitHub releases, drift recorded; memoryfields’ spec gained optional page status fields, a partial contradiction-handling answer (matrix cell moved); cognee’s $1.00/1M cut holds [glm-5.3-flash]
  • Context, review, evaluation, executions material: graft +800 stars in three days (8,121, fastest in category); graphify v0.9.62 with monthly billing added to its plans (matrix cell moved); rtk’s v0.50.0 RC train explicit; open-code-review jumped 23,041 to 29,564 stars in three days with v1.12.x releases; graphite-diamond’s reference comment count fixed (253); qodo’s unlimited-hero-versus-credit-FAQ tension recorded; deepeval 4.2.3, phoenix 20.12.0, plannotator v0.27.15, agentsview v0.43.0; openshell gained its formal-methods post thread as a reference; github-agentic-workflows gained the missed security advisory (GHSA-8h78-hpm7-29gg, releases 0.83.3-0.85.4 retired, prerelease train at v0.89.15); instructor downloads 21.2M/month [glm-5.3-flash]
  • Small categories material: spec-kit v1.0.7 at 137.1k stars; beads shipped v1.3.0 stable Sep 15; har v1.14.2 after a cadence-breaking week; machinist 418 stars with zero-coverage re-dated; paperclip 80,789 with 5,481 issues; people: latent-space two new issues, Interconnects over 83,000, simon-willison posting through Sep 15, the-pragmatic-engineer’s OpenAI agentic-factory deep dive added [glm-5.3-flash]
  • Entrant resolution, nine accepted, each written to citation standards with all sources fetched this run: opensandbox (15,337 stars, Apache-2.0, E2B-class platform whose traction is GitHub/Trendshift-driven rather than HN) and cubesandbox (Tencent, 12,537 stars, RustVMM/KVM microVMs, multi-submission HN pattern recorded) into Sandboxing; agents-observe (77-point launch thread found that earlier scans missed, fills the matrix’s named live-observation gap) into Session analytics; langfuse (34.7k stars, MIT core, ClickHouse acquisition, ee/ split as the caution) into Evaluation and review; ordewell (50-point launch, typed plan artifacts, the AI-written-replies episode recorded as the transparency caution) into Task management; gsd (the archived 64.5k-star get-shit-done redirecting to open-gsd/gsd-core, crypto-scam claim kept attributed not asserted) into Spec-driven development; pion (Andon Labs’ 483-point closed research preview where agents run a real business, placed in Automated research as a capability-research instrument, assistant runtimes considered and rejected since nothing self-hostable exists) into Automated research; ag-ui (CopilotKit’s agent-to-frontend protocol, 15.9k stars, 7.17M monthly npm downloads) into Protocols; chonkie (4,751 stars, 1.06M monthly PyPI downloads, Launch HN 151 points, commercial arm dead with the maker pivoted to Feyn Labs) into Retrieval [glm-5.3-flash]
  • Matrices extended in the same run per the matrix rule: sandboxing 6 to 8, session-analytics 2 to 3, evaluation 5 to 6, task-management 3 to 4, spec-driven 4 to 5, automated-research 6 to 7, protocols 5 to 6, retrieval 4 to 5; all 22 re-dated to 2026-09-16 with cells updated only where facts moved (memory, context-engines, code-review, orchestration, assistant-runtimes, provider Moonshot cells, surface Windsurf reference correction) [glm-5.3-flash]
  • Essays: model-selection caught the day’s price move, kimi-k2.7-code cut to $0.71/$3.21 on OpenRouter with cache reads up to $0.18 while Moonshot’s platform page now lags on both Kimi models; every other tabled price re-verified unchanged; provider matrix Moonshot workhorse and cached-input cells updated; landscape as-of bumped with Orca about 70k folded in; context-management-patterns and the-tells-are-structural link-checked clean and re-dated [glm-5.3-flash]
  • Entrant rejections with evidence: panel (53 stars day-old), pizza bot (four citable sources), agent-launcher (zero pushes since creation), ai-data-extractor (star-farm signature), birdview, life-recorder, baize, ivyclaw, aegisops, gpty (Godot multiplexer, wrong domain), thurbox (4 points); standing exclusions held (community-skill wave, gateways trio, okf-agent-memory, oh-my-pi, hermes, cursor CLI and routines covered inside their sibling notes per the vendor-family precedent; xgrammar and vLLM structured outputs stay engine internals cited inside the outlines note, below the category’s application-layer scope) [glm-5.3-flash]
  • _index.md: nine new notes listed alphabetically, six matrix one-liners re-counted and all dated lines re-dated to 2026-09-16, one-liners refreshed (DeepSeek 225k and v0.1.6-alpha.1, claude-mem 94k, Hermes 246k, session-analytics gap filled) [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus nine entrant notes with their matrix columns filled the day at quality [glm-5.3-flash]
  • Verification: 192 files pass front-matter parsing with no type field and mandatory tags, 1,630 internal link targets all resolve on disk, all 21 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (provider matrix at 13 rows by 8 cells), no em-dashes in article files, no banned terms beyond the documented ante crate-name exception and the log’s own historical mentions, one sentence per line in prose, and every URL cited by new or edited content was fetched 200 this run or its absence recorded honestly (pypistats 429, devin.ai/pricing 429, x.ai changelog 403, junie pricing 404, science.org 403, picoclaw.io expired certificate) [glm-5.3-flash]

2026-09-18 (daily refresh, full parallel re-verification + harnesstax entrant)
#

  • Full parallel refresh under the no-stalest-ranking rule: 21 sub-runs across 3 waves re-verified all 163 research notes across the 21 categories, the four essays, and all 22 matrices against live sources (GitHub via gh api, npm/PyPI/crates.io, HN Algolia, vendor and pricing pages); every “Facts below verified as of” line bumped to 2026-09-18, notes with material revisions got updated set [glm-5.3-flash]
  • Entrant accepted: HarnessTax written into Evaluation and review (the UC Berkeley harness-swap study, 217-point HN 2026-09-16, seven models x three harnesses on SWE-bench Lite and Terminal-Bench 2.0, harness moves cost up to 5x while success stays within a few points, six sources fetched this run); evaluation-review-feature-matrix extended from six to seven columns in the same run per the matrix rule [glm-5.3-flash]
  • Harnesses material: ante shipped v0.2.0 (2026-09-17) closing the v0.preview.N series; kimi-code caught kimi-k3’s OpenRouter cut to $2.10/$10.95 with cache hits at $0.23 (Moonshot’s platform page still lists $3/$15, lag recorded); warp-agent-cli published Factories early-access pricing (PAYG at a 20 percent markup, credits in Build/Max/Business tiers); deepseek-harness 228,399 stars with dsh-v0.1.6-alpha.2 still prerelease-only; kilo-code v7.7.4, openhands v1.20.0 (2026-09-17), gemini-cli v0.60.0, jcode 19.9k stars, grok-build alpha v1.0.36; aider stall re-confirmed [glm-5.3-flash]
  • Surfaces material: trae repriced its whole ladder upward between the 16th and today (Lite $3 tier dropped, Pro $10 to $20, Pro+ $30 to $60, Ultra $100 to $200; the note’s price-advantage thesis rewritten); openchamber v1.24.0/1.24.1 crossed 10k stars with a v2-preview tag on OpenCode v2; kiro CLI 2.22.0 with Claude Fable 5.1 Enterprise Preview; windsurf’s devin.ai pricing route moved 429 to 404, prices still unverified; dead records (continue, roo-code, void) intact with install counters still ticking [glm-5.3-flash]
  • Orchestration material: cmux pricing moved to monthly-only (Pro $50, new Max $200 tier) for the third change in nine days, release v0.64.25; vibe-kanban’s community commits resumed in earnest (seven substantive default-branch commits 2026-09-16 to 18 after the recorded single 09-15 bump, still no release); worktrunk v0.78.0 with a breaking hook-context-key rename; gastown’s active-but-cooling status verified still holding (57 days quiet on the default branch) [glm-5.3-flash]
  • Assistant runtimes material: Anthropic announced Claude Cowork merging into Claude chat (2026-09-16, 230-point thread), recorded in the eigent and openwork Cowork-alternative notes; openclaw shipped a July-line patch v2026.7.33 dated today while the 9.x line is current (pattern recorded); hermes v2026.9.14 at 247k [glm-5.3-flash]
  • People material: addy-osmani’s site now states he is Member of Technical Staff at Anthropic working on Claude Code (note and people-matrix vantage cell updated); steve-yegge’s Gas Town shut down, Yegge admitting he never successfully built anything with it, per Dan Luu (2026-09-15) via AINews (2026-09-17), both the note and matrix cell updated; caleb-writes-code 112k subs and 120 videos; simon-willison posting through September 17; latent-space new AIUC interview issue [glm-5.3-flash]
  • Elsewhere material: open-code-review jumped to 35,673 stars (+6,109 in two days, still no HN catalyst) with v1.12.5; qodo launched the Software Map beta; ellipsis’s free-individual tier reworded to Claude Code or Codex subscriptions; langfuse v4.38.0, phoenix 20.14.0 (today), plannotator v0.27.16; graft 8,458 stars; graphify 119,153 stars with a new AMBIGUOUS edge tag recorded; rtk’s v0.50.0 RC train at rc.442; engrim v1.4.7; skills-sh leaderboard refresh with azure-validate now Pass on all three scanners; ag-ui shipped 1.0 (2026-09-17) with AWS Bedrock AgentCore first-party; paperclip v2026.916.0 at 80,962 stars; ctx v1.4.10 with about 40 documented harnesses (session-matrix cell moved); instructor downloads 20.3M/month [glm-5.3-flash]
  • Essays: model-selection moved with the kimi-k3 cut (table, prose, decision bullet, Changes bullet); landscape as-of bumped and the Gas Town line rewritten for the shutdown with inline AINews and Dan Luu links; context-management-patterns and the-tells-are-structural link-checked clean, untouched; provider matrix Moonshot flagship cell moved to $2.10/$10.95 with the platform lag noted [glm-5.3-flash]
  • Entrant rejections with evidence, all resolved in-run: Apprentice (6 points, Common Lisp harness), Agentbox (431 stars but a 3-point Show HN and no independent coverage), Skillsync (55 points, session portability, rejected by four category runs as out of scope), skillbay.sh (23 points, under five verifiable sources), skillcrossroads (scope plus 2 points), thruwire/foreman (237 stars but one day old, the coleam00 bar), MCPJam evals (10 points, protocol tooling), agentmemoryleaderboard (2 points), plus the weak tails; standing exclusions held (community-skill wave, gateways trio, okf-agent-memory, oh-my-pi, vendor-family folds) [glm-5.3-flash]
  • Reconciliation this run: two sub-runs fetched kimi-k3’s OpenRouter price and reported different figures ($2.00/$11.20 versus $2.10/$10.95); the coordinator re-fetched the live OpenRouter models API, confirmed $2.10/$10.95 with cache hits at $0.23, and corrected the kimi-code note and its Changes bullet [glm-5.3-flash]
  • _index.md: HarnessTax listed in Evaluation and review, the evaluation matrix one-liner now seven columns, eleven one-liners refreshed (DeepSeek 228k/alpha.2, Træ repricing, cmux tiers, Gas Town shutdown, Vibe Kanban commits, Spec Kit 137.7k, Graft 8.4k, Osmani at Anthropic, Yegge shutdown, AG-UI 1.0), all dated lines re-dated to 2026-09-18 [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus the HarnessTax note and the Træ repricing filled the day at quality [glm-5.3-flash]
  • Verification: 190 article files plus 5 control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, all 21 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (provider matrix at 13 rows by 8 cells), no em-dashes, no banned terms beyond the documented exceptions (ante’s protocol-shape crate name, the log and rules files’ own historical mentions, emdash’s append-only Changes history), one sentence per line, and every URL cited by new or edited content was fetched 200 this run or its absence recorded honestly (x.ai changelog 403, devin.ai/pricing 404, pypistats 429 for two notes whose dated claims stand, venturebeat 429) [glm-5.3-flash]

2026-09-18 (owner-prompted, jev note)
#

  • Owner question in chat (“why aren’t we covering jev yet?”); the answer is a process defect now fixed: Jev launched 2026-09-15, and both the 09-16 and 09-18 entrant scans started their windows at 09-16, so the 1,915-point launch thread never entered a scan; the 09-18 hybrid-execution sub-run did surface a Jev commentary repost and rejected it as commentary without tracing it to the product [glm-5.3-flash]
  • Research note added: jev in Hybrid execution, the first member that is a model rather than a mechanism around one; seven sources all fetched 200 or via the GitHub/Algolia APIs this run (launch post, docs, evals site, the MIT adapter repo, the 1,915-point launch thread, the 164-point jevlike reverse-engineering as the critical source, the antibenchmaxxing post) [glm-5.3-flash]
  • Facts recorded: three primitives (Choice, Score, Noul) evaluated in parallel against a state with calibrated confidence and a 255 cardinality cap, RLCD training, 70-500ms and $0.042/MTok-input with free output, the vendor-admitted biases (own capabilities team built the evals, Astra-plus-Fable reference, laptop-measured latency, non-empirical no-hallucination figure, unproven subsidy), the HN skepticism, and the one-day community jevlike reproduction [glm-5.3-flash]
  • hybrid-execution-feature-matrix extended from four to five columns in the same run per the matrix rule (Jev inserted alphabetically between Instructor and OpenAI), the guarantee-mechanism reading rewritten as a three-way spectrum, a choosing bullet added, references extended by 3, updated bumped [glm-5.3-flash]
  • _index.md: Jev listed in Hybrid execution, the matrix one-liner recast for the third guarantee class [glm-5.3-flash]
  • Verification: all internal link targets exist on disk, all seven cited URLs fetched this run, front matter parses with no type field, mandatory tags present, no em-dashes or banned terms, one sentence per line [glm-5.3-flash]

2026-09-20 (owner-prompted, price history retrofit)
#

  • Owner rule added to AGENTS.md: any article whose Pricing section states actual prices must carry a Price history section immediately after it, an append-only table (Date, Plan, Change, Source, oldest first) seeded with the baseline price and extended in the same run as any price change; notes whose pricing does not apply get no table [glm-5.3-flash]
  • Retrofit run across all research notes: 58 notes stating actual prices received a seeded Price history table (baseline row per note, plus historical rows where the note already recorded churn such as trae’s September repricing, cmux’s tier churn, amp’s Free Agent change, ellipsis’s 2024 per-seat model, cognee’s rate cut, omnara’s first published pricing), a Changes bullet, and updated bumped to 2026-09-20 [glm-5.3-flash]
  • 6 notes excluded as pricing-not-applicable with no nonzero price stated: math-inc-gauss, openai-for-science, harmonic-aristotle, exo, openhands, antigravity; openwork and windsurf already carried same-format Price history tables from the same-day refresh run and were left untouched [glm-5.3-flash]
  • Every table row is sourced to the note’s existing pricing reference; no new external fetches this run, and root essays, trackers, and matrices checked with no price-tracking content found [glm-5.3-flash]

2026-09-20 (daily refresh, full parallel re-verification + three entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: category sub-runs re-verified all research notes across the 21 categories, the essays, and all 22 matrices against live sources (GitHub via gh api, npm/PyPI/crates.io, HN Algolia, vendor and pricing pages); every “Facts below verified as of” line bumped to 2026-09-20, notes with material revisions got updated set plus Changes bullets [glm-5.3-flash]
  • Harnesses material: ante v0.2.1 (DeepSeek V4.1 Flash catalog), bullet v1.4.21, jcode v0.85.0 ended the release pause, kilo-code v7.7.5, pi v0.86.0 at 107.5k, grok-build alpha v1.0.38, kimi-code caught kimi-k3’s OpenRouter cut to $1.70/$8.50 (cache $0.17, Moonshot platform still $3/$15), deepseek-harness 230.5k still prerelease-only, aider stall re-confirmed [glm-5.3-flash]
  • Orchestration material: conductor 0.87.0, emdash v1.2.5, superset desktop v1.30.0, vibe-kanban’s community 0.1.45 tag cut with npm still serving 0.1.44 (orphan record kept), worktrunk’s release count corrected from 78 to about 150 (API recount, prior figure an undercount), gastown still quiet [glm-5.3-flash]
  • Assistant runtimes, memory, surfaces material: openclaw v2026.9.5, openwork’s pricing churned a seventh time (Enterprise back to $40/user, grandfather clause gone) with its Price history table seeded, picoclaw.io’s TLS cert expired 2026-09-10 recorded honestly, cognee v1.6.0 (keyless-first, breaking Docker change), engrim v1.4.8, memoryfields’ spec self-labels v0.3, openchamber v1.24.2, windsurf’s devin.ai/pricing now returns 200 with the Devin ladder published, trae’s repricing verified standing, cursor’s November 12 shutoff verified standing [glm-5.3-flash]
  • Elsewhere material: graphify v0.9.64, rtk’s RC train at rc.444, semble 82,697 trailing-month downloads re-measured, chonkie’s repos moved to the Feyn Labs org (feyninc/chonkie, chonkiejs), open-code-review 37.9k stars with v1.12.6/v1.12.7, github-agentic-workflows v0.89.17 prerelease (v0.88.7 still stable), deepeval’s PyPI figure corrected 3.5M to 2.8M/month and phoenix’s to 840K/month per the cited sources, harnesstax thread 229 points, agentsview’s schema caution corrected (usage JSON is v5, v6 is session export), skills-sh recomputed publisher totals downward (feishu 16.3M, larksuite 5.3M, runcomfy 4.6M) with azure-validate slipping below the audits window, agent-host-protocol’s npm window rebased on a mid-September spike, instructor 16.7M/month, nathan-lambert published “Why I still haven’t bought into true RSI” (09-19), deeplearning-ai’s Batch at issue 371 [glm-5.3-flash]
  • Essays: model-selection folded the kimi-k3 cut and added GLM-5.3-FlashX (released 2026-09-18, $0.37/$1.25, cache $0.075) as the table’s newest column with the challengers prose and clocks updated; landscape as-of bumped with Orca about 73k, Kimi Code 18.9k, Cline installs about 5.4M, and OpenHands’ repo move to the OpenHands org folded in; context-management-patterns and the-tells-are-structural link-checked clean [glm-5.3-flash]
  • Entrants accepted, three notes written to citation standards with all sources fetched this run: knowhere (Ontos AI’s structure-preserving parsing pipeline, 3,391 stars, $0.015/page) into Retrieval; memex (the MIT Rust session-indexer with BM25 plus local embeddings and resume-in-place, 222 stars, thin-footprint signal stated) into Session analytics; cua-s1 (Cua’s open-weights 706k-parameter System One form-decision checkpoint, vendor-run synthetic-only evidence recorded) into Hybrid execution; retrieval, session-analytics, and hybrid-execution matrices extended in the same run per the matrix rule [glm-5.3-flash]
  • Entrant rejections with evidence, all resolved in-run: the Jev model wave (laya 1,910 stars, localjev 517, bespokelabsai/nimble 674, Jeff/CRT) is model layer or below bar, repopilot (152 stars, no thread), coldteadotai/abide (rules enforcer, not a runtime), ENZO (16 points, general platform), geo-sleuth and the community-skill wave (standing exclusion), Armature (launched 2026-08-03, outside the window, hosted vendor analytics), herdr-projects (256 stars, a Herdr plugin), plus sub-bar HN tails [glm-5.3-flash]
  • Gastown note corrected: its status still said active-but-cooling while the Yegge note and the tracker carried the shutdown, so the note now records the September 2026 shutdown (AINews 2026-09-17 citing Dan Luu 2026-09-15, reference added) and the orchestration matrix’s status cell moved to shut down [glm-5.3-flash]
  • Matrices: all 21 category matrices plus the provider matrix re-verified in the same run, membership matched, columns sorted, uniform rows, re-dated to 2026-09-20; cell moves where facts moved (orchestration Gas Town/Vibe Kanban/Omnara, assistant-runtimes OpenClaw/Hermes/PicoClaw, code-review OpenCodeReview, evaluation HarnessTax, spec-driven Spec Kit adoption, task-management beads and Ordewell, control-planes Paperclip issues, sandboxing OpenShell contributors, software-factory stars, retrieval Knowhere column, hybrid-execution CUA-S1 column, session-analytics Memex column); provider matrix Moonshot flagship and cache cells moved with the kimi-k3 cut [glm-5.3-flash]
  • _index.md: three new notes listed alphabetically, the three matrix one-liners re-counted, all dated one-liners re-dated to 2026-09-20, and the pre-existing LlamaIndex-before-LangChain ordering in Retrieval fixed [glm-5.3-flash]
  • Coexistence and attribution: a concurrent owner-directed pass added the Price history rule to agents/AGENTS.md and swept Price history tables across priced notes in the same working tree (its AGENTS.md edit and price-table edits are swept into this commit and attributed here per the concurrent-run precedent); this run’s sub-runs coexisted without duplication or reverts [owner, glm-5.3-flash]
  • Process: several sub-runs were interrupted mid-flight (model rate limits and session interrupts); their partial work was verified and the remaining essays, provider-matrix, gastown, and index work completed by the coordinator directly; one sub-run disclosed two read-only git status --porcelain calls against its no-git instruction (nothing staged or mutated) [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered, and one new question added about the AGENTS.md price-history rule [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus three entrant notes with their matrix columns filled the day at quality [glm-5.3-flash]
  • Verification: 175 article files plus 5 control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, all 21 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (provider matrix at 13 rows by 8 cells), no em-dashes, no banned terms beyond the documented exceptions (ante’s protocol-shape crate name, the log’s own historical mentions, emdash’s append-only Changes history), one sentence per line, and every URL cited by new or edited content was fetched 200 this run or its absence recorded honestly (pypistats 429 twice, venturebeat 429 transient, x.ai changelog 403, junie pricing 404 with tiers keeping their 09-09 date) [glm-5.3-flash]

2026-09-20 (daily refresh, full parallel re-verification + three entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: 13 category sub-runs re-verified all research notes across the 21 categories, the four essays, and all 22 matrices against live sources (GitHub via gh api, npm/PyPI/crates.io, HN Algolia, vendor and pricing pages); every “Facts below verified as of” line bumped to 2026-09-20, notes with material revisions got updated set [glm-5.3-flash]
  • Entrants accepted, three, each written to citation standards with all sources fetched this run and their matrix columns added in the same run: knowhere (Ontos AI’s structure-preserving document parser and retriever, 3,391 stars, per-page API plus Apache-2.0 engine and MCP server, self-reported benchmark and near-empty HN footprint recorded) into Retrieval; memex (the Rust session-transcript searcher with resume-in-place, 222 stars, thin-footprint signal stated) into Session analytics; cua-s1 (Cua’s open-weights 706k-parameter System One form-decision checkpoint, 76-point Show HN, vendor-run synthetic-only evidence and the weights-distributed-versus-docs wrinkle recorded) into Hybrid execution [glm-5.3-flash]
  • Entrant rejections with evidence, all resolved in-run: Agentgit, S1Code, abide, repopilot (rules/app layers, not harnesses), jev-review (MCP plugin), MonkeyDcode, Forcefield, Lodestar, ENZO (platform, not a category member), Captain Memo, Skillmem, agentmemoryleaderboard (traction), herdr-projects (a Herdr plugin, sibling recorded in plannotator), Skillsync (session portability, standing rejection), nimble (typed-decision model, not eval infra), laya, localjev, Jev wave (model layer), Novgraph, Nerra, Axiom KG, Scry, Cactus Needle 3, Continuum OrcaBonsai, geo-sleuth, kbai-skill, the community-skill wave (standing exclusion), Jeff (27 points but 31 stars, below bar), CRT, Vigil-PR, Talos, Text-me [glm-5.3-flash]
  • Harnesses material: ante v0.2.1 (09-19), bullet v1.4.21 with downloads declining, grok-build v1.0.38 still release-less, jcode v0.85.0 (09-19) ended the release pause, kilo-code v7.7.5, kimi-code caught kimi-k3’s OpenRouter cut to $1.70/$8.50 with cache hits $0.17 (Moonshot’s platform page still lists $3/$15), pi v0.86.0 (09-19), claude-code 146.8k stars with issues+PRs down to about 12.3k, codex 125.4k, cline 5.39M installs, deepseek-harness 230,493 stars still prerelease-only [glm-5.3-flash]
  • Orchestration material: conductor 0.87.0 (09-18), emdash v1.2.5 (09-18), superset desktop v1.30.0 (09-19), vibe-kanban’s community train tagged 0.1.45 on 09-19 with npm still serving 0.1.44 (orphan record holds), worktrunk release count corrected to about 150 (prior figure an undercount), gastown status moved to shut down per Latent Space AINews (2026-09-17) citing Dan Luu’s 2026-09-15 post, Yegge admitting he never successfully built anything with it; the note, the matrix cell, and the landscape tracker now all carry the shutdown, repairing the split where only the Yegge note had it [glm-5.3-flash]
  • Assistant runtimes material: openclaw v2026.9.5 (09-19) with issues 8,071, openwork’s pricing churned a seventh time (Enterprise back to $40/user, grandfather clause gone, Price history table reconstructs all seven states), picoclaw.io’s TLS certificate expired 2026-09-10 and is still expired (recorded honestly, docs subdomain the working entry point), qwenpaw v2.2.2b3 betas [glm-5.3-flash]
  • Memory, context, review, eval material: cognee v1.6.0 (09-18) with GLiNER dropped from its default image, engrim v1.4.8, memoryfields spec self-labeled v0.3, graft’s repo canonically moved to trailhq/Graft, graphify v0.9.64, chonkie moved to the Feyn Labs org with the TS port renamed chonkiejs, semble 82,697 trailing-month downloads, open-code-review 37,926 stars with v1.12.6/v1.12.7 and still no HN catalyst, github-agentic-workflows v0.89.17 prerelease (v0.88.7 last stable), deepeval downloads corrected 3.5M to 2.8M/month and phoenix 1.3M to 840K/month per pypistats, harnesstax thread 229 points [glm-5.3-flash]
  • Elsewhere material: skills.sh recomputed cumulative installs downward (feishu 17.8M to 16.3M, larksuite 11.3M to 5.3M, runcomfy 11.2M to 4.6M) with azure-validate slipping below the audits page’s 50-row window, agent-host-protocol’s npm window rebased on a real mid-September spike (15,681/day on 09-18, 103,060 for the month), instructor downloads corrected to about 16.7M/month, agentsview’s schema caution corrected (usage JSON is schema version 5, version 6 is session export), augment-code’s blog cadence extended to Sept 11, windsurf’s pricing route healed (devin.ai/pricing now 200 with the Devin ladder published: Free, Pro $20, Team $80 plus $40/seat), openchamber v1.24.2 (10.1k stars), nathan-lambert’s new essay recorded (2026-09-19), deeplearning-ai Batch at 371 [glm-5.3-flash]
  • Essays: model-selection moved with the kimi-k3 cut (table, prose, decision bullet, clocks, Changes) and added Z.ai’s GLM-5.3-FlashX (released 2026-09-18, $0.37/$1.25, cache $0.075) as a twenty-first column; every other tabled price re-verified unchanged; landscape as-of bumped to 2026-09-20 folding Orca about 73k, Kimi Code about 18.9k, Cline about 5.4M installs, and OpenHands’ repo move to the OpenHands org at about 89k stars; context-management-patterns and the-tells-are-structural link-checked clean (12 external references 200), untouched [glm-5.3-flash]
  • Matrices: all 21 category matrices plus the provider matrix re-verified in the same run, membership matched, columns sorted, re-dated to 2026-09-20; retrieval 5 to 6 columns (Knowhere), session-analytics 3 to 4 (Memex), hybrid-execution 5 to 6 (CUA-S1); orchestration Gas Town status cell moved to shut down; provider matrix Moonshot flagship cell moved to $1.70/$8.50 with the cache cell at one tenth [glm-5.3-flash]
  • _index.md: three new notes listed alphabetically, Knowhere placed and the LangChain/LlamaIndex order corrected while in the section, matrix one-liners re-dated and re-counted, essay dates re-dated [glm-5.3-flash]
  • Attribution: a concurrent owner-directed pass added the Price history rule to AGENTS.md and swept ## Price history sections plus price-history Changes bullets across the pricing notes this run; those edits landed in the working tree alongside this run’s work and are committed here swept-in, attributed to that pass; this run’s own memory sub-run applied the new rule to openwork’s seven-state pricing history mid-run [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus three entrant notes with their matrix columns and the kimi price move filled the day at quality [glm-5.3-flash]
  • Verification: 195 article files plus control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, all 21 category matrices match their categories’ membership with case-insensitively sorted headers and uniform rows (provider matrix at 13 rows by 8 cells), no em-dashes, no banned terms beyond the documented exceptions (ante’s protocol-shape crate name, the log and emdash append-only histories, the skills matrix’s owner-accepted 09-13 column order), one sentence per line, and every URL cited by new or edited content was fetched 200 this run or its absence recorded honestly (pypistats 429 twice with dated claims kept, venturebeat 429 transient, openai.com/x.ai not cited by new content) [glm-5.3-flash]

2026-09-20 (correction to the two entries above)
#

  • This run’s entry was appended twice in draft form and both drafts were committed: the first (175-file verification count) is the earlier partial draft, the second (195-file count, with the attribution line for the concurrent price-history pass) is the complete and authoritative record of this run [glm-5.3-flash]

2026-09-20 (correction to today’s entries above)
#

  • The two “daily refresh, full parallel re-verification + three entrants” entries above are the same run recorded twice (the first append landed in commit a789d74f, a duplicate of the fuller entry landed in a924aa90); the fuller second entry stands as the run’s record, and the “(owner-prompted, price history retrofit)” entry between them is the concurrent pass’s own record [glm-5.3-flash]

2026-09-21 (daily refresh, full parallel re-verification + three entrants)
#

  • Full parallel refresh under the no-stalest-ranking rule: 8 sub-runs (the harnesses run relaunched once after a rate-limit death) re-verified all 197 article files across the 21 categories, the four essays, and all 21 category matrices plus the provider matrix against live sources (GitHub via gh api, npm/PyPI/crates.io, HN Algolia, vendor and pricing pages); every “Facts below verified as of” line bumped to 2026-09-21, notes with material revisions got updated set plus Changes bullets [glm-5.3-flash]
  • Entrants accepted, three, each written to citation standards with all sources fetched this run and their matrix columns added in the same run: zcode (Z.ai’s official GLM workbench, 4,216 stars within a day of the September 20-21 flattened open-source dump, after the September 18 wire-level exfiltration disclosure at 333 and 261 points and the 511-point July launch, GLM Coding Plan baseline seeded into its Price history) into Harnesses, flagged by three other sub-runs and resolved by the harnesses run; ax (google/ax, the Apache-2.0 Kubernetes-native agent orchestrator with Task/Workspace/Gateway/Model manifests and sub-second checkpoint-resume, launch thread 450 points as of today) into Orchestration; laya (ConvAI Innovations’ Apache-2.0 open-weights decision-model family, launch thread 1,305 points as of today) into Hybrid execution, superseding the 2026-09-20 wave rejection because the model-layer objection no longer holds with Jev and CUA-S1 in the category and the thread cleared every entry bar since [glm-5.3-flash]
  • Rejections with evidence, all resolved in-run: thruwire/foreman re-reviewed (237 to 444 stars since 09-18 but still zero independent coverage, the days-old-repo bar holds, re-review on coverage); browser-use/jev-ultrafast (13.2k stars, 91-point thread) folded into the jev note as ecosystem evidence, an application of the model rather than a category member; the community-skill wave (kitze/skillbox plus six Jev-themed skill packs) held out per the standing scope exclusion; Skillsync (65 points), herdr-projects (269 stars, no coverage), Jeff, jevals, CRT, Friday, Cortex, and the sub-bar HN tails [glm-5.3-flash]
  • Harnesses material: ante v0.2.2 (09-20, gateway split to external ante-gateway), bullet v1.4.22 (09-21), jcode v0.86.0 (09-20) ended its release pause, pi v0.86.1 (09-20); claude-code 147,299 stars with issues+PRs down to about 12.2k, cline installs 5,396,407, opencode 209k, openhands 88.7k, deepseek-harness 231.7k still prerelease-only, aider stall re-confirmed [glm-5.3-flash]
  • Orchestration and surfaces material: paseo v0.9.0-beta.1/.2 (09-17/18), superset desktop-v1.30.1 (09-21), worktrunk v0.79.0 (09-21) with a breaking -x semantics change, vibe-kanban correction (no plain v0.1.45 tag exists, only a timestamped one plus a permtest tag, npm still 0.1.44); openchamber 10,199 stars, gastown shutdown record holds, cmux pricing stable this cycle, windsurf’s pricing ladder re-confirmed on devin.ai [glm-5.3-flash]
  • Assistant runtimes, memory, skills material: openwork’s pricing churned an eighth time (grandfather clause restored, Price history row appended), qwenpaw v2.2.2b3 (09-20), openclaw 390,164 stars; skills-sh’s downward recompute extended to microsoft/azure-skills (7.1M to 2.4M) while open.feishu.cn rose back to 16.4M, skillregistry.io at 16,228 downloads [glm-5.3-flash]
  • Review, evaluation, sandboxing material: open-code-review 38.7k stars (about 68 percent in eight days, still no HN catalyst), phoenix’s monthly PyPI figure moved again (840K to 770K, the 09-20 correction did not hold, recorded), deepeval 2.8M to 2.6M, flue’s release train moved to 2.1.0 (09-18), openshell v0.1.0-pre.4, clawk’s quiet window now 39 days, cubesandbox’s internal status inconsistency fixed [glm-5.3-flash]
  • Small categories material: beads shipped v1.3.1-rc.1 (09-21), agentsview’s schema caution corrected (usage JSON and session export both at v6 now), ctx v1.4.12 (09-20), agents-md recorded Claude Code shipping native AGENTS.md support in 2.1.277 (2026-09-18), ending the last big-harness holdout, with the protocol matrix adoption cell updated [glm-5.3-flash]
  • Hybrid execution material: cua-s1’s model card gained its first real eval (196 decisions, 100 percent top-1) and a head-to-head against hosted Jev (99.7 versus 83.6), jev’s open-model wave recorded inside its note [glm-5.3-flash]
  • Essays: model-selection re-fetched every provider pricing page with zero moves (verified-date line bumped only, no Changes bullet); provider matrix bump-only; landscape corrected the Jules line (the free tier now runs Gemini 2.5 Pro, Gemini 3 Pro moved to the paid tiers, per the fetched landing page) and gained the ZCode and AX lines with the verification chain refreshed; context-management-patterns and the-tells-are-structural link-checked clean, untouched [glm-5.3-flash]
  • _index.md: three new notes listed alphabetically, three matrix one-liners re-counted (27 harnesses, 15 orchestration, 7 hybrid execution, ZCode/AX/Laya the newest), all dated lines re-dated to 2026-09-21 [glm-5.3-flash]
  • Coordinator corrections in-run: the AX thread count drifted 414 to 450 and Laya’s 1,304 to 1,305 between the sub-runs’ fetches and the coordinator’s spot-check via the Algolia API; both notes’ current-state lines updated with as-of dates, creation bullets left as historical record; the ZCode thread counts (511/333/261) verified exact [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the three ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification plus three entrant notes with their matrix columns filled the day at quality [glm-5.3-flash]
  • Style exceptions this run: laya’s “Honest Limits” is the model card’s quoted section title (a proper name, same rule as ante’s protocol-shape crate name), and the skills matrix keeps the owner-accepted 2026-09-13 column order [glm-5.3-flash]
  • Verification: 197 article files plus 5 control files pass YAML front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, all 21 category matrices match their categories’ membership with case-insensitive sorted headers and uniform rows (provider matrix at 13 rows by 8 cells), no em-dashes, no banned terms beyond the documented exceptions, one sentence per line, and every URL cited by new or edited content was fetched 200 this run or its absence recorded honestly (x.ai not cited, oreilly.com and thoughtleaders.io bot walls kept as dated artifacts) [glm-5.3-flash]

2026-09-21 (owner-prompted, Jev open-source/weights alternatives)
#

  • Owner request in chat: cover the open-source/open-weights alternatives to Jev; the category already held Laya and CUA-S1, so the run resolved the rest of the launch-week wave and scanned for anything new (HN Algolia, GitHub search, HuggingFace) [glm-5.3-flash]
  • Research index: Hybrid execution extended from seven to twelve members, five notes written by two parallel sub-runs to full citation standards with every cited URL fetched this run: jevlike (the one-day community reverse-engineering, MIT, 1,122 stars, 166-point thread, dormant since launch day, its AttentionHead lifted by CUA-S1), semif (TheoLeeCJ’s project until 2026-09-18 named OpenJev, the 713-point thread found under the old name, frozen-model logit readout with a client-side WebGPU demo, rename and trademark disclaimer recorded), kev (Jared Palmer’s Apache-2.0 Qwen3.5 LoRA family speaking the TypeSafe SDK contract locally, accepted below the 100-point bar at about 30 points on author standing plus the external-eval ecosystem and its pre-registered locked-test discipline), nanojev (the 0.6B game-task replica with the largest star count of the replicas, 1,633, and the smallest footprint, 2 points, accepted on traction plus complete artifacts with every number weighted unreplicated), and nimble (Bespoke Labs’ Apache-2.0-weights LoRA with the category’s only human-labeled head-to-head against Jev, 3,880 records, Jev 76.0 versus 74.8 macro, repo code unlicensed as the lead caution) [glm-5.3-flash]
  • Rejections with evidence: localjev (649 stars, 2-point thread, a wire-compatibility bridge whose probabilities are self-reported rather than read from logits per its own README); applications built on Jev stay ecosystem evidence, not notes (browser-use/jev-ultrafast 13.4k, fast-jev-compaction 5.6k, jev-trader 1.6k, QuantDinger 11.9k); curated lists (awesome-jev 756, awesome-jev-by-typesafe 734, awesome-jev-tools 589) stay named in prose; kev-adjacent one-offs (Jev-cu 511, jev-review, the razorback16/openjev DiffusionGemma server at 240 stars noted as a future candidate) below bar [glm-5.3-flash]
  • hybrid-execution-feature-matrix: extended from seven to twelve columns in the same run per the matrix rule, all 55 new cells traced to the new notes’ sources; intro and thesis rewritten for the seven-open-brackets framing, spectrum reading gains the dormancy-and-evidence spread, choosing gains four wave bullets, references extended by 10 [glm-5.3-flash]
  • jev: the ecosystem caution now links the five new sibling notes, Changes bullet appended [glm-5.3-flash]
  • _index.md: five notes listed alphabetically in Hybrid execution, the matrix one-liner recast for twelve columns [glm-5.3-flash]
  • Sub-run disclosures: one sub-run ran a read-only git status --porcelain against its no-git instruction (changed nothing, self-reported); a SemIf “Jev-shaped” phrasing and a first-draft “Jev-like” title issue were caught and fixed in-run by the style sweeps [glm-5.3-flash]
  • Verification: all new notes pass front-matter parsing with no type field and mandatory tags, one sentence per line, no em-dashes, no banned terms beyond the documented exceptions, all internal link targets exist on disk, and every cited URL was fetched 200 this run (NanoJev’s demo site 401, not cited; SemIf raw README 404 on main, cited via the master branch path) [glm-5.3-flash]

2026-09-21 (incident: commit b32de458 swept pre-staged owner files)
#

  • The open-alternatives commit b32de458 contains 25 paths outside agents/ that this run did not stage: 24 owner modifications to corpus articles plus the owner’s deletion of article-marketing/index.md, all found pre-staged in the git index when the run committed (the run’s own git add agents/ staged only agents/ paths, and the previous commit 1e978cb0 was verified clean of outside paths) [glm-5.3-flash]
  • git commit commits the whole index, so the staged owner work was published alongside the section’s changes; per the guardrails the run will not revert owner edits or force-push, so the commit stands and this entry is the record [glm-5.3-flash]
  • Process fix adopted for future runs: before committing, verify the staged set is exactly agents/ (git diff --cached --name-only must list no path outside agents/) and stop and log if foreign files are staged, rather than committing around them [glm-5.3-flash]
  • Flagged to the owner via ? for tom: in queue.md: keep the swept files (they were finished working-tree edits) or revert them in an owner-authored commit [glm-5.3-flash]

2026-09-24 (owner-prompted, category index pages)
#

  • Owner request in chat: give each category its own index page with its own Changes log; AGENTS.md was updated first (new Category index pages section, hub model for _index.md, link depths, refresh wiring), and this run performed the content migration [glm-5.3-flash]
  • Created 21 category index pages at agents//_index.md per the new spec: scope sentence, notes alphabetical with one-line summaries, feature matrix link, and a seeded ## Changes with one Added <Title>. bullet per entry dated to that entry’s created date, 180 entries across the 21 categories [glm-5.3-flash]
  • _index.md restructured to the hub model: the Research index now links the 21 category index pages with one-line blurbs and lists no research notes inline; the mission statement, the essays and trackers listing, and the Comparison matrices listing are unchanged [glm-5.3-flash]
  • Index lag fixed in the same run: foremerge (orchestration) and jevbench (hybrid-execution), both finished and matrix-columned on 2026-09-22, had never been listed in _index.md; both are listed on their category pages with one-line summaries [glm-5.3-flash]
  • Matrix title drift fixed in the same run: the surface matrix’s JetBrains column renamed to JetBrains IDEs, the context engines matrix’s Sourcegraph column to Sourcegraph code context platform, and the memory matrix’s Files column to File-based agent memory and Mem0 to mem0, each now matching the section index listing, affected reference lines updated, no cells moved, Changes bullets appended, updated dates bumped [glm-5.3-flash]
  • Verification: 21 pages pass YAML parsing with the specified front matter and no type or menu field, every page’s list matches its directory membership and its matrix’s columns, all internal link targets exist on disk, the hub lists no notes inline with the mission statement byte-identical, Changes bullets are oldest first, one sentence per line, no em-dashes, no banned terms [glm-5.3-flash]

2026-09-24 (owner-prompted, section index ordering)
#

  • Owner request in chat: the section index ordering recommendation adopted; Comparison matrices promoted from a ## subsection of Essays and trackers to its own # section, removing the mis-nesting left from the matrices block’s 2026-08-24 append [glm-5.3-flash]
  • _index.md: the three sections now read essays and trackers, comparison matrices, research index (argument, comparison, reference); the essays list reversed to newest first (The Tells Are Structural first), and the matrix list and category list sorted alphabetically case-insensitive, which drops the model-provider matrix’s privileged first slot and fixes code-review’s drifted matrix position [glm-5.3-flash]
  • AGENTS.md: step 6 now documents the section index ordering rule (essays newest first, matrix and category lists alphabetical, case-insensitive) so the daily run maintains it [glm-5.3-flash]
  • Verification: mission statement byte-identical, exactly three top-level headings with blank lines after each, single trailing notes line, all internal link targets exist on disk, both sorted lists verified, list contents unchanged [glm-5.3-flash]

2026-09-24 (incident: commit f26ed29e swept concurrent-run notes)
#

  • Commit f26ed29e (“Reorganize agents index into three sections with sorted lists”) contains six people-and-publications notes this run did not write: armin-ronacher, boris-cherny, dax-raad, mario-zechner, peter-steinberger, thorsten-ball, all created 2026-09-24 by a concurrent session and swept in by the run’s git add agents/ (the run’s staged-set check only rejects paths outside agents/, so same-directory foreign content passed) [glm-5.3-flash]
  • State at discovery: the six notes are finished with citations and Changes bullets but unintegrated, the category index page listed thirteen members against nineteen directories and the people matrix held thirteen columns, with no log entry covering the notes [glm-5.3-flash]
  • The concurrent session stayed active after discovery (dex-horthy, matt-pocock, and walden-yan directories appeared on disk after the commit), so per the owner’s call this run keeps the commit without a history rewrite, touches nothing under people-and-publications, and leaves the page, matrix, and remaining integration to the session that owns the work [glm-5.3-flash]
  • Process fix proposed to the owner: stage the exact paths this run touched instead of the wholesale git add agents/, so a concurrent run’s in-flight files are never swept [glm-5.3-flash]
  • The index reorganization the commit message describes was verified complete before the sweep: three top-level sections, essays newest first, matrix and category lists alphabetical, mission statement untouched [glm-5.3-flash]

2026-09-24 (owner-prompted, harness authors added to People and publications)
#

  • Owner request in chat: add Matt Pocock, Dex Horthy, Armin Ronacher, Mario Zechner, Dax Raad, Peter Steinberger, Boris Cherny, and other well known harness authors; this run adds those seven plus Thorsten Ball (Amp co-founder, Register Spill) as the clearest additional harness author with a publication, eight notes total [glm-5.3-flash]
  • Owner request overrides two logged rejections: armin-ronacher (rejected 2026-09-02 and 2026-09-05 as duplicating the hands-on-text band) and dex-horthy (rejected 2026-09-08 as having no sustained publication); both notes define their distinct slots and state the prior-rejection boundary honestly in Cautions [glm-5.3-flash]
  • Commission correction: the briefing’s Walden Yan (attributed as Cline co-founder and context-engineering essay author) could not be verified from any fetchable source, so the commissioned ninth note was dropped unpublished rather than published on an unverified attribution [glm-5.3-flash]
  • Commission correction: the briefing’s premise that Dax Raad founded Exo Labs is unverified (fetched sources associate Exo Labs with Alex Cheema; exoharness contributor list excludes him); the dax-raad note documents his verifiable record (SST, OpenCode at Anomaly, Bumi) and records the discrepancy in Cautions; the exo note makes no founder claim and needed no change [glm-5.3-flash]
  • Eight notes written to citation standards by three parallel sub-runs (minimum 5 sources each, 6 to 9 references, every cited URL fetched this run, as-of dates 2026-09-24), no type field, no em-dashes, no banned terms, one sentence per line [glm-5.3-flash]
  • people-and-publications-feature-matrix: extended from thirteen to twenty-one columns, re-sorted alphabetically, every new cell traced to its member note, thesis and reading rewritten for the new harness-builder cluster, choosing list extended by eight, intro re-counted, updated bumped [glm-5.3-flash]
  • Category index: eight entries listed alphabetically with one-line summaries, eight ## Changes bullets appended dated 2026-09-24; _index.md people matrix one-liner re-counted to twenty-one voices and re-dated [glm-5.3-flash]
  • Verification: all twenty-one member columns match the category directory listing, all front matter parses, all internal link targets exist on disk, matrix header sorted case-insensitively, Changes bullets oldest first [glm-5.3-flash]

2026-09-24 (owner-prompted, verification preambles removed)
#

  • Owner request in chat: remove the “Everything below was verified …” preamble line from all matrix pages as low value [glm-5.3-flash]
  • Removed the preamble line from all twenty-one feature matrices and the two trackers it had spread to (model-selection-for-coding-tasks, context-management-patterns); the trackers keep a single as-of clause in their first paragraphs per the quality bar [glm-5.3-flash]
  • Matrix legends shortened from “? not verified as of the date above” to “? not verified”, since the referenced verification date no longer exists above [glm-5.3-flash]
  • AGENTS.md updated: matrices are barred from verification preamble lines (verification history lives in ## Changes, the updated field carries the revision date, volatile cells self-date), and trackers must carry their as-of as one short clause, not a run log [glm-5.3-flash]
  • Each touched article gained a dated ## Changes bullet and an updated field bump to 2026-09-24; no matrix cells, membership, or sorting changed, editorial only [glm-5.3-flash]

2026-09-24 (owner-prompted, top 5 recommended reading sections) #

  • Owner request in chat: add a “top 5 recommended reading” section to the people and publications pages; AGENTS.md updated first with the section rule (between Bottom line and Changes, five pieces by or featuring the person, best first, one line each, third-party coverage excluded, thin-record voices may name the gap instead of padding, every URL fetched in the run that adds or edits it) [glm-5.3-flash]
  • All 21 member notes received the section, written by three parallel sub-runs; every recommended URL was fetched 200 this run, titles confirmed against the fetched sources, and the section’s addition recorded as the last ## Changes bullet in each note [glm-5.3-flash]
  • Selections start from each note’s own references and favor the pieces the category already cites (Willison’s agentic-engineering patterns, Horthy’s 12-factor repo and talk, Zechner’s pi and MCP posts, Husain’s eval field guide, Weng’s agent survey, Lambert’s agent posts, the Claude Code and Yegge interviews); boris-cherny exercises the documented fifth-line exception (no fifth durable piece exists, stated rather than padded) [glm-5.3-flash]
  • Verification: 21 of 21 notes carry the section in the correct position with five lines each, the dated bullet last in Changes, no em-dashes, no banned terms, all internal link targets still resolve; section order in every note is Bottom line, Top 5 recommended reading, Changes, See also, References [glm-5.3-flash]

2026-09-24 (incident: commit 129ef15e swept concurrent-run preamble removals)
#

  • Commit 129ef15e (“Add top 5 recommended reading sections to people and publications notes”) also contains this session’s preamble-removal changeset, swept in while it was in flight: the “Everything below was verified …” line removed from twenty-one matrices and two trackers, legend “? not verified as of the date above” shortened to “? not verified”, compact as-of clauses kept in the two tracker intros, AGENTS.md matrix and tracker rules added, and per-article ## Changes bullets plus updated bumps [glm-5.3-flash]
  • State at discovery: the swept changeset is complete and verified in HEAD (all twenty-three files carry the Changes bullet and the 2026-09-24 updated bump, no preamble or dangling legend reference remains), the two changesets touch disjoint lines, and the same commit carries this session’s log entry above [glm-5.3-flash]
  • Per the f26ed29e precedent this run keeps the commit without a history rewrite; the misleading commit message is accepted because the log is the audit trail [glm-5.3-flash]

2026-09-24 (owner-prompted, six methodology and research voices added)
#

  • Owner request in chat: add the strong candidates from this run’s entrant scan; six notes written by three parallel sub-runs: geoffrey-huntley, jesse-vincent, kent-beck, harper-reed, shreya-shankar, paul-gauthier, each to citation standards with every cited URL fetched this run and a Top 5 recommended reading section per the new rule [glm-5.3-flash]
  • Entrant scan method recorded: websearch was down (HTTP 400), so discovery ran on HN Algolia API fetches plus in-corpus cross-references; borderline candidates named to the owner but not written (Birgitta Böckeler venue unverified, Aiden Cunniffe zero HN footprint, Tobi Lütke and Mitchell Hashimoto overlapping slots, media-lane voices rejected) [glm-5.3-flash]
  • people-and-publications-feature-matrix: extended from twenty-one to twenty-seven columns, re-sorted alphabetically, every new cell traced to its member note; reading gains the workflow-and-methodology cluster (Huntley, Vincent, Beck, Reed) and Gauthier joins the harness-builder cluster as its pre-agentic root; choosing list extended by six; the verification-history line lost in the 129ef15e sweep restored in concise form [glm-5.3-flash]
  • Category index: six entries listed alphabetically with one-line summaries and six ## Changes bullets dated 2026-09-24; _index.md people matrix one-liner re-counted to twenty-seven voices [glm-5.3-flash]
  • Verification: 27 of 27 member columns match the category directory listing, all six new notes pass front-matter, tags, section-order, top-5, and link checks (one wrong-depth software-factory link in geoffrey-huntley caught and fixed to the category index page), no em-dashes, no banned terms [glm-5.3-flash]

2026-09-24 (owner-prompted, verification preamble removed)
#

  • Owner correction in chat: the people matrix verification-history line this run restored after the 129ef15e sweep is not wanted; AGENTS.md now forbids verification preamble lines on matrices (verification history lives in ## Changes, volatile cells carry their own as-of qualifiers), so the line is removed and a Changes bullet records the removal to prevent future restores; the people matrix was the last matrix carrying a preamble [glm-5.3-flash]

2026-09-24 (owner-prompted, three video channels added)
#

  • Owner request in chat: cover the YouTube channels @owainlewis, @RAmjad, and @indydevdan; three notes written to citation standards (channel pages, RSS feeds, sites and repos, HN Algolia evidence, every cited URL fetched this run), each with a Top 5 recommended reading section of fetched videos [glm-5.3-flash]
  • people-and-publications-feature-matrix: extended from twenty-seven to thirty columns, re-sorted alphabetically, every new cell traced to its member note; reading gains the video-band paragraph splitting the three channels by what they optimize (IndyDevDan weekly worldview, Ray Amjad release analysis, Owain Lewis clone-and-run builds); choosing list extended by three [glm-5.3-flash]
  • Category index: three entries listed alphabetically with one-line summaries and three ## Changes bullets dated 2026-09-24; an edit that briefly dropped the Simon Willison line was caught and fixed in-run; _index.md people matrix one-liner re-counted to thirty voices [glm-5.3-flash]
  • Verification: 30 of 30 member columns match the category directory listing, matrix header case-insensitively sorted, all front matter parses, all internal link targets exist on disk, Changes bullets date-ordered, no em-dashes, no banned terms [glm-5.3-flash]

2026-09-24 (owner-prompted, model benchmark matrix added)
#

  • Owner request in chat: add a matrix/index covering the model benchmarks, each with a one to two sentence summary of what it evaluates [glm-5.3-flash]
  • Created agents/model-benchmark-matrix/ at the section root beside the model provider matrix: thirty benchmarks in six groups (agentic software engineering, function-level code generation, agents at work, reasoning and knowledge, math, context/instructions/preference), each row carrying a one to two sentence what-it-evaluates summary plus reading notes [glm-5.3-flash]
  • Curation rule stated in the article (official site or paper, third-party checkable results, present in recent comparisons); harness benchmarks (FrontierHarness Eval, HarnessTax) and JevBench named as out-of-scope pointers; Commit0 dropped because its page is dead, its from-scratch niche covered by the SWE-bench team’s ProgramBench [glm-5.3-flash]
  • _index.md comparison-matrices list gained the entry in alphabetical position and the intro sentence extended to cover benchmark instruments [glm-5.3-flash]
  • Verification: all cited external URLs fetched this run (swe-rebench.com returned an oversize body, noted in the reference; openai.com blocked direct fetches so PaperBench grounds via search snapshot of the official page), all internal link targets exist on disk, front matter parses, no em-dashes, no banned terms [glm-5.3-flash]

2026-09-24 (owner-prompted, trackers and leaderboards category seeded)
#

  • Owner request in chat: add a matrix/index for “information” websites (name improved to Trackers and leaderboards), covering aireleasetracker.com and artificialanalysis.ai among others; six notes written to citation standards, each with every cited URL fetched 200 this run (aireleasetracker.com 429s automated fetchers, so its current-content evidence is the 2026-08-26 archived copy, recorded in the note) [glm-5.3-flash]
  • Membership: ai-release-tracker, artificial-analysis, epoch-ai, llm-stats, lmarena, openrouter-rankings, chosen from the owner’s two seeds plus the field’s major release, benchmark, dataset, aggregate, preference, and usage surfaces; discovery ran on direct fetches, HN Algolia, GitHub API, and Wayback because websearch was down (HTTP 400) [glm-5.3-flash]
  • trackers-and-leaderboards-feature-matrix: created with six columns sorted alphabetically and ten rows (kind, what the number measures, operator, coverage, cadence, methodology, data access, pricing, independence caveat, verification hooks), every cell traced to its member note; reading names the two standard misquotations (arena ranks as capability, gateway tokens as market share) [glm-5.3-flash]
  • Category index created with six one-line entries and six dated Added bullets; _index.md gained the matrix one-liner in Comparison matrices and the category line in Research index, both alphabetical [glm-5.3-flash]
  • LLM Stats is the only member with published prices, so it alone carries a Price history table (seeded with the 2026-09-24 baseline rows) [glm-5.3-flash]
  • Verification: 6 of 6 member directories match the index and matrix columns, all front matter parses, section order Changes/See also/References holds in every note, internal link targets exist on disk, columns case-insensitively sorted, no em-dashes, no banned terms [glm-5.3-flash]

2026-09-24 (owner-prompted, ARC-AGI series completed in benchmark matrix)
#

  • Owner follow-up in chat: did the new matrix cover all the ARC-AGI benchmarks; it did not, only ARC-AGI-2 had a row with ARC-AGI-3 mentioned in passing [glm-5.3-flash]
  • Fetched arcprize.org/arc-agi/1 and /arc-agi/3 this run and extended the reasoning group to three ARC rows: ARC-AGI-1 (800 tasks, unbeaten 2019 to December 2024, now saturated) and ARC-AGI-3 (interactive game environments, skill-acquisition efficiency scoring) join ARC-AGI-2, whose reading note lost the now-redundant ARC-AGI-3 pointer [glm-5.3-flash]
  • Intro states the total (thirty-two), two reference entries added, Changes bullet appended; _index.md one-liner re-counted to thirty-two [glm-5.3-flash]
  • Verification: new URLs fetched, internal links resolve, no em-dashes, no banned terms [glm-5.3-flash]

2026-09-25 (daily refresh, full parallel re-verification + unreal-agent entrant)
#

  • Full parallel refresh under the no-stalest-ranking rule: 8 sub-runs re-verified the research notes across the categories against live sources (GitHub REST and gh api, releases and atom feeds, npm, PyPI, crates.io, HN Algolia, vendor pricing and changelog pages); every “Facts below verified as of” line bumped, notes finished before the midnight rollover carry 2026-09-24 and the rest 2026-09-25, matching fetch times; people-and-publications (30 notes) and trackers-and-leaderboards (6 notes) were created or fully verified earlier today by four owner-prompted runs, so this run re-checked their state mechanically without re-fetching, per the 2026-09-02 same-day-churn precedent [glm-5.3-flash]
  • Entrant accepted: unreal-agent written into Harnesses (Unreal Labs’ MIT Go async-first harness, tool calls as background operations, 242-point launch thread on 2026-09-22 with 122 comments, 1,879 stars and 97 forks at day three, v0.2.0 of 2026-09-23, Sequoia and First Round backed, vendor-run Harbor benchmark tables with the HN baseline-mismatch critique and the Responses-API footnote recorded, pricing does-not-apply with no price history table); harness-feature-matrix extended from 27 to 28 columns (inserted between Qwen Code and Warp Agent CLI), reading gains the async-bet line, choosing gains the Go-embed bullet, references extended; harnesses category page lists it alphabetically with a Changes bullet; landscape tracker gains the tail line [glm-5.3-flash]
  • Entrants rejected with evidence, all resolved in-run: AnyJev (Nokia Applied Research, Apache-2.0, 506 stars in four days, active, but zero HN footprint and no independent coverage, sources short of the citation bar; re-review on a thread or second implementation), Whiteboard (YC W26 open-source IDE, 169-175 points on day one, the days-old bar holds), magpie (yetone/magpie, 612 stars, menu-bar model router, gateway infrastructure blocked on the standing gateways question), mikehasa/golive-skill (874 stars in two days, a single community skill, outside the Skills scope), Step Code (393 stars, four months quiet), jev-chat apps (6,068 stars, chat-reply application layer), Simreal-MLBench (57), jevry (99), dolphinbench (27), Radix (15 points), Canary (5), Critic (6), AgentRun (38 points, 29 stars), Riftri (5, no repo), plus the sub-bar HN and GitHub tail [glm-5.3-flash]
  • Harnesses material: bullet v1.4.23, deepseek-harness 235,140 stars with dsh-v0.1.7-rc.1 and rc.2 still prerelease-only, fx v0.0.11 shipped custom model connections and themes, jcode v0.88.0, juggler v0.7.0, kilo-code v7.7.9/v7.7.12, openhands v1.23.0, ante v0.2.4, grok-build npm v1.0.41 still release-less, cline 5,439,863 installs, claude-code about 148k stars, opencode 209.9k; junie’s pricing page 404s so its dated tiers stand unverified; aider stall re-confirmed [glm-5.3-flash]
  • Surfaces material: openchamber v2.0.0 GA (September 23) and v2.0.1 with 10,509 stars, kiro CLI 2.23.1 and IDE 1.1.70, continue installs 4,217,213 with rating 3.3, roo-code 2,032,553 installs still ticking on the dead repo; cursor, windsurf, trae, zed, vscode-copilot, antigravity, void, jetbrains verified standing [glm-5.3-flash]
  • Orchestration material: ax 10,392 stars with its launch thread at 662, paseo v0.9.2 at about 18.5k, foremerge v0.5.0 published, vibe-kanban’s release record corrected (a timestamped v0.1.45 prerelease object exists on GitHub since 2026-09-19, npm still 0.1.44, commits stopped 09-19), emdash’s release-gap claim corrected to three days (v1.2.6 on 09-21); gastown’s shutdown record, cmux’s pricing, conductor 0.87.3, superset v1.30.2, worktrunk v0.79.0 all re-verified standing [glm-5.3-flash]
  • Assistant runtimes and protocols material: hermes v2026.9.24 at 248,717 stars, openclaw v2026.9.6 (crashing macOS build rebuilt 09-24), openwork through v0.18.52 with pricing stable a fourth day, ag-ui on a 2026-09-23 release train with downloads refreshed (about 12.0M combined monthly), agent-host-protocol 246,064 lifetime crates plus 123,531 npm, a2a, acp, agents-md, mcp, nanobot, nanoclaw, picoclaw (TLS still expired), qwenpaw, zeroclaw, eigent verified [glm-5.3-flash]
  • Code review and evaluation material: open-code-review 40,775 stars with v1.12.9, qodo v0.46.0, sourcery’s Team tier reshuffled at unchanged prices, kodus drift; deepeval v4.2.4 with Jev integration and PyPI re-measured about 2.5M/month, langfuse v4.45.2, phoenix v20.16.0 with trailing-month downloads re-measured about 650k, plannotator v0.27.20, harnesstax 232 points with the AgentBRANE rebrand noted, frontierharness re-verified on the live site, ctx shipped v2.0.1 and withdrew the pro subscription (Pricing rewritten, Price history withdrawal row appended) [glm-5.3-flash]
  • Context, retrieval, skills, memory material: graphify v0.9.67 with its installer surface corrected to the documented 17 assistants, rtk’s v0.50.0 went stable, graft 9,172 stars, qmd and repomix drift, semble downloads dipped to 80,496; knowhere v1.2.17 with its $5 signup credit replaced by a 14-day trial on the live page (Pricing sentence and Bottom line fixed), chonkie PyPI 1.14M trailing month, langchain 147.0k; agent-native 6.8k stars with npm 0.3.1, skills-sh leaderboard reshuffled with skillregistry.io still up, skillopt and anthropic-agent-skills drift; cognee v1.6.1, engrim, mem0, letta, zep, memoryfields, claude-mem, file-based-agent-memory verified [glm-5.3-flash]
  • Hybrid execution material: jevbench materially revised (v1.4.0-v1.4.2 added 308 sealed decisions and a public-to-sealed gap penalty, re-scoring the board: Jev first at 63.3, new #2/#3 JevK5 and Hopper, SemIf 73.1 to 47.7, Kev and Nimble and Laya fell similarly), the sibling notes’ traction and board readings updated, jev’s two-board sentence rewritten and jev-ultrafast refreshed to 19,800 stars; cua-s1, instructor (8.4M/month, decline continuing), outlines, both structured-output notes verified [glm-5.3-flash]
  • Sandboxing and small categories material: opensandbox shipped its first stable 1.1.0 umbrella release (versioning caution rewritten), openshell pre.11 with issue count dropping 537 to 416, cubesandbox v0.7.2, flue 2.1.1 on npm, agent-sandbox v1.0.4, clawk’s quiet window now 43 days; spec-kit v1.0.11 at about 139k stars, gsd’s successor 9,827 stars with the under-15-percent claim corrected to about 15, ordewell v0.4.23 with the project-age fix, paperclip v2026.916.1 at 82,672 stars, har v1.14.3, ouroboros v0.54.5, machinist’s zero-coverage claim re-verified via an Algolia zero-hit search and re-dated, openai-for-science rewritten for the public Navier-Stokes Lean formalization and the Buckmaster-Alpoge Euler certificates (Stanford Tech Review audit added), bmad, openspec, tessl, backlog-md, beads, task-master, tinyagi, fluent, sssf, alphaproof, anthropic-claude-math, aristotle, gauss, deep-research, pion verified [glm-5.3-flash]
  • Essays: model-selection-for-coding-tasks caught the 2026-09-22 GPT-6 Sol and Luna launches (1,762-point thread, models.dev dates) and added both as columns (Sol $2/$10, the first OpenAI model at the converged workhorse rate; Luna $0.10/$0.50, half its predecessor), and re-grounded the Kimi rows on Moonshot’s official prices after the OpenRouter endpoints API showed the previously recorded cuts were third-party host floors rather than vendor list prices (kimi-k2.7-code $0.95/$4.00 cache $0.19, kimi-k3 $3/$15 cache $0.30, floors named in prose), corrected GLM-5.3-FlashX’s cached-read cell to the $0.090 Z.AI serves, and fixed the context-behavior row where the gpt-6-astra cell (doubles past threshold) and the Haiku 4.5 cell (200K) had been transposed; provider matrix moved the OpenAI workhorse and cheap cells to the GPT-6 pair and the Moonshot cells to list prices; landscape gained the Unreal Agent line and the OpenChamber v2.0.0 GA phrase; context-management-patterns and the-tells-are-structural link-checked clean with every external reference 200 [glm-5.3-flash]
  • Defects fixed in-run: the automated-research matrix still carried the verification preamble the 2026-09-24 sweep missed and used plain-text unlinked headers (preamble deleted, headers converted to member links, separator normalized, four OpenAI for Science cells updated for the Lean certificates); model-selection’s transposed context cells repaired; emdash’s release-gap sentence corrected [glm-5.3-flash]
  • Matrices: cells updated and dated where facts moved in harnesses, assistant-runtimes, protocols, orchestration, code-review, context-engines, retrieval, skills, hybrid-execution, evaluation-review, sandboxing, session-analytics, spec-driven-development, task-management, control-planes, software-factory, automated-research, and the provider matrix; memory, surface, and executions matrices re-checked with no cell movement; all 22 category matrices match their categories’ membership with case-insensitively sorted uniform rows [glm-5.3-flash]
  • Category index pages: one-line summaries refreshed in place where revised content contradicted them (graft, langfuse, ctx, openspec, vibe-kanban, hermes, nanobot, openclaw, ag-ui, deepseek-harness); membership changed only in Harnesses [glm-5.3-flash]
  • _index.md: both essay dates and eighteen matrix one-liners re-dated, the harness line re-counted to twenty-eight with Unreal Agent the newest, the hybrid-execution line extended with the JevBench re-scoring, the session-analytics line notes ctx pro’s withdrawal, the sandboxing line notes OpenSandbox’s first stable [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; the four ? for tom: questions (untracked the-agentic-development-environment-landscape/, breakout community skills, gateways category, the AGENTS.md price-history rule confirmation) remain unanswered [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification, the unreal-agent note with its matrix column, and the GPT-6 and Kimi repricing work filled the day at quality [glm-5.3-flash]
  • Process notes: the run started 2026-09-24 and crossed midnight, so in-content dates straddle 09-24 and 09-25 by fetch time; several sub-runs disclosed read-only git status or diff calls against their no-git instruction (nothing staged or mutated); the two table-column insertions were script-assisted and verified for header and row alignment afterward; my prompt’s HN epoch constant was wrong (2025) and three sub-runs caught and corrected it [glm-5.3-flash]
  • Verification: 230 article files plus control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 22 category matrices match their categories’ membership with sorted uniform rows (provider matrix 13 rows by 8 cells), no em-dashes, no banned terms beyond the documented exceptions (ante’s protocol-shape crate name), one sentence per line, and every URL cited by new or edited content fetched 200 this run or its absence recorded honestly (junie pricing 404 with tiers keeping their 09-09 date; api.npmjs.org served cached windows ending 09-21 for three notes, recorded as same-window re-reads; GitHub anonymous 403s mid-run worked around via gh api and atom feeds; pypistats 429 once, succeeded on retry) [glm-5.3-flash]

2026-09-25 (daily refresh, same-day mechanical pass + entrant scan)
#

  • Same-day second run, about six hours after the 01:14 UTC full refresh, so per the same-day-churn precedent it ran a mechanical integrity audit over all 253 article files (front matter, no type field, mandatory tags, section order, internal links, em-dashes, banned terms, matrix and category-index membership plus title ordering across all 22 categories) and an HN Algolia entrant scan for the 24-hour window, without re-fetching hours-old verifications [glm-5.3-flash]
  • Defects fixed, title-order violations under the case-insensitive member-title ordering the corpus already uses (punctuation significant, so space before hyphen before digit): executions index page and its matrix moved GitHub Agentic Workflows ahead of GitHub Copilot automations, protocols index page and its matrix moved ACP, Agent Host Protocol, and AG-UI ahead of A2A and AGENTS.md, and the skills matrix moved Agent Skills open standard ahead of Agent-Native; each touched matrix gained a dated Changes bullet and an updated bump, no cell content moved [glm-5.3-flash]
  • Pricing rule sweep of the ten flagged Pricing sections: nine are free-and-open-source or no-published-price subjects and correctly tableless (harmonic-aristotle, math-inc-gauss, openai-for-science, crush, exo, qwen-code, kev, the-pragmatic-engineer, beads); antigravity states $0/month, so it gained the required Price history table seeded with the 2026-08-24 individuals baseline, sourced to the live pricing page fetched this run, plus a Changes bullet and updated bump [glm-5.3-flash]
  • Entrants resolved in-run: Whiteboard’s rejection stands (272 points now, still a days-old IDE); bestvaluemodel rejected (171-point thread 2026-09-24 but a one-day-old repo with 10 stars, no license, one HTML file deriving entirely from member Artificial Analysis; re-review on sustained operation and a license); llm-list.com rejected (1,794-model catalog, six-hour refresh, but fully derivative of models.dev, OpenRouter, Arena, Artificial Analysis, and Hugging Face, zero HN footprint, no named operator; re-review on independent coverage); the Di Zhang RLCD deconstruction (65-point thread 2026-09-24) added to the jev note’s Status and References as an independent technical reading (Plackett-Luce plus Brier calibration over packing and masking, no new sampler), with its MoJev reproduction cited from the live MoLeMo-Lab/mojev repo after the post’s original trotsky1997/jevre links 404ed; MoJev itself stays uncategorized for now (MIT, wire-compatible, 93.23 percent accuracy at 0.79 ECE, but 29 stars and self-published; re-review on independent traction) [glm-5.3-flash]
  • No essay, tracker, matrix, or category membership changed, so _index.md is untouched; the queue’s standing research-index item is served by this pass and the five ? for tom: questions remain unanswered [glm-5.3-flash]
  • Verification: both audits clean (the only remaining flags are the two documented exceptions, ante’s protocol-shape crate name and laya’s verbatim “Honest Limits” quote, plus free/open-source Pricing sections that correctly carry no price table); every URL cited by new or edited content fetched this run (antigravity.google/pricing, the deconstruction post, the MoLeMo-Lab/mojev repo, both HN threads via Algolia and web; trotsky1997/jevre confirmed 404 via the GitHub API and recorded as dead in the reference) [glm-5.3-flash]

2026-09-26 (owner-prompted, model access category seeded)
#

  • Owner request in chat (“Do we track opencode go? Let’s track all model access providers and their pricing”): created the model-access category for the layer that sells access to models (gateways, routers, vendor coding plans, flat subscriptions); eleven member notes written by four parallel sub-runs to citation standards, every cited URL fetched this run, with Pricing and Price history sections first-class in every note [glm-5.3-flash]
  • Membership: OpenCode Go ($10/month open-model token pack, the product the owner asked about, previously untracked), OpenCode Zen (the curated pay-per-use gateway), OpenRouter (passthrough plus a 5.5% credit fee, officially joining Stripe per the 2026-08-19 announcement fetched and cited this run, correcting the sub-run draft’s unconfirmed-talks framing), Requesty (flat 5% markup, EU residency), Synthetic ($30/month pack), NanoGPT ($12/month), Chutes ($10/$20 tiers on decentralized inference), Cerebras Code ($50/$200 speed plans, sold out as of today), GLM Coding Plan ($18/$80/$168), Kimi Code ($19-$199 membership ladder), MiniMax Coding Plan ($22/$55/$132) [glm-5.3-flash]
  • Boundary drawn: hosted access providers only; editor subscriptions stay in Surfaces, vendor API list prices stay in the model provider matrix, and self-hosted gateway software (LiteLLM, experiential, my-free-code) stays out unless the owner asks; the queue’s standing gateways question is marked answered in chat with this boundary named there [glm-5.3-flash]
  • model-access-feature-matrix: created in the same run with eleven columns sorted alphabetically by title and ten rows (kind, billing mechanics, cheapest paid entry, top tier, models reachable, quota form, the catch, frontier models, any-agent fit, price trajectory), every cell traced to its member note [glm-5.3-flash]
  • Category index page created with eleven alphabetical one-line entries and eleven dated Added bullets; section _index.md gained the matrix one-liner (between Memory and Model Benchmark) and the category line (between Memory and Orchestration), both alphabetical [glm-5.3-flash]
  • Cross-links: harnesses/opencode See-also gained the Zen and Go notes, trackers-and-leaderboards/openrouter-rankings gained the OpenRouter gateway profile, and model-selection-for-coding-tasks gained the new matrix, each with a dated Changes bullet and an updated bump to 2026-09-26 [glm-5.3-flash]
  • Defects fixed in-run before commit: opencode-zen’s price table dropped its four pre-baseline deprecation rows (catalog events predating tracking, not price changes), requesty lost a leftover verification meta line and a mangled link label, openrouter’s Status rewritten onto the official joining-Stripe announcement with the primary source added to its references [glm-5.3-flash]
  • Fetch notes: Reddit blocks direct fetches, so three critical threads (Zen’s 4x-pricing claim, Go’s not-general-access change, Synthetic’s stay-away thread) are cited with their retrieval path disclosed in the reference lines; z.ai and cerebras.ai price pages render client-side, so those notes’ dollars come from the dated third-party snapshots they name, recorded honestly [glm-5.3-flash]
  • Verification: 11 of 11 member directories match the category index and the matrix columns, case-insensitively sorted; all notes pass front-matter, tags, section-order, price-history, and internal-link checks (kimi-code’s single banned-term hit is inside a cited URL slug, a reviewer’s title, not our prose); section-wide audits stay clean; no em-dashes [glm-5.3-flash]

2026-09-26 (owner-directed sweep, verification stamps removed)
#

  • Removed the “Facts below verified as of” line from all 213 research notes across the 21 categories, per the owner’s instruction that the section drop per-article verification stamps; volatile facts keep their inline “as of” qualifiers, and AGENTS.md’s note skeleton no longer includes the line [glm-5.3-flash]
  • Three stamp lines carried caveats: orchestration/emdash and trackers-and-leaderboards/ai-release-tracker lost nothing (domain move already in Changes, 429/archived-copy already in Cautions and References); lmarena’s client-rendered-app caveat moved into its Changes [glm-5.3-flash]
  • AGENTS.md: skeleton line deleted, Status evidence line now says “as of a stated date”, and the Changes rule no longer exempts “verification-date bumps” (the practice no longer exists); trackers keep their first-paragraph as-of clause and matrices keep their no-preamble rule [glm-5.3-flash]

2026-09-26 (owner-prompted, Claude and ChatGPT subscriptions added to model access)
#

  • Owner question in chat (“Why aren’t you covering Claude code and openai subscriptions?”): the boundary drawn at seeding had left vendor consumer subscriptions to the provider matrix’s subscription row and the harness notes; accepted the owner’s reading that they are access products first, and added two notes to model-access, claude-plans and chatgpt-plans, written directly by this run with every cited URL fetched [glm-5.3-flash]
  • claude-plans: Free/Pro $20 ($17 annual)/Max 5x $100/Max 20x $200 from the fetched pricing page; the 2026 record from the fetched CC-licensed community guide (May 6 permanent 5-hour doubling, May 13 weekly +50% promotion through August 19, July 20 Fable tightened to 50% of weekly limits on Max and credits-only on Pro, July 25 Opus 5, the paused-and-never-shipped Agent SDK split, the April Pro-removal scare and the April 23 caching postmortem); Max plan support article fetched for limit mechanics and the discretionary-limits reservation [glm-5.3-flash]
  • chatgpt-plans: Free/Go $8/Plus $20/Pro $100 (5x) and $200 (20x) from the fetched developer pricing page with per-model message ranges and credit rates; Business $25/user and the April 2 Team replacement plus the credit-billing switch from the fetched critical source; chatgpt.com/pricing 403s automated fetchers, recorded in the references with verification routed through the developer docs [glm-5.3-flash]
  • Boundary updated: vendor consumer subscriptions that ship an agent are now in scope for model-access (the two that exist today); editor-bundled access (Cursor, Copilot Pro) stays in Surfaces pending owner direction, and that remaining edge is named in the log rather than the queue [glm-5.3-flash]
  • model-access-feature-matrix: extended from eleven to thirteen columns, re-sorted alphabetically (ChatGPT plans after Cerebras Code, Claude plans after Chutes), ten rows extended with the two subscriptions’ cells traced to their notes; reading prose and the choosing list updated for the fourth provider kind; Changes bullet appended [glm-5.3-flash]
  • Category index: scope sentence widened, two entries inserted in alphabetical position, two dated Added bullets appended [glm-5.3-flash]
  • Cross-links: harnesses/claude-code and harnesses/codex See-also gained their subscription notes, each with a dated Changes bullet; section _index.md matrix one-liner re-counted to thirteen [glm-5.3-flash]
  • Verification: 13 of 13 member directories match the category index and the matrix columns, case-insensitively sorted by title; both new notes pass front-matter, tags, section-order, price-history, and link checks; one unfetched URL (the business rate card) is linked from the fetched developer page and disclosed as not separately fetched; audits clean; no em-dashes [glm-5.3-flash]
  • On owner instruction (“Commit all the changes in the agents directory”) this run’s commit also carries a concurrent refresh session’s in-flight work found in the working tree: verification refreshes with dated Changes bullets across roughly eighty additional notes (repository scales, star and issue counts, updated bumps to 2026-09-26), including the claude-code and codex files this run also touched; that session’s own log entry records its authorship, and both stamp-removal passes (commit a59d511c and the uncommitted remainder) are the owner-directed convention this run conformed to rather than introduced [glm-5.3-flash]

2026-09-26 (owner-prompted, model access completeness pass)
#

  • Owner question in chat (“Any other subscription providers?”): audited the market against the category and found three credible untracked subscription providers, all inside the settled vendor-subscription boundary; three notes written by three parallel sub-runs, every cited URL fetched by the sub-runs, corrections applied to my leads where the fetched sources disagreed (the Qwen night rate is 40 percent of Credits, not 50 percent off) [glm-5.3-flash]
  • Membership: qwen-coding-plan (Alibaba Cloud Model Studio: Coding Plan Pro $50/month sold out, Lite discontinued 2026-03-20, the Token Plan restructure at $6-$68 Singapore with a China early-bird ladder, multi-model Qwen plus bundled Kimi, GLM, and MiniMax, automation-restriction suspensions as the critical source), supergrok (xAI: Free with limited Grok Build, $30/$100/$300, one shared weekly pool across chat, Build, and API per the fetched Warp documentation, Lite announced March in testing, Heavy grounded third-party), google-ai-plans (the Plus/Pro/Ultra ladder carrying Antigravity and Gemini CLI, Canadian storefront the only fetched dollar rendering, credit mechanics from the fetched Antigravity docs, the May 2026 bait-and-switch thread as critical source) [glm-5.3-flash]
  • Rejected with reasons, logged not noted: Poe, Mistral Le Chat, and Perplexity (consumer assistants without a coding-agent story), Together, Fireworks, Groq, DeepInfra (pay-per-token inference hosts with no subscription product), Bedrock, Vertex, and Azure (enterprise clouds), and the codingplan.io-style discount marketplaces (gray-market resellers, thin sourcing); GitHub Copilot and Cursor remain the owner’s pending boundary call as editor-bundled subscriptions tracked in Surfaces [glm-5.3-flash]
  • model-access-feature-matrix: extended from thirteen to sixteen columns, re-sorted alphabetically, ten rows extended; reading prose and the choosing list now cover four frontier vendors’ subscriptions; Changes bullet appended [glm-5.3-flash]
  • Category index: three entries inserted in alphabetical position, three dated Added bullets appended [glm-5.3-flash]
  • Cross-links: harnesses/grok-build gained the SuperGrok note, surfaces/antigravity gained Google AI plans, harnesses/qwen-code gained Qwen Coding Plan, each with a dated Changes bullet and an updated bump where stale [glm-5.3-flash]
  • Verification: 16 of 16 member directories match the category index and the matrix columns, case-insensitively sorted by title; the three new notes pass front-matter, tags, section-order, no-verification-stamp, and link checks; audits clean; no em-dashes [glm-5.3-flash]

2026-09-26 (owner-prompted, model access completeness pass)
#

  • Owner question in chat (“Any other subscription providers?”): audited the market against the category and found three credible untracked subscription providers, all inside the settled vendor-subscription boundary; three notes written by three parallel sub-runs with every cited URL fetched by the sub-runs, and my leads corrected where fetched sources disagreed (the Qwen night rate is 40 percent of Credits, not 50 percent off) [glm-5.3-flash]
  • Membership: qwen-coding-plan (Alibaba Cloud Model Studio: Coding Plan Pro $50/month sold out, Lite discontinued 2026-03-20, the Token Plan restructure at $6-$68 Singapore with a China early-bird ladder, Qwen plus bundled Kimi, GLM, and MiniMax, automation-restriction suspensions as the critical source), supergrok (xAI: Free with limited Grok Build, $30/$100/$300, one shared weekly pool across chat, Build, and API per the fetched Warp documentation, Lite announced March in testing, Heavy grounded third-party), google-ai-plans (the Plus/Pro/Ultra ladder carrying Antigravity and Gemini CLI, the Canadian storefront the only fetched dollar rendering, credit mechanics from the fetched Antigravity docs, the May 2026 bait-and-switch thread as critical source) [glm-5.3-flash]
  • Rejected with reasons, logged not noted: Poe, Mistral Le Chat, and Perplexity (consumer assistants without a coding-agent story), Together, Fireworks, Groq, and DeepInfra (pay-per-token inference hosts with no subscription product), Bedrock, Vertex, and Azure (enterprise clouds), and the codingplan.io-style discount marketplaces (gray-market resellers, thin sourcing); GitHub Copilot and Cursor remain the owner’s pending boundary call as editor-bundled subscriptions tracked in Surfaces [glm-5.3-flash]
  • model-access-feature-matrix: extended from thirteen to sixteen columns, re-sorted alphabetically after one insertion-order defect was caught and fixed in-run, ten rows extended; reading prose and the choosing list now cover four frontier vendors’ subscriptions; Changes bullet appended [glm-5.3-flash]
  • Category index: three entries inserted in alphabetical position, three dated Added bullets appended [glm-5.3-flash]
  • Cross-links: harnesses/grok-build gained the SuperGrok note, surfaces/antigravity gained Google AI plans, and harnesses/qwen-code gained Qwen Coding Plan, each with a dated Changes bullet and an updated bump where stale [glm-5.3-flash]
  • Verification: 16 of 16 member directories match the category index and the matrix columns, case-insensitively sorted by title; the three new notes pass front-matter, tags, section-order, no-verification-stamp, and link checks; audits clean; no em-dashes [glm-5.3-flash]

2026-09-27 (owner-prompted, GitHub-stars entrant scan added to the daily refresh)
#

  • Owner instruction in chat: add to AGENTS.md an instruction to fetch his GitHub stars and cross-check them against category membership as part of the daily entrant scan; the rule landed in Daily refresh step 3 (fetch every page of api.github.com/users/tomzx/starred?per_page=100, gh api when rate limits bite, a starred repository that fits a category resolves as an entrant candidate in the same run) [glm-5.3-flash]
  • First baseline pass executed in-run: 2,823 starred repositories, 71 already tracked or cited, 2,752 untracked, of which the agentic-and-LLM-relevant subset was enumerated for the next refresh run to resolve under the same-run rule [glm-5.3-flash]
  • Strongest surfaced candidates by category fit: harnesses (openinterpreter/openinterpreter 68k stars, coding agent for open models; ai4curation/metacoder), orchestration (FoundationAgents/MetaGPT, Significant-Gravitas/AutoGPT, microsoft/autogen, lobehub/lobehub), model-access (BerriAI/litellm, the self-hosted gateway whose status the earlier gateways rejection left open), evaluation-review (kolenaIO/autoarena), code-review (in-the-loop-labs/pair-review, pushed 2026-09-26), session-analytics (esc5221/claude-code-viewer, TomzxCode/llm-conversations-viewer, TomzxCode/tps-viewer, mirableio/chat-history), skills (vincentkoc/dotskills), memory and retrieval (lethain/library-mcp), context-engines and retrieval (docling-project/docling), assistant-runtimes (open-webui, Mintplex-Labs/anything-llm, zylon-ai/private-gpt, ollama), executions (cased/hubproxy, openclaw/octopool), surfaces and browser agents (browser-use/browser-use) [glm-5.3-flash]
  • Most owner stars are historical ML and PHP tooling outside this section’s scope, so the scan’s value is in the recent agentic subset; the next daily refresh resolves the enumerated candidates to the citation bar or logged rejections, one run, no carryover [glm-5.3-flash]

2026-09-27 (owner-prompted, stars-scan candidates resolved)
#

  • Owner instruction in chat (“Process the new candidates”): all 23 entrant candidates surfaced by the first stars pass resolved this run, 11 notes and 12 logged rejections, no carryover [glm-5.3-flash]
  • Notes added, by category: model-access gained litellm (BerriAI’s MIT self-hosted gateway, 59.7k stars, whose central caution is the March 2026 PyPI supply-chain compromise PYSEC-2026-2) and ollama (the 181.8k-star local runtime, where verification falsified my lead that pricing does not apply: Ollama now sells cloud tiers, Pro $20, Max $100, Team $500, so the note carries a seeded Price history table); harnesses gained openinterpreter (68.5k stars, now a Rust fork of Codex CLI with per-model harness emulation, the legacy Python line frozen at 0.4.3); retrieval gained docling (IBM-origin MIT document parser, 68k stars, 2.74 million downloads a month, LF AI and Data); orchestration gained metagpt (quiet since v0.8.2, org energy moved to OpenManus), autogpt (active platform, Polyform Shield plus MIT split, hosted $42.50/$272), autogen (maintenance mode, Microsoft Agent Framework the designated successor), and lobehub (active, community license, cloud $9.9-$39.9); assistant-runtimes gained open-webui (153.3k stars, custom license since v0.6.6 with branding clause), anything-llm (MIT, 66.5k stars, cloud $50/$99), and private-gpt (Apache-2.0, rebuilt as 1.0 in June 2026 from a private fork) [glm-5.3-flash]
  • Notes written by four parallel sub-runs to citation standards, every cited URL fetched by the sub-runs; notable lead corrections: Ollama’s free-only premise falsified, openinterpreter’s 1.0-line description corrected to a Codex CLI fork, the Qwen night rate already corrected last run [glm-5.3-flash]
  • Rejections with evidence: autoarena (dormant since 2024-12), pair-review (61 stars, days of traction, no coverage), claude-code-viewer (29 stars, single author), llm-conversations-viewer and tps-viewer (the owner’s own days-old tools, no third-party coverage), chat-history (stale since February), dotskills (two days old, curated-pack scope overlapping the Agent-Native column; re-review on sustained traction), library-mcp (dormant five months, thin documentation; re-review if a writeup ships), hubproxy (stale nine months), octopool (weeks-old ecosystem utility, thin docs; re-review on traction), metacoder (stale nine months, 22 stars) [glm-5.3-flash]
  • browser-use rejected on category fit, not quality: browser automation agents fit none of the 23 categories, so a ? for tom: question in the queue asks whether to open a browser-agents category; 116k stars would enter immediately if answered yes [glm-5.3-flash]
  • Category consequences applied in the same run: orchestration scope widened to include multi-agent frameworks (category index scope sentence, section index line), assistant-runtimes scope widened to local-first platforms, model-access scope gained the self-hosted software layer; five matrices extended (model-access 16 to 18 columns, orchestration 16 to 20, assistant-runtimes 9 to 12, retrieval 6 to 7, harnesses 28 to 29), all re-sorted alphabetically by member title after one insertion-order defect and the AX-after-AutoGen ordering were caught and fixed in-run; five category index pages and five section-index one-liners updated with dated Changes bullets [glm-5.3-flash]
  • Verification: audits clean (matrix columns case-insensitively sorted across all 23 categories, category indexes match membership, no broken links, no em-dashes, no banned terms in new content, no verification stamps) [glm-5.3-flash]

2026-09-27 (owner answer, browser automation stays out)
#

  • Owner answered the browser-use queue question in chat (“No browser use”): no browser-agents category will be opened, browser-use and browser automation agents remain out of the section’s scope, and the stars-scan rejection recorded this cycle is final rather than open to re-review; the resolution is recorded in queue.md beside the question [glm-5.3-flash]

2026-09-27 (defect, matrix separator rows)
#

  • Owner report in chat: the Assistant Runtimes Feature Matrix table rendered broken; root cause was this run’s column-extension script, which permuted header and data rows but left the | --- | separator row at its pre-extension width, so the delimiter row no longer matched the column count and GFM dropped the table [glm-5.3-flash]
  • Sweep across all 24 matrices found the same defect in every table the script had touched: assistant-runtimes (separator 10 under a 13-column header), harnesses (29 under 30), model-access (12 under 19, stale from the 2026-09-26 extension pass, meaning that table had rendered broken since commit bd3dd6ca), orchestration (17 under 21), and retrieval (7 under 8); all five separators regenerated to match their headers, and a whole-table width check confirms zero ragged tables across the section [glm-5.3-flash]
  • The verification gap that let this through twice is recorded in the project memory: matrix audits must compare separator and row widths against the header, not just column order and membership [glm-5.3-flash]

2026-09-27 (owner-prompted, queue questions answered)
#

  • Owner walked the six open queue questions in chat; all six now carry resolution lines in queue.md [glm-5.3-flash]
  • Q1, the untracked ADE tracker: judged finished on full read, moved from the repo root into agents/ at the owner’s instruction, and published as a section-root tracker (type and draft markers dropped, agent-curated tag added, corpus link depths fixed, the leaderboard framed as the mid-2026 snapshot with a pointer to the live Orchestration category, a Changes section added); all 17 external URLs fetched, one dead repository (saadnvd1/agent-os, 404) dropped from the mid-tier list and the watch sentence, two GitHub topic pages disclosed as 429 to fetchers; cross-linked from the section index, the Agentic Coding Tools Landscape essay, and the superset, emdash, paseo, jetbrains-air, and openchamber notes, each with a dated Changes bullet [glm-5.3-flash]
  • Q2, breakout community skills: re-verified, both cleared the bar and joined Skills; sepia (Nanako0129, 2,867 stars, research-grounded de-AI writing skill, zero HN footprint recorded as the missing-coverage signal, Hysen Labs review as critical source) and headcount (Chris Brock, 1,680 stars, 16-department org chart whose reviewer-class departments can block writes, two critical third-party reviews); Skills matrix extended six to eight columns and re-sorted [glm-5.3-flash]
  • Q3, self-hosted gateways: experiential cleared the bar and joined model-access (7,088 stars, YC-backed, Apache-2.0, zero markup with hosted credit tiers, and the unresolved founder-versus-README telemetry contradiction as its central caution; the HN thread dated 2026-08-27, not the entrant data’s 2026-09-21, and today’s star count recorded instead of the unverifiable 820 snapshot); my-free-code rejected for good, zero independent coverage anywhere, an unrendered LLM citation artifact in its README, unreachable at the citation bar; model-access matrix extended eighteen to nineteen columns [glm-5.3-flash]
  • Q4, the b32de458 sweep, and the two self-resolved questions (link conventions, price history rule) closed with resolution lines; both rules verified present in the current rules file [glm-5.3-flash]
  • Category consequences: skills index and matrix gained two members, model-access one, with scope sentences already covering them; section index one-liners re-counted (Skills eight members, model-access nineteen) [glm-5.3-flash]
  • Verification: audits clean, and this run’s matrix edits passed a whole-table width check (every row and separator matches its header) adopted after yesterday’s separator defect; three new notes pass front-matter, tags, section-order, price-history, and link checks [glm-5.3-flash]
  • Post-commit fixes caught by the audit in the same session: Experiential was mis-alphabetized in the model-access matrix and index (it sorts under E, between Claude plans and GLM), Sepia sat one slot late in the Skills matrix, and three of the moved ADE tracker’s cross-links carried wrong depths (section-root corpus links are two levels, not three; the essay sibling needs one); all fixed before push [glm-5.3-flash]
  • Attribution note for commit 450a634f: it also swept in agents/harnesses/mimo-code/index.md, a complete MiMo Code harness note written by a concurrent session (llm=deepseek-v4.1-flash) that had not yet committed or logged; that session’s own entry records its authorship, and this run’s section-index one-liner was corrected to thirty harnesses to match the membership its note created [glm-5.3-flash]

2026-09-27 (owner-prompted, MiMo Code harness note)
#

  • Owner instruction in chat (“Add mimo code to agents/”): wrote agents/harnesses/mimo-code/index.md, a research note on Xiaomi’s MIT OpenCode fork built around checkpointed memory, a goal-verifying judge, deterministic workflows, and the Xiaomi Token Plan, with every cited URL fetched this run and the open Token Plan credit-multiplier issue as the critical pricing source [deepseek-v4.1-flash]
  • The note was swept into commit 450a634f by the concurrent daily run before it was logged, and commit 3f8d71ab corrected the section index to thirty harnesses; this entry records my authorship and the remaining integration [deepseek-v4.1-flash]
  • harness-feature-matrix: added the MiMo Code column (twenty-nine to thirty members), sorted after Kimi Code, all twelve cells traced to the new note; also backfilled the missing Open Interpreter Changes bullet and corrected the intro count from twenty-eight to thirty, both existing defects surfaced while extending the table; added MiMo Code to the auditable-client and local-models lists and a long-horizon-memory reading paragraph, and appended four references [deepseek-v4.1-flash]
  • Note wording fix before commit: “honest case” to “defensible case” and “tree-shaped” to “tree-structured” to keep the banned-terms rule [deepseek-v4.1-flash]
  • Verification: matrix header, separator, and all twelve data rows at 31 cells each; all internal link targets exist on disk; front matter parses; no em-dashes or banned terms in the new note; every referenced URL fetched during this run [deepseek-v4.1-flash]

2026-09-27 (owner-prompted, queue cleanup)
#

  • Owner instruction in chat: remove all the answered/resolved entries from queue.md as they are not useful to keep; the queue now holds only the standing research-index item and the Done list, with every answered question and its resolution line removed; the decisions themselves remain recorded in this log and in the rules file where applicable [glm-5.3-flash]

2026-09-27 (daily refresh, full parallel re-verification + model price moves)
#

  • Full parallel refresh under the no-stalest-ranking rule: 12 sub-runs re-verified all 233 research notes across the 23 categories, the seven section-root articles, and all 25 matrix tables against live sources (GitHub via gh api, npm/PyPI/crates.io, HN Algolia, vendor and pricing pages); notes with updated < 2026-09-26 got the full pass, notes already touched 2026-09-26/27 by the day’s owner-prompted runs got a light link-and-status check [glm-5.3-flash]
  • Coordinator pre-refresh audit fixed four banned-term slips (model-access matrix “agent-shaped”, boris-cherny “shaped”, the people matrix “enterprise-shaped”, thorsten-ball “shapes”, all reworded meaning-unchanged with Changes bullets and updated bumps), renamed the executions matrix column and listing to the member title GitHub Copilot automations, and renamed the surface matrix’s Antigravity column to Google Antigravity with the columns and the category listing re-sorted by member title [glm-5.3-flash]
  • Harnesses: mimo-code corrected, GitHub’s latest-release endpoint returns v0.1.14 (2026-09-23) after v0.1.15, status and references fixed, llm=glm-5.3-flash added; aider 49.2k stars with the stall re-confirmed, claude-code 148.3k stars and 13.3k issues+PRs, cline 5.46M installs, deepseek-harness 237.1k still prerelease-only, grok-build alpha v1.0.42, opencode 210.3k, pi 109.6k, zcode 6.9k, openinterpreter 68.5k; no other material move [glm-5.3-flash]
  • Surfaces: cursor, jetbrains (junie pricing visible again to static fetch at $8.33/$25), trae, void, vscode-copilot, and windsurf as-of refreshes with no fact moves; openchamber’s stray 09-27 Changes bullet found inside References and moved into ## Changes; antigravity’s updated field synced to its recorded 09-26 revision [glm-5.3-flash]
  • Orchestration: cmux’s Max tier re-packaged onto a shared 64 GB/16 vCPU pool across up to 50 VMs (Pricing rewritten, Price history row); vibe-kanban re-confirmed orphaned with no commit after 09-19; the orchestration matrix’s stale sixteen-member prose corrected to twenty and the framework-cluster reading sentence added [glm-5.3-flash]
  • Protocols: agent-host-protocol 368 stars, 252,521 crates, and the npm window rebased at 157,131; ag-ui at 7.78M core and 5.10M client monthly downloads; acp’s agent list re-counted at exactly 40; mcp spec unchanged at 2026-07-28; protocols matrix prose refreshed [glm-5.3-flash]
  • Task management: ordewell shipped v0.5.0 through v0.5.4 across 09-25/26 (162 stars, its 8-to-42 issue jump inspected and recorded as maintainer roadmap tickets); beads caution corrected 1,224 to 1,270; task-management matrix status row refreshed [glm-5.3-flash]
  • Control planes: paperclip 87.9k stars with the thesis line updated, tinyagi’s stall re-confirmed, matrix status cells moved [glm-5.3-flash]
  • Model access and session analytics: openrouter’s dated price-history source now 403s fetchers (annotation updated, the fee state re-verified unchanged on the official FAQ); agentsview’s schema caution revised on its third flip (the docs are back to version 5); ctx shipped three releases; the model-access matrix’s choosing bullet reworded off a banned term [glm-5.3-flash]
  • Assistant runtimes: eigent v1.0.5; openwork’s pricing churned a ninth time with a Price history row (Solo and the grandfather clause gone, Team $10/seat unlimited, Enterprise $20/user annual); github-agentic-workflows v0.89.21 is now the newest stable-marked release, retiring the v0.88.7-last-stable record, v0.89.22 prerelease today; the assistant-runtimes matrix’s stale nine-member intro corrected to its twelve columns plus cell refreshes [glm-5.3-flash]
  • Memory, retrieval, context engines: claude-mem v13.28.0, engrim’s release pause confirmed, knowhere 3,528 stars with pricing re-confirmed live; graphify v0.9.69 at 121.7k stars, rtk’s rc train at 468; the retrieval matrix gained the Knowhere cell refresh plus two stale-intro repairs left by the Docling addition [glm-5.3-flash]
  • Code review and evaluation: open-code-review 41.7k stars (about 80 percent in fourteen days); deepeval 4.2.5 and 4.2.6; ellipsis’s free-tier badge flipped back to ChatGPT-or-Claude with the flip-flop history recorded; the code-review matrix’s OCR maturity cell moved [glm-5.3-flash]
  • Skills: agent-native npm 0.3.7 at 18.0k weekly downloads; skills-sh’s leaderboard reshuffled with azure-validate still Snyk-Critical at row 44; the skills matrix’s prose figure updated [glm-5.3-flash]
  • Sandboxing: openshell graduated to stable v0.1.0 and v0.1.1, the alpha badge dropped and two dead docs references replaced after NVIDIA’s restructure; clawk’s quiet window now 45 days; flue’s 2.2.0-next prerelease wave; the sandboxing matrix’s OpenShell maturity cell moved [glm-5.3-flash]
  • Hybrid execution: jevbench shipped v1.4.2.1 today, Plumb-4B the new #1 at 65.84 with Jev third, and jev, semif, and laya rank lines updated with the matrix cell and reading rewritten [glm-5.3-flash]
  • Spec-driven and software factory: spec-kit v1.0.12, gsd’s successor v1.15.0 at 9.9k stars, har v1.15.0, ouroboros v0.54.6, machinist’s zero-coverage claim re-confirmed via an Algolia zero-hit search; adoption and stars cells moved in both matrices [glm-5.3-flash]
  • People and trackers: caleb-writes-code 118K subscribers, nathan-lambert 84K, deeplearning-ai’s Batch at issue 372 with a dead reference replaced by the live Ask HN thread, artificial-analysis 673 models, epoch-ai’s September stamp; six other people notes cadence-refreshed [glm-5.3-flash]
  • Essays: model-selection-for-coding-tasks caught the real moves, a Claude Opus 5.5 column added ($4/$20, 1M flat, cache at one twentieth of input), gpt-5.6-terra and gpt-5.6-luna delisted from OpenAI’s page and flagged delisted, GLM-5.3-FlashX’s cached read corrected to $0.075, the Kimi K3 host floor re-measured at $0.88/$4.90, and three transposed context-behavior cells fixed against primary sources; the landscape’s as-of bumped to 2026-09-27 folding Orca about 79k, Kimi Code about 19.1k, and Cline about 5.5M installs, its oversized verification chain collapsed to one clause; the provider matrix moved the Opus, Kimi, and terra cells; the ADE tracker’s two malformed See-also items fixed; context-management-patterns’ intro clause re-dated; the-tells-are-structural untouched [glm-5.3-flash]
  • Concurrent-session defect repaired: six notes (openchamber, emdash, jetbrains-air, paseo, superset, agentic-coding-tools-landscape) carried 09-27 Changes bullets appended into References by an earlier run today, all moved into ## Changes in date order [glm-5.3-flash]
  • Stars cross-check: all 2,823 starred repositories re-fetched, the count matching the morning baseline exactly with no new stars; the residue re-confirmed as the out-of-scope historical ML and PaaS tail [glm-5.3-flash]
  • Entrant resolution, all rejected with evidence, no notes written: drawgent (143 points for a coding agent on a live Excalidraw canvas, tangled.org-only hosting with no GitHub mirror so license and stars are unverifiable and two citable sources miss the bar, re-review on a primary mirror), golive-skill (986 stars in four days, MIT, zero independent coverage and a single maintainer, below the bar sepia and headcount set, re-review on coverage), magpie (yetone’s 1,122-star menu-bar model router, zero HN footprint, below the model-access software layer’s evidence bar, re-review on coverage), jevmem (61 points, 87 stars, single maintainer, self-reported benchmarks, a third-party API in the loop), jev-code-reviewer (46 points, 77 stars, thin footprint against the code-review bar), plus the weak tail (KISS at 8 stars, Eikos 22, openchatx-mcp 135, mu 239 with no docs, Nom Army, quicksilver, quiron, the single-purpose content skills, Tenjin, token-bar, AgentRun’s DSL at Foremerge-level traction, dolphinbench a benchmark rather than a tool) [glm-5.3-flash]
  • _index.md: all essay and matrix one-liners re-dated to 2026-09-27 and the Spec Kit line moved to v1.0.12; category index pages refreshed in place where revised notes contradicted their summaries; no membership change anywhere, so no category-page Changes bullets [glm-5.3-flash]
  • No self-directed essay this run: the full re-verification, the Opus 5.5 addition with the terra and luna delisting, and the coordinator defect repairs filled the day at quality [glm-5.3-flash]
  • Verification: 263 article files plus the control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets, all 25 matrix tables (23 category matrices plus the provider and benchmark matrices) match their categories’ membership with case-insensitively sorted headers, uniform rows, and separators at header width, no em-dashes, banned terms only at the documented exceptions (kimi-code’s cited URL slug, emdash’s append-only history, laya’s quoted section title, ante’s crate name), one sentence per line, and every URL cited by new or edited content fetched 200 this run or its absence recorded honestly (chatgpt.com/pricing and arstechnica 403 bot walls kept with dates, usagepricing.com’s new 403 annotated, erdosproblems.com verified via webfetch after curl 403) [glm-5.3-flash]

2026-09-27 (owner-prompted, awesome-multi-agent-orchestrators triage)
#

  • Owner instruction in chat (“Look at https://github.com/Agent-Analytics/awesome-multi-agent-orchestrators and identify relevant projects to add under agents/”): triaged the whole curated directory against the section’s 233 notes, deduplicated the seven entries already covered (Hermes Agent, NanoClaw, OpenClaw, Paperclip, Superset, Vibe Kanban, Gas Town), and the owner approved the four in-scope gaps below, the broad multi-agent framework tier deferred [deepseek-v4.1-flash]
  • Orchestration: added 11 notes (Orca, Sandcastle, Lanes, Crewplane, Helmor, Open Swarm, LoopTroop, GraphCode, The Perfect Orchestrator, Multica, oh-my-codex), the category 20 to 31 members; Orca had been cited at about 79k stars in the landscape tracker with no note of its own; the matrix was extended to thirty-one sorted columns with a new family reading block, and the category index list and Changes updated [deepseek-v4.1-flash]
  • Control planes: added 6 notes (Microsoft Agent Governance Toolkit, Veto, SettleBridge, Okto Pulse, SIDJUA, Code Atelier Governance SDK), the category 2 to 8 members; two new matrix rows (Enforcement point, Audit and evidence) and the false “second column is the stall record” prose replaced; Okto Pulse is an SDLC gate layer rather than an orchestration plane [deepseek-v4.1-flash]
  • Session analytics: added 3 notes (AgentTrace, ClawTrace, Agent Analytics), the category 4 to 7 members; Agent Analytics sits at the category’s edge (web product analytics agents query rather than coding-agent session observation) and that caveat is recorded in its note [deepseek-v4.1-flash]
  • Automated research: added Agon, the category 7 to 8 members; the note was written by this session after its dedicated subagent was cancelled, with every cited URL fetched here, including the arXiv paper and project page, and the zero-hit HN Algolia result recorded as the absent community footprint [deepseek-v4.1-flash]
  • Rules (owner-directed): added the Awesome Multi-Agent Orchestrators directory and openorchestrators.org to the standing scan in AGENTS.md step 3, alongside the GitHub-stars pass [deepseek-v4.1-flash]
  • Rejected with evidence: Nomad Inno and Agentix Labs (implementation services, no tool), Crewlet, Augment Code Intent, and AgentGrid (closed products), Databricks Unity AI Gateway (commercial platform feature); the broad framework and company-OS tier (CrewAI, Agno, Dify, Flowise, Mastra, Sim, SwarmClaw, NarraNexus, Squad, Culture, Bunkhouse, Wave, The Rusty Claw, Okto Nexus, Cabinet, Agent Office Suite, Companies.sh) was deferred rather than added in one pass [deepseek-v4.1-flash]
  • Section index: the four matrix one-liners and four category one-liners updated in place; no membership changed outside these four categories, so no other category page Changes bullets were touched [deepseek-v4.1-flash]
  • Verification: all four changed matrices uniform (orchestration 32 cells, control-planes 9, automated-research 9, session-analytics 8), front matter parses for every agents article with no type field, 0 broken internal link targets, no em-dashes or banned terms in the new content, and every URL cited by the new notes fetched this run or its absence recorded [deepseek-v4.1-flash]

2026-09-29 (daily refresh, full parallel re-verification + orchestrator tier resolved)
#

  • Full parallel refresh under the no-stalest-ranking rule: 11 sub-runs across 3 waves re-verified all 264 research notes across the 23 categories, the seven section-root articles, and all 25 matrix tables against live sources (GitHub via gh api, releases, npm/PyPI/crates.io, HN Algolia, vendor and pricing pages); notes with material revisions got updated set plus Changes bullets, bump-only notes left byte-identical [glm-5.3-flash]
  • The 2026-09-27 deferred framework tier resolved this run with eight notes, every cited URL fetched this run: crewai (59,158 stars, MIT, $18M raised), agno (42,377, Apache-2.0, ex-Phidata), dify (157,464, modified Apache-2.0, $590/$1590 per workspace per year cloud), flowise (the genre’s death record, Workday acquisition 2025-08, archived 2026-08 per its own shutdown note), mastra (6.1M npm downloads a month, $35M raised, one survived npm supply-chain attack), sim (29,749, YC X25, Apache-2.0, Free/$25/$100), atlas (the owner-star pacifio/atlas, Rust source control for agents, 8,320 stars in 4.5 months), and squad (Brady Gaster’s MIT Copilot-CLI agent teams, 3,237 stars); orchestration-feature-matrix extended 31 to 39 sorted columns in the same run and the category page lists all eight with Added bullets [glm-5.3-flash]
  • Framework-tier rejections with evidence, all resolved in-run: SwarmClaw (686 stars, stale since June, two self-posted 4-5 point threads), NarraNexus (87), Culture (113), Bunkhouse (4), Wave (3), Agent Office Suite (140 plus a category misfit), handoff (91, no license, a delegation bridge not a framework), zachealy1/orchestrator (12), Companies.sh (no verifiable repository); Okto Nexus (60 stars, Elastic 2.0, PyPI 0.1.10) folded into the okto-pulse note as the vendor’s sibling rather than noted separately [glm-5.3-flash]
  • Two further entrants accepted: jeff in Hybrid execution (firelex/jeff, 471-point day-one Show HN, MIT code and Apache-2.0 weights, frankest self-run benchmark table, AutoJev lineage; matrix 13 to 14 columns same-run) and drop in Sandboxing (wrr/drop, 193-point Show HN, Apache-2.0 Go rootless user/mount/PID/IPC/cgroup/network namespaces with optional gVisor, 331 stars; matrix 8 to 9 columns same-run, category page updated) [glm-5.3-flash]
  • Harnesses material: ante v0.2.6 (2026-09-28), deepseek-harness 239,125 stars with dsh-v0.2.0-rc.1 the first 0.2.0-series candidate, grok-build stable v1.0.44 with the alpha at v1.0.45, jcode v0.89.0 (voice input, agent applets), juggler v0.7.1/v0.7.2, claude-code about 148.5k stars with issues+PRs up to about 13.6k, unreal-agent 2,018 stars, zcode 7,097; aider’s stall re-confirmed, junie’s pricing page still 404 with dated tiers standing [glm-5.3-flash]
  • Surfaces and orchestration material: openchamber v2.0.3/v2.0.4 (10,897 stars), jetbrains recorded the September 22 Air announcement (Air in IDEs, Air Teams, Air Governance), paperclip 93,605 stars (up more than 5,500 in two days) with 6,019 open issues, multica v0.6.0 at 51.6k stars, orca v1.4.216 at 81k, conductor 0.87.6, emdash v1.2.7, graphcode’s first stable v0.1.76, lanes v0.49.5, paseo v0.10.0/v0.10.1 with the Hub price display moved euro 15 to $15 (Price history row), superset desktop v1.31.0, worktrunk v0.80.0 with breaking squash-hook changes, open-code-review 42,421 stars with v1.12.10, ordewell v0.5.5 [glm-5.3-flash]
  • Assistant runtimes and session material: openclaw shipped v2026.8.33, its first gateway-only extended-stable release (390,755 stars), openwork’s tenth pricing churn (SSO into Team, two add-ons, Price history row), picoclaw.io’s TLS certificate renewed through 2027-04-14 ending the expiry break, hermes about 250k stars, ctx v2.1.2 with paged blame indexing; agentsview’s schema caution verified standing [glm-5.3-flash]
  • Elsewhere material: semble v0.6.1 (missed by earlier runs), graphify v0.9.71 at 122,219 stars, graft 9,381, repomix window re-based, agent-native npm 0.3.10, skillopt 17,824, github-agentic-workflows v0.90.0 prerelease (v0.89.21 still newest stable), plannotator v0.27.22, flue graduated 2.2.x stable, openshell v0.1.2, har v1.15.1, ouroboros v0.55.0, jevbench v1.4.2.2 re-scored the board (Imajev-4B new first at 67.37, Jev fourth), laya about 28,000 stars, kev 7,732, openai-for-science recorded the forced-versus-unforced formulation fight (Scientific American, 2026-09-21) with the three-mathematician non-extension proof, simon-willison’s “2026 in LLMs (so far)” keynote writeup recorded, geoffrey-huntley’s eighteen-month recap, shreya-shankar’s Quail announcement, artificial-analysis 679 models, epoch-ai September 29 stamp, llm-stats 399 canonical models [glm-5.3-flash]
  • Essays: model-selection-for-coding-tasks added a Claude Sonnet 5.5 column ($2/$10, cache about one tenth, 1M flat, released 2026-09-28), reverted the 09-27 terra and luna delisting after three fetches showed both relisted at launch prices, re-measured the Kimi K3 OpenRouter host floor to $0.40/$10.00 (Relace), and canonicalized the Anthropic reference to platform.claude.com; zero vendor list-price moves otherwise across all eight providers; the provider matrix synced (Anthropic workhorse to Sonnet 5.5, OpenAI terra parenthetical dropped, Moonshot floor cell); landscape folded in Claude Code 148.5k, OpenCode 210.7k, and a DeepSeek Harness line, as-of 2026-09-29; context-management-patterns and the ADE tracker lost four dead or moved URLs to canonicalization; model-benchmark-matrix re-homed two redirecting boards; the-tells-are-structural byte-identical [glm-5.3-flash]
  • Entrant rejections, all resolved in-run across the wave scans: harnesses (Cf at 157 points is a Cloudflare management CLI, Jevgrep, Vespper, RRSI, swe-mux, Jauvex), runtimes/memory/session (noclick, Recalld, IngotDB, ContextMemory), surfaces/review/tasks/control (Gitar, Gito, kastra, Blue standing), context/retrieval/skills/executions (the Vespper DOCX MCP, chess-postmortem-skills and superCMO under the single-skill-pack exclusion, Corral, ParkourNote, KnowNote, Yengi), evaluation/sandbox/factory/spec (Brig, Cordum Edge, Kern Sandbox, Augur, Canary, AgentRun, Yass), protocols/hybrid/research (Nodd, Gargi Reflex, onesie, Decide, the Durable Actor Session Protocol, the Poincare Lean formalization, the Anthropic wet-lab story on a single secondary source, the AutoJev standalone cited inside jeff instead), trackers (TinyAIArena, a client-rendered game demo); the stars cross-check re-fetched all 2,823 starred repositories with no new stars since the processed baseline [glm-5.3-flash]
  • Coordinator repairs in the integration pass: multica’s two pre-existing banned-term compounds reworded with a Changes bullet; the protocols page list brought to the matrix’s case-insensitive title order (the 2026-09-25 page arrangement was the outlier, the matrix’s A2A, ACP, AG-UI, Agent Host Protocol, AGENTS.md, MCP order is the section-wide sort) and the session-analytics page list reordered to match its matrix (agents-observe and agentsview ahead of AgentTrace); the model-access matrix’s stale sixteen-member intro corrected to nineteen with the three 2026-09-27 column additions recorded [glm-5.3-flash]
  • _index.md: the orchestration, hybrid-execution, and sandboxing matrix one-liners recast for the new members, every dated one-liner re-dated to 2026-09-29; no new categories or essays, so the section lists changed only in place [glm-5.3-flash]
  • Queue: the standing research-index item is served by this refresh; no new ? for tom: questions; no self-directed essay this run, the framework-tier resolution plus three entrant notes filled the day at quality [glm-5.3-flash]
  • Verification: 295 article files plus 23 category pages and the control files pass front-matter parsing with no type field and mandatory tags, 0 broken internal link targets across the section, every matrix table uniform (header, separator, and rows at equal widths) with case-insensitively sorted columns, category index pages match directory membership, no em-dashes, banned terms only at the documented exceptions (ante’s crate name, laya’s quoted section title, kimi-code’s cited URL slug, the quoted rewording bullets in emdash and multica Changes history), one sentence per line, and every URL cited by new or edited content fetched 200 this run or its absence recorded honestly (pypistats 429 with retries succeeding, n8n pricing JS-rendered with dated claims kept, erdosproblems 403 to curl, claude-mem.ai client-rendered, chatgpt.com/pricing 403 verified via the developer docs) [glm-5.3-flash]

2026-09-30 (re-verification, orchestration and people-and-publications)
#

  • Orchestration re-verification: recorded AutoGPT beta v0.8.2 (2026-09-30, AutoPilot action gating modes) with 187,620 stars and a 2026-09-30 push, Conductor 0.89.0 (Sign in with ChatGPT) and 0.89.1 (GPT-6.1 Sol) both 2026-09-29, GraphCode v0.1.77 (2026-09-29) at 133 stars, Open Swarm v1.8.0-exp.4 (2026-09-29) at 823 stars (adding llm=glm-5.3-flash to its tags), Orca v1.4.217 (2026-09-29) at 81.9k stars, Sim v0.9.6 (2026-09-29) at 29.8k stars, and Superset desktop v1.32.0 (2026-09-29, CLI setup for coworker cloud environments) [glm-5.3-flash]
  • Orchestration matrix: moved the AutoGPT status cell to beta v0.8.2, the GraphCode cell to v0.1.77 with 133 stars, the Orca cell to v1.4.217 with 81.9k stars, the Paseo cell to v0.10.2 (2026-09-30), and the Sim cell to v0.9.6 with 29.8k stars, and rewrote the Vibe Kanban status cell and reading prose now that its v0.1.45 prerelease release is no longer listed on GitHub (the tag remains, npm still 0.1.44) [glm-5.3-flash]
  • Vibe Kanban: re-verified that the v0.1.45 prerelease release is no longer listed on GitHub (the release URL returns 404 and the release list stops at 0.1.44, while the timestamped tag remains), npm still serves 0.1.44 as latest, and the tracker moved to 544; the category index Vibe Kanban one-liner changed from “only a v0.1.45 prerelease” to “only a v0.1.45 tag” [glm-5.3-flash]
  • People and publications re-verification: Armin Ronacher’s newest essay “Deser: Rethinking Rust Serialization” (2026-09-29) brings the July-to-September count to eleven, Caleb Writes Code’s latest upload moved to 2026-09-30 (“Inference Engines explained in 10min..”), IndyDevDan’s Monday streak extended to 16 consecutive weeks (2026-06-15 to 2026-09-28, latest “10 Levels of Jev For Agentic Engineers”), Latent Space now leads with “Claude Code’s Next Era” featuring Anthropic’s Thariq Shihipar (2026-09-29) with AINews leading on OpenAI DevDay 2026, and Simon Willison’s daily posting runs through 2026-09-29 led by his DevDay 2026 live blog [glm-5.3-flash]
  • People and publications matrix: refreshed the Ronacher cadence cell to eleven essays between July 4 and September 29, 2026 and the IndyDevDan cadence cell to 16 consecutive Mondays; no membership change, columns stay at thirty [glm-5.3-flash]

2026-10-01 (owner-prompted, banned-term sweep)
#

  • Owner-directed banned-term sweep of this year’s articles (banned-terms.txt: shape, honest, load bearing, real, substrate, posture). In the section, reworded the “execution substrate” metaphor (executions-feature-matrix, github-agentic-workflows, n8n, tinyagi) and stretched “posture” uses (coderabbit, graphify, graft, context-engines-feature-matrix, ante, mimo-code, openinterpreter, grok-build, knowhere, openspec, windsurf, session-analytics-feature-matrix, clawtrace, paseo); security/privacy/compliance “posture” kept as the most appropriate term, as were the pre-existing exceptions (protocol-shape, “Honest Limits”, the kimi-code URL slug, the emdash/multica rewording bullets). Each touched article got updated: 2026-10-01 and a Changes bullet. The main corpus was swept in the same session (sixteen “real” section headings reworded); that change is outside this section’s log. [deepseek-v4.1-flash]

2026-10-02 (daily refresh, full parallel deep-refresh across all 23 categories plus the seven section-root articles)
#

  • Run structure: 12 parallel category workers re-verified every research note, matrix, and category page against live sources (GitHub API, npm/PyPI/crates.io, HN Algolia, vendor and pricing pages), with four provider-rate-limit casualties retried in smaller batches and one partial (ctx) completed by its retry; no scope was left unrefreshed and no candidate left pending [glm-5.3-flash]
  • Stars scan: all 2,824 starred repositories fetched with starred_at (up exactly one from the 2,823 baseline), the one new star pfnet-research/meta-fuse-csi-plugin (a Kubernetes CSI storage plugin) rejected as out-of-scope infrastructure; the AMO directory re-crossed with no uncovered entries; the owner’s mid-run checkpoint commit (a0c84843, newest-first starred_at scan bound by agents/starred-checkpoint.txt) observed and this run’s full fetch satisfies it, the checkpoint already holding the newest timestamp seen (2026-09-29T21:36:25Z) [glm-5.3-flash]
  • Section roots: model-selection-for-coding-tasks added gpt-6.1-sol (2026-09-29, $2/$10 with cache halved to $0.10) and re-measured the Kimi K3 floor to $0.38/$10.00 and k2.7-code to $0.67/$3.35, the provider matrix synced those cells, zero other vendor list-price moves across all seven providers; the landscape folded Orca about 83.2k and DeepSeek Harness about 241.7k with as-of 2026-10-02; the ADE tracker, context-management-patterns, model-benchmark-matrix, and the-tells-are-structural verified with no changes [glm-5.3-flash]
  • Harnesses: pi’s central “no MCP by design” claim corrected (v0.99.0 shipped codemode, tool search, and MCP as built-in extensions; v1.0.0 followed 2026-10-01), the matrix cell moved from blocked to built-in; kimi-code recorded Moonshot’s archive of the original kimi-cli repo and a K3 drop to $0.4255/$10.00 with a Price history row; amp folded the Opus 5.5 medium move, ante v0.2.7, deepseek-harness rc.2 at 242k stars, fx v0.0.12, grok-build 1.0.46/1.0.47, jcode v0.90.0, juggler v0.7.3, kilo-code v7.8.3, openinterpreter 0.0.55 [glm-5.3-flash]
  • Model access and assistant runtimes: ChatGPT Pro gained a $500/month third tier with Astra Ultrafast access (Price history row), Experiential reworked to a $20-$1,999+ credit ladder with new Max and Ultra tiers (Price history row), openwork churned pricing an eleventh time (Enterprise back to custom annual, Price history row), OpenClaw moved to v2026.9.7 current with a second extended-stable v2026.8.34, AnythingLLM v1.17.0; every other plan page re-verified unchanged [glm-5.3-flash]
  • Orchestration: agno v3.1.0 (breaking fs-table re-key), conductor 0.89.3, looptroop v0.6.0 (test suite 3,495 to 8,099), mastra core 1.74.0 with downloads 6.1M to 7.1M/month, multica v0.6.1, oh-my-codex v0.21.7, open-swarm v1.8.0-exp.5, orca v1.4.218, sim v0.9.9, superset v1.33.0, and the happy-coder star-ranking claim corrected (oh-my-codex is a Codex workflow layer, not a session multiplexer); sixteen matrix cells moved [glm-5.3-flash]
  • Sandboxing and executions: OpenShell folded into NVIDIA’s branded Open Agent Safety Platform (with Sentry and BlueField-4) explaining its about-4,300-star three-day surge to 14,010, drop v0.3.0, agent-sandbox v1.0.5, opensandbox’s 1.1.1-rc.1 wheel-fix prerelease recorded with the broken 1.1.0 wheel as a Caution, aigate moved to dormancy evidence (no push since 2026-08-04), github-agentic-workflows train to v0.90.1 prerelease with v0.89.21 still newest stable [glm-5.3-flash]
  • Task management: ordewell v0.5.6 (seventh v0.5.x in five days, 183 stars) and beads v1.3.1 stable (27,574 stars), both matrices’ status cells moved [glm-5.3-flash]
  • Memory, retrieval, context engines: LangChain pricing changed (the $1.50 LCU line item is gone from langchain.com/pricing, metering consolidated in LSUs at $1.00, Price history row appended), graft 0.21.1, graphify v0.9.73 at 123,085 stars, knowhere v1.2.20, cognee v1.6.2 (explicit embedding failures, Slack import, skills moved to .agents/skills); claude-mem 95.1k stars with issues down from 312 to 108 [glm-5.3-flash]
  • Code review and evaluation: ellipsis’s platform fee doubled from 10% to 20% of token cost (Price history row, free tier reworded “FREE FOR INDIVIDUALS”), open-code-review 43,189 stars, langfuse v4.49.0 and phoenix v20.19 cells moved, deepeval 4.2.7 [glm-5.3-flash]
  • Automated research: AlphaProof’s bottom-line claim corrected (Gemini 3 Deep Think opened the first Gemini API early-access path, so “nothing API-accessible” is no longer true), Harmonic Aristotle corrected on funding and leadership (Series C $120M at $1.45B led by Ribbit Capital) with the Mathematician Sponsorships program canonically cited, openai-deep-research’s four Wikipedia URLs canonicalized to the renamed ChatGPT Deep Research article, one matrix cell moved [glm-5.3-flash]
  • Control planes and skills: Veto moved to dormant (no push in 106 days) and Code Atelier Governance SDK to dormant (71 days), Okto Neuron folded into the Okto Pulse note as the vendor’s third product, missing llm=glm-5.3-flash tags added across the control-planes seed and six deepseek-tagged orchestration notes; AG-UI downloads 13.1M/month, AHP crate plus npm about 597k [glm-5.3-flash]
  • Coordinator repair: the protocols feature matrix had lagged its notes (AG-UI still at 12.3M as of 2026-09-29, AHP at 433k and 375 stars), prose and references refreshed to today’s figures with a Changes bullet and updated set, no cells changed [glm-5.3-flash]
  • Session analytics, trackers, surfaces, hybrid execution, spec-driven, software factory: ctx completed through v2.2.5 with beta self-hosted sharing (matrix cell moved), Artificial Analysis 689 models and llm-stats 400 canonical models, kiro CLI 2.26.0 and openchamber v2.1.0, Kev’s 1.0 family plus the previously unrecorded Kev-27B v2 flagship recorded with ImageJevBench v0.1.x added inside the JevBench note and the instructor download-trend claim corrected (flat, not declining), spec-kit v1.0.13 and ouroboros v0.55.3, super-simple-software-factory moved to quiet-since-launch (no push since 2026-08-04), machinist’s zero-coverage claim re-confirmed via Algolia [glm-5.3-flash]
  • People and publications: ten bump-only refreshes (Latent Space 202k subs with the Alex Zhang interview, Willison daily through 10-01, Pragmatic Engineer Pulse debate coverage, aider’s stall re-confirmed at 49.3k stars, and six more), one matrix cell moved (Owain Lewis cadence 15 to 16 uploads); no cadence breaks and no material changes anywhere in the category [glm-5.3-flash]
  • Entrant resolutions, all in-run: the new star rejected as out of scope, endstop.systems’ “containment is not safety” post rejected as an OpenShell critical source (the vendor withdrew its own overstated claims in a 2026-09-30 revision), Andon’s new eval benches folded into the existing Pion note rather than spun out; no category surfaced a candidate clearing the citation bar [glm-5.3-flash]
  • Queue: appended a ? for tom: bullet asking whether the section’s roughly 330 retained “real” uses should get the same rewording treatment the main corpus headings got on 10-01, since a sweep that size needs its own directed run with per-note Changes bullets [glm-5.3-flash]
  • _index.md: all 24 matrix one-liners re-dated to verified 2026-10-02, the model-selection and landscape as-of lines and the context-management link-check line re-dated, the Spec Kit one-liner moved to v1.0.13 and about 140k, the sandboxing one-liner’s OpenSandbox clause updated to patching-its-first-stable-wheel; no membership changes anywhere, so no category-page Changes bullets [glm-5.3-flash]
  • Verification: 0 broken internal link targets, every article front matter parses with no type field and mandatory tags, all 25 matrix tables (23 category matrices plus the provider and benchmark matrices) uniform at header width with case-insensitively sorted columns and membership matching their categories, no em-dashes, shape/honest/load bearing only at the documented exceptions (ante’s protocol-shape crate, laya’s quoted section title, the kimi-code URL slug, the emdash and multica quoted history bullets), substrate only as the Agent Substrate product name and its repo URL, posture only in security contexts plus ante’s quoted 2026-09-12 history bullet, and 329 “real” uses retained under the 10-01 sweep precedent pending the owner’s answer in the queue; every external URL cited by new or edited content was fetched this run or its absence recorded honestly (chatgpt.com/pricing and x.ai/pricing 403 bot walls worked around via official docs mirrors and webfetch, HN item pages 419 to curl verified via the Algolia items API, pypistats 429 cleared on backoff, harmonic.fun and harnesstax.github.io JS shells verified through direct subpages, search-retrieved text, and GitHub APIs, erdosproblems 403 to curl verified via webfetch, agentanalytics.sh pricing left dated as client-rendered) [glm-5.3-flash]