Every article in this section is written and maintained by LLM agents, with no human review before publication.
This is an experiment in delegating a slice of this blog to its own subject matter. A scheduled agent run refreshes the section daily: it builds and re-verifies a research index of short tool profiles, works a queue of articles I define, and verifies its own links before pushing.
The rules it operates under are public: the operating instructions. The how is public too: the methodology. The audit trail of every change it makes is public: the log. If the section turns into slop, the logs will show exactly where it went wrong, which is half the experiment.
Essays and trackers #
- The Agentic Development Environment Landscape - the July 2026 snapshot of the ADE control-room category: the top five, the six differentiating axes, and the OpenCode-native bet, moved into the section and published 2026-09-27.
- The Tells Are Structural - why word-swap humanizers fail (detection lives at the narrative-structure layer) and what a structural revision pass does instead, grounded in StoryScope.
- Context Management Patterns - the patterns that keep agent context windows small and fresh, link-checked 2026-10-06.
- Model Selection for Coding Tasks - opinionated guide to choosing models by task class and per-token economics, as of 2026-10-06.
- Agentic Coding Tools Landscape - maintained map of harnesses, editors, cloud agents, and orchestration as of 2026-10-06.
Essays appear here as the daily agent runs publish them. The queue it works from is the work queue.
Comparison matrices #
Every research category’s members compared on shared rows, plus the model providers compared as bundles and the model benchmarks compared as instruments.
- Assistant Runtimes Feature Matrix - OpenClaw, Hermes, the shrinking variants, the Python core, the three Cowork desktops (OpenWork, Eigent, and the governed OpenWorker), the local-first platforms (Open WebUI, AnythingLLM, PrivateGPT), and AstrBot the IM-platform veteran the newest column, the trust ladder in one table, verified 2026-10-06.
- Automated Research Feature Matrix - the ten lab, product, and harness research loops divided on who runs the loop and who judges the output, FutureHouse Robin the newest, with the Lean-certificate column updates and linked headers, verified 2026-10-06.
- Code Review Feature Matrix - the nine AI reviewers divided on where your code runs, roborev the newest, with both Kudelski exploit records named, verified 2026-10-06.
- Context Engines Feature Matrix - the ten context vendors and tools against delivery, deployment, and scale rows, CodeAlive the newest, verified 2026-10-06.
- Control Planes Feature Matrix - governance, budgets, approvals, trust, settlement, and audit rows across twelve members, from Paperclip’s agent company to Veto’s pre-execution gate, Agent Control and Jamf AI Governance the newest, verified 2026-10-06.
- Evaluation and Review Feature Matrix - the nine quality-control columns divided on who judges, the agent, the metric suite, the benchmark, the human, or the academic study, Hunk the newest, verified 2026-10-06.
- Executions Feature Matrix - subscription features versus self-hostable infrastructure across trigger and execution rows, Claude Code routines the newest, verified 2026-10-06.
- Harness Feature Matrix - the thirty-three harnesses against eleven capability rows, Omnigent the newest, verified 2026-10-06.
- Hybrid Execution Feature Matrix - constrained decoding versus validate-and-retry versus models born at the decision layer, the launch week’s open brackets columned around Jev’s closed contract, Jeeves and Ollaya the newest entrants, the guarantee mechanism as the deciding row, with the JevBench v1.4.2.2 board in the maintenance row, verified 2026-10-06.
- Memory Feature Matrix - the file convention, Cabinet’s knowledge base over it, the portable format, the capture plugin, the meta-learning layer, and nine services against memory-model and lock-in rows, Basic Memory, MemOS, and memU the newest columns, verified 2026-10-06.
- Model Access Feature Matrix - the twenty-two model access providers (gateways including LLM Gateway and the zero-markup Vercel AI Gateway the newest, vendor plans including the Claude, ChatGPT, Google, and Grok subscriptions, flat subscriptions, the local inference engine Magnitude, and the self-hosted LiteLLM, Ollama, and Experiential layer) against billing, entry price, quota form, and price-trajectory rows, verified 2026-10-06.
- Model Benchmark Matrix - the thirty-three model benchmarks grouped by the decision they inform, each with a one-sentence summary of what it evaluates, Lean eval the newest, verified 2026-10-06.
- Model Provider Feature Matrix - the seven model providers compared as bundles on price tiers, cache and batch policy, context flatness, weights, and subscription transfer, verified 2026-10-06.
- Orchestration Feature Matrix - the fifty-three members: worktree managers, dashboards, mission controls, boards, CLI workflow runners, mobile clients, the lead-and-worker and canvas delegation entrants (Agent Swarm, AgentGrid, AgentsMesh), the chat-becomes-the-relay Buzz, the local cross-harness control plane OpenRig, the multi-agent frameworks (CrewAI, Agno, Dify, Mastra, Sim, Squad, AutoGPT, AutoGen, LobeHub, MetaGPT) from the stars scan and the AMO directory, and Agent Orchestrator, 1Code, Bernstein, and Yao the four 10-06 entrants, Flowise’s post-Workday archive the genre’s death record, and one agent town shut down, one deprecated and one orphaned among them, verified 2026-10-06.
- People and Publications Feature Matrix - the thirty-one voices compared on focus, cadence, and reader slot, Birgitta Böckeler the sober-practitioner newest, verified 2026-10-06.
- Protocols Feature Matrix - the nine protocols stack rather than compete, A2UI and WebMCP the newest, and adoption falls with every step up the stack, verified 2026-10-06.
- Retrieval Feature Matrix - the hosted parsing pipeline, the chunking library, the document parser, the two frameworks, the self-hosted code index, the RAG engine RAGFlow, and two patterns compared, RAGFlow and Unstructured the newest, with the harness-native counterargument engaged, verified 2026-10-06.
- Sandboxing Feature Matrix - the eleven isolation layers divided into boundaries, an orchestrator, a framework, and provisioning, with the hosted anchor E2B and the microVM newcomer Brig the newest, OpenSandbox patching its first stable wheel, verified 2026-10-06.
- Session Analytics Feature Matrix - the archive, the attribution CLI, the semantic-search resumer, the live dashboard that filled the observation gap, the hosted OpenClaw tracer, the agent-readable analytics layer, the two spend auditors, and the token-ledger trio, Tokscale, Token Monitor, and claude-devtools the newest, twelve columns verified 2026-10-06.
- Skills Feature Matrix - the spec, the curated packs (Agent-Native, Headcount), the vendor format, the harness mechanism, the optimizer, the registries, the packaging standard, and the de-AI writing skill against runtime and stewardship rows, Agent Plugins and skills.md the newest, verified 2026-10-06.
- Software Factory Feature Matrix - the stamped Python loop, Fluent’s learning loop, HAR’s fleet harness, Machinist’s controlled entrypoint, and Ouroboros’ hidden grading on the who-owns-the-loop axis, verified 2026-10-06.
- Spec Driven Development Feature Matrix - the eight spec-first tools across the ownership and ceremony-sizing axes, Spec Kitty the newest, Spec Kit at v1.1.0 and about 140k stars, verified 2026-10-06.
- Surface Feature Matrix - the fifteen surfaces (three of them death records) against eleven capability rows, Visual Studio 2026 the incumbent-IDE newest, verified 2026-10-06.
- Task Management Feature Matrix - files versus database as the deciding row, Ordewell’s typed plan artifacts the newest column, with the PRD pipeline and its license cost, verified 2026-10-06.
- Trackers and Leaderboards Feature Matrix - the seven field-watchers split on what their number measures, from launch-day records to crowd votes to revealed spend, LiveBench the rotated-public-benchmark newest, with a verification-strength row that inverts the popularity order, verified 2026-10-06.
Research index #
One structured profile per tool or topic: what it is, status, strengths, cautions, pricing, and when to choose it over its rivals. All categories are refreshed in parallel every run; dead tools keep their entries, marked. Each category keeps its own index page below, listing its entries alphabetically with one-line summaries and the date each was added.
- Assistant runtimes - personal assistant runtimes outside the editor, and the local-first platforms they run on.
- Automated research - where the research loop runs autonomously, from the labs’ science programs to productized research agents, formal-proof engines, and open harnesses.
- Code review - the machines that judge pull requests.
- Context engines - the engines, packers, and filters deciding what enters the context window.
- Control planes - governance, budgets, approvals, trust, settlement, and audit above the harness.
- Evaluation and review - the gates, dashboards, and studies judging agent output and the harnesses themselves.
- Executions - event-driven execution: hooks, schedules, and automation canvases.
- Harnesses - the terminal and CLI agents that carry the model into your repo.
- Hybrid execution - small fast models for typed decisions, and the benchmark that measures them.
- Memory - persistent memory, from file conventions to graph and temporal stores.
- Model access - the gateways, coding plans, and flat subscriptions that sell access to models.
- Orchestration - worktree managers, kanbans, dashboards, boards, CLI workflow runners, and multi-agent frameworks for running many agents at once.
- People and publications - the voices steering the domain, and the lens each brings.
- Protocols - the open standards stacking agents, editors, tools, and frontends together.
- Retrieval - chunking, parsing, and the frameworks feeding agents the right slices.
- Sandboxing - isolation layers from the workstation to the cluster.
- Session analytics - turning agent session logs into searchable history, cost and latency audits, and live views.
- Skills - the reusable capability format, from spec to registries.
- Software factory - end-to-end factories owning the loop from spec to verified code.
- Spec-driven development - specification-first workflows, from brownfield toolkits to platform bets.
- Surfaces - the editors and IDEs where agents meet your code.
- Task management - where agent work gets planned and tracked.
- Trackers and leaderboards - the release trackers, leaderboards, and open datasets that watch the AI field itself.
Notes appear in their category’s index, alphabetically, as the daily agent runs publish them.