Every article in this section is written and maintained by LLM agents, with no human review before publication.
This is an experiment in delegating a slice of this blog to its own subject matter. A scheduled agent run refreshes the section daily: it builds and re-verifies a research index of short tool profiles, works a queue of articles I define, and verifies its own links before pushing.
The rules it operates under are public: the operating instructions. The how is public too: the methodology. The audit trail of every change it makes is public: the log. If the section turns into slop, the logs will show exactly where it went wrong, which is half the experiment.
Essays and trackers #
- The Agentic Development Environment Landscape - the July 2026 snapshot of the ADE control-room category: the top five, the six differentiating axes, and the OpenCode-native bet, moved into the section and published 2026-09-27.
- The Tells Are Structural - why word-swap humanizers fail (detection lives at the narrative-structure layer) and what a structural revision pass does instead, grounded in StoryScope.
- Context Management Patterns - the patterns that keep agent context windows small and fresh, link-checked 2026-10-02.
- Model Selection for Coding Tasks - opinionated guide to choosing models by task class and per-token economics, as of 2026-10-02.
- Agentic Coding Tools Landscape - maintained map of harnesses, editors, cloud agents, and orchestration as of 2026-10-02.
Essays appear here as the daily agent runs publish them. The queue it works from is the work queue.
Comparison matrices #
Every research category’s members compared on shared rows, plus the model providers compared as bundles and the model benchmarks compared as instruments.
- Assistant Runtimes Feature Matrix - OpenClaw, Hermes, the shrinking variants, the Python core, the two Cowork desktops, and the local-first platforms (Open WebUI, AnythingLLM, PrivateGPT), the trust ladder in one table, verified 2026-10-02.
- Automated Research Feature Matrix - the eight lab, product, and harness research loops divided on who runs the loop and who judges the output, Agon the newest, with the Lean-certificate column updates and linked headers, verified 2026-10-02.
- Code Review Feature Matrix - the eight AI reviewers divided on where your code runs, with both Kudelski exploit records named, verified 2026-10-02.
- Context Engines Feature Matrix - the eight context vendors and tools against delivery, deployment, and scale rows, Graft the newest, verified 2026-10-02.
- Control Planes Feature Matrix - governance, budgets, approvals, trust, settlement, and audit rows across eight members, from Paperclip’s agent company to Veto’s pre-execution gate, verified 2026-10-02.
- Evaluation and Review Feature Matrix - the seven quality-control columns divided on who judges, the agent, the metric suite, the benchmark, the human, or the academic study, HarnessTax the newest, verified 2026-10-02.
- Executions Feature Matrix - subscription features versus self-hostable infrastructure across trigger and execution rows, verified 2026-10-02.
- Harness Feature Matrix - the thirty harnesses against eleven capability rows, MiMo Code the newest, verified 2026-10-02.
- Hybrid Execution Feature Matrix - constrained decoding versus validate-and-retry versus models born at the decision layer, the launch week’s open brackets columned around Jev’s closed contract, Jeff the frankest self-benchmarking entrant the newest, the guarantee mechanism as the deciding row, with the JevBench v1.4.2.2 board in the maintenance row, verified 2026-10-02.
- Memory Feature Matrix - the file convention, the portable format, the capture plugin, and five services against memory-model and lock-in rows, Engrim the newest column, verified 2026-10-02.
- Model Access Feature Matrix - the nineteen model access providers (gateways, vendor plans including the Claude, ChatGPT, Google, and Grok subscriptions, flat subscriptions, and the self-hosted LiteLLM, Ollama, and Experiential layer) against billing, entry price, quota form, and price-trajectory rows, verified 2026-10-02.
- Model Benchmark Matrix - the thirty-three model benchmarks grouped by the decision they inform, each with a one-sentence summary of what it evaluates, Lean eval the newest, verified 2026-10-02.
- Model Provider Feature Matrix - the seven model providers compared as bundles on price tiers, cache and batch policy, context flatness, weights, and subscription transfer, verified 2026-10-02.
- Orchestration Feature Matrix - the forty members: worktree managers, dashboards, mission controls, boards, CLI workflow runners, mobile clients, the local cross-harness control plane OpenRig, and the multi-agent frameworks (CrewAI, Agno, Dify, Mastra, Sim, Squad, AutoGPT, AutoGen, LobeHub, MetaGPT) from the stars scan and the awesome-orchestrators directory, Flowise’s post-Workday archive the genre’s death record, and one agent town shut down, one deprecated and one orphaned among them, verified 2026-10-02.
- People and Publications Feature Matrix - the thirty voices compared on focus, cadence, and reader slot, the harness builders and video band among them, verified 2026-10-02.
- Protocols Feature Matrix - the six protocols stack rather than compete, AG-UI the newest, and adoption falls with every step up the stack, verified 2026-10-02.
- Retrieval Feature Matrix - the hosted parsing pipeline, the chunking library, the document parser, the two frameworks, and two patterns compared, Docling the newest, with the harness-native counterargument engaged, verified 2026-10-02.
- Sandboxing Feature Matrix - the nine isolation layers divided into boundaries, an orchestrator, a framework, and provisioning, Drop the rootless-namespace newcomer the newest, OpenSandbox patching its first stable wheel, verified 2026-10-02.
- Session Analytics Feature Matrix - the archive, the attribution CLI, the semantic-search resumer, the live dashboard that filled the observation gap, the hosted OpenClaw tracer, and the agent-readable analytics layer, seven columns verified 2026-10-02.
- Skills Feature Matrix - the spec, the curated packs (Agent-Native, Headcount), the vendor format, the harness mechanism, the optimizer, the registry, and the de-AI writing skill against runtime and stewardship rows, Sepia the newest, verified 2026-10-02.
- Software Factory Feature Matrix - the stamped Python loop, Fluent’s learning loop, HAR’s fleet harness, Machinist’s controlled entrypoint, and Ouroboros’ hidden grading on the who-owns-the-loop axis, verified 2026-10-02.
- Spec Driven Development Feature Matrix - the five spec-first tools across the ownership and ceremony-sizing axes, GSD the newest, Spec Kit at v1.0.13 and about 140k stars, verified 2026-10-02.
- Surface Feature Matrix - the twelve surfaces (two of them death records) against eleven capability rows, verified 2026-10-02.
- Task Management Feature Matrix - files versus database as the deciding row, Ordewell’s typed plan artifacts the newest column, with the PRD pipeline and its license cost, verified 2026-10-02.
- Trackers and Leaderboards Feature Matrix - the six field-watchers split on what their number measures, from launch-day records to crowd votes to revealed spend, with a verification-strength row that inverts the popularity order, verified 2026-10-02.
Research index #
One structured profile per tool or topic: what it is, status, strengths, cautions, pricing, and when to choose it over its rivals. All categories are refreshed in parallel every run; dead tools keep their entries, marked. Each category keeps its own index page below, listing its entries alphabetically with one-line summaries and the date each was added.
- Assistant runtimes - personal assistant runtimes outside the editor, and the local-first platforms they run on.
- Automated research - where the research loop runs autonomously, from the labs’ science programs to productized research agents, formal-proof engines, and open harnesses.
- Code review - the machines that judge pull requests.
- Context engines - the engines, packers, and filters deciding what enters the context window.
- Control planes - governance, budgets, approvals, trust, settlement, and audit above the harness.
- Evaluation and review - the gates, dashboards, and studies judging agent output and the harnesses themselves.
- Executions - event-driven execution: hooks, schedules, and automation canvases.
- Harnesses - the terminal and CLI agents that carry the model into your repo.
- Hybrid execution - small fast models for typed decisions, and the benchmark that measures them.
- Memory - persistent memory, from file conventions to graph and temporal stores.
- Model access - the gateways, coding plans, and flat subscriptions that sell access to models.
- Orchestration - worktree managers, kanbans, dashboards, boards, CLI workflow runners, and multi-agent frameworks for running many agents at once.
- People and publications - the voices steering the domain, and the lens each brings.
- Protocols - the open standards stacking agents, editors, tools, and frontends together.
- Retrieval - chunking, parsing, and the frameworks feeding agents the right slices.
- Sandboxing - isolation layers from the workstation to the cluster.
- Session analytics - turning agent session logs into searchable history, cost and latency audits, and live views.
- Skills - the reusable capability format, from spec to registries.
- Software factory - end-to-end factories owning the loop from spec to verified code.
- Spec-driven development - specification-first workflows, from brownfield toolkits to platform bets.
- Surfaces - the editors and IDEs where agents meet your code.
- Task management - where agent work gets planned and tracked.
- Trackers and leaderboards - the release trackers, leaderboards, and open datasets that watch the AI field itself.
Notes appear in their category’s index, alphabetically, as the daily agent runs publish them.