<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>metrics on tomrochette.com</title>
    <link>https://tomrochette.com/tags/metrics/</link>
    <description>Recent content in metrics on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Tue, 06 Oct 2026 02:31:53 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/metrics/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>Metrics for a Software Factory: Optimize Autonomy, Guard the Trust</title>
      <link>https://tomrochette.com/metrics-for-a-software-factory/</link>
      <pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/metrics-for-a-software-factory/</guid>
      <category>ai</category><category>llm</category><category>ai-agents</category><category>software-engineering</category><category>software-factory</category><category>metrics</category><category>okr</category><category>fully-ai-generated</category><category>llm=deepseek-v4.1-flash</category>
      <description>&lt;p&gt;A software factory is the infrastructure that lets LLM agents carry a change through the whole software lifecycle, from an observed signal to a shipped fix, instead of handing one step to a model and the rest to me.&#xA;Once the factory exists, the question that decides whether it was worth building is how much of that lifecycle runs without me.&#xA;&lt;strong&gt;The metrics I put on it decide what it optimizes for, and the obvious ones (agents spawned, tokens burned, pull requests merged, lines written) are the ones an agent can increase without producing anything I can use.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Output metrics measure the wrong thing&#xA;    &lt;div id=&#34;output-metrics-measure-the-wrong-thing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#output-metrics-measure-the-wrong-thing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Every measurement I inherited from human development counts effort.&#xA;Tickets closed, story points burned, pull requests merged, lines written: each one counted human activity because human activity was the expensive part, which is the argument I made in &lt;a href=&#34;../outcome-driven-development/index.md&#34; &gt;Outcome-Driven Development&lt;/a&gt;.&#xA;An agent produces all of it at almost no cost.&#xA;Point one at a backlog and it will empty the backlog, open fifty pull requests before lunch, and every one of them will look like a morning of work.&lt;/p&gt;&#xA;&lt;p&gt;This is &lt;a href=&#34;https://en.wikipedia.org/wiki/Goodhart%27s_law&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=en.wikipedia.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Goodhart&amp;rsquo;s law&lt;/a&gt; at its most extreme.&#xA;When a measure becomes a target it stops being a good measure, and an agent can turn any output metric into a target within a single run.&#xA;The failure is not theoretical.&#xA;&lt;a href=&#34;https://lucumr.pocoo.org/2026/9/7/astra-why/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=lucumr.pocoo.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Armin Ronacher&amp;rsquo;s 35 hours with GPT-6 Astra&lt;/a&gt; produced 79 commits, 75,000 lines, about a billion tokens, and roughly $1,200 of API spend over a weekend.&#xA;By volume alone it looked like his most productive weekend ever.&#xA;His verdict was that it delivered nothing of value.&lt;/p&gt;&#xA;&lt;p&gt;The problem is not only that volume misleads.&#xA;It is that volume looks like progress to the person running the factory too.&#xA;&lt;a href=&#34;https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=metr.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;METR&amp;rsquo;s randomized trial&lt;/a&gt; found that experienced developers using early-2025 AI tools took 19% longer on their own tasks, and those same developers still believed the tools had made them 20% faster.&#xA;&lt;strong&gt;When measured time and perceived speed disagree about the same work, self-reported productivity measures a feeling, not a result.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;The first step in building a factory is therefore to retire the output metrics, because an agent can increase every one of them for free.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Measure the factory the way a plant is measured&#xA;    &lt;div id=&#34;measure-the-factory-the-way-a-plant-is-measured&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#measure-the-factory-the-way-a-plant-is-measured&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Manufacturing has measured itself for a century and almost never treats gross output as success.&#xA;A plant that ships a thousand units at 10% yield produces 900 units nobody can use, and what matters is good units per unit of input, not units.&#xA;&lt;a href=&#34;https://en.wikipedia.org/wiki/First_pass_yield&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=en.wikipedia.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Yield&lt;/a&gt; is the fraction of what enters a process that comes out usable on the first pass.&#xA;A software factory has the same two numbers with different names.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Autonomy rate&lt;/strong&gt; is the share of accepted changes that reach production with no human intervention: no question answered mid-run, no correction, no manual fix, no human-triggered rerun, and no approval.&#xA;&lt;strong&gt;Yield&lt;/strong&gt; is the share of the autonomous output that survives: it passes verification, stays merged, and remains valid during the observation window after release.&#xA;Multiply the two and you get the number the factory exists to raise, &lt;strong&gt;the rate of trustworthy unattended work.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;Autonomy is easy to raise without doing the work, so the definition has to be strict about what counts as intervention.&#xA;A change that needed one clarification at hour three is not autonomous, and counting it as autonomous is how a dashboard reaches 90% while I am still involved in every hard case.&lt;/p&gt;&#xA;&lt;p&gt;A signal is an observed event that starts work: a bug report, a failing test, a support ticket, a security alert, or a request from a customer.&#xA;The loop below starts from that signal and shows where each metric is read.&lt;/p&gt;&#xA;&lt;pre class=&#34;not-prose mermaid&#34;&gt;flowchart LR&#xA;    S[Signal] --&amp;gt; P[Produce]&#xA;    P --&amp;gt; V{Independent verification}&#xA;    V --&amp;gt;|fails| P&#xA;    V --&amp;gt;|passes| G{Accept}&#xA;    G --&amp;gt;|human needed| H[Human gate]&#xA;    G --&amp;gt;|no human needed| M[Accepted change]&#xA;    H --&amp;gt; M&#xA;    M --&amp;gt; O[Observe in production]&#xA;    O --&amp;gt;|regression| R[Rework]&#xA;    R --&amp;gt; P&lt;/pre&gt;&#xA;&lt;p&gt;Autonomy counts the changes that travel from Produce to Accepted change without passing through the Human gate.&#xA;Yield counts the accepted changes that survive Observe.&#xA;Cost divides the whole path by the accepted changes, which is why the word &amp;ldquo;accepted&amp;rdquo; matters.&lt;/p&gt;&#xA;&lt;p&gt;Armin priced his run at about $15.50 per commit.&#xA;A commit is output, not a usable unit, so that price counts work whether or not it survived.&#xA;The denominator that survives is the accepted change: the change that passed verification and held up in production, not the change that got written.&#xA;&lt;strong&gt;Cost per commit counts output.&#xA;Cost per accepted change counts results.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;The guardrail that keeps autonomy trustworthy&#xA;    &lt;div id=&#34;the-guardrail-that-keeps-autonomy-trustworthy&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#the-guardrail-that-keeps-autonomy-trustworthy&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Autonomy is easy to raise by shipping worse work, so it cannot be the only metric.&#xA;The &lt;a href=&#34;https://dora.dev/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=dora.dev&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;DORA program&lt;/a&gt; spent years showing that delivery performance has two independent axes, throughput and stability, and that a team can improve one while the other gets worse.&#xA;Stability is change failure rate and time to restore, and in a factory it also includes defect escape rate, the defects that reach production past the gates, and rework rate, the accepted changes that had to be redone.&lt;/p&gt;&#xA;&lt;p&gt;The failure mode when stability goes unmeasured is well documented.&#xA;Addy Osmani documents a &lt;a href=&#34;https://addyosmani.com/blog/software-factories/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=addyosmani.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;&amp;ldquo;dark&amp;rdquo; factory&lt;/a&gt; that shipped for about four months with no human reading the code, passed its own tests the whole way, and then needed painstaking manual debugging to recover.&#xA;The tests passed because the factory wrote both the code and the tests, and the loop had no independent signal to catch what that matching pair got wrong together.&#xA;&lt;strong&gt;A factory with high autonomy and low yield produces unverified volume, and it accumulates comprehension debt faster than anyone can repay it.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;The two metrics are not interchangeable, and the grid shows why.&lt;/p&gt;&#xA;&lt;figure&gt;&lt;img&#xA;    class=&#34;my-0 rounded-md&#34;&#xA;    loading=&#34;lazy&#34;&#xA;    decoding=&#34;async&#34;&#xA;    fetchpriority=&#34;low&#34;&#xA;    alt=&#34;A two by two grid of autonomy against yield, with quadrants for a manual shop, unverified volume, a careful assistant, and a software factory&#34;&#xA;    src=&#34;https://tomrochette.com/metrics-for-a-software-factory/images/autonomy-yield.svg&#34;&#xA;    &gt;&lt;/figure&gt;&#xA;&lt;p&gt;The rule that keeps a factory out of the bottom-right quadrant is to pair the metrics.&#xA;&lt;strong&gt;Never set an autonomy target without a stability guardrail, because the cheapest way to raise autonomy is to lower the standard for done.&lt;/strong&gt;&#xA;DORA&amp;rsquo;s 2025 report found that 90% of the technology professionals it surveyed now use AI and more than 80% believe it made them more productive, while 30% report little or no trust in the code it generates, which is the gap between the two metrics stated as a statistic.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;The two numbers that decide whether the factory is worth running&#xA;    &lt;div id=&#34;the-two-numbers-that-decide-whether-the-factory-is-worth-running&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#the-two-numbers-that-decide-whether-the-factory-is-worth-running&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Once autonomy and yield are high enough, two more numbers decide whether the factory is a good use of money.&lt;/p&gt;&#xA;&lt;p&gt;Throughput comes first: lead time from signal to production, and deployment frequency.&#xA;The DORA speed axis still applies, but in a factory it is a consequence, not a goal.&#xA;Raising throughput without raising verification capacity only grows the queue in front of the gate, which is the back-pressure problem Osmani names: generation runs without limit while verification stays slow, so the extra output becomes waiting, not more usable code.&#xA;&lt;strong&gt;Track queue depth alongside throughput, because a growing backlog of unverified changes is the first sign that the factory produces faster than it can check.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;Unit cost comes second, and it has two parts.&#xA;The metered part is tokens and compute divided by accepted changes.&#xA;The part that matters more is human minutes per accepted change, because the point of the factory is to remove my minutes, and a factory that ships cheap code while I spend the same hours supervising it has automated only part of the work.&#xA;Armin&amp;rsquo;s run is the cautionary version of both: a weekend of compute for a yield near zero, with nobody supervising because supervising nobody was the point.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;The leading indicators you can control&#xA;    &lt;div id=&#34;the-leading-indicators-you-can-control&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#the-leading-indicators-you-can-control&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Outcome metrics lag by weeks.&#xA;Autonomy and yield tell me where I ended up, not what to do on Monday, so the OKR needs a layer of inputs that change first.&#xA;Four inputs predict the outcome metrics well enough to act on.&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;Spec coverage&lt;/strong&gt; is the share of work items that carry a machine-checkable outcome (what becomes true, how it is checked, what it may cost) rather than a task description.&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Gate coverage&lt;/strong&gt; is the share of change types covered by an independent machine check, as opposed to a check the producing agent wrote for itself.&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Context readiness&lt;/strong&gt; is whether a cold session dropped into the project can find everything it needs, the exit condition I use in &lt;a href=&#34;https://tomrochette.com/nine-months-of-llm-agents-on-large-projects/&#34; &gt;Nine Months of LLM Agents on Large Projects&lt;/a&gt;.&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Failure-to-gate capture&lt;/strong&gt; is the share of escaped defects that become a new automated check, the only mechanism that makes a factory improve instead of merely run.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;These are the inputs that turn model capability into stable throughput.&#xA;It is the same argument I made in &lt;a href=&#34;https://tomrochette.com/the-foundation-predicts-the-house-of-cards/&#34; &gt;Team Maturity Explains the Friction, the Foundation Predicts the House of Cards&lt;/a&gt;: verification infrastructure is the highest-return investment for a team shipping with agents, and in a factory it separates a loop that improves from one that only repeats.&lt;/p&gt;&#xA;&lt;p&gt;The layers stack like this, with the guardrails holding the two primary metrics in place.&lt;/p&gt;&#xA;&lt;figure&gt;&lt;img&#xA;    class=&#34;my-0 rounded-md&#34;&#xA;    loading=&#34;lazy&#34;&#xA;    decoding=&#34;async&#34;&#xA;    fetchpriority=&#34;low&#34;&#xA;    alt=&#34;Four metric layers under one objective: autonomy and yield on top, throughput and unit cost below them, a guardrail band that must be held flat, and the leading indicators that steer the rest&#34;&#xA;    src=&#34;https://tomrochette.com/metrics-for-a-software-factory/images/metric-layers.svg&#34;&#xA;    &gt;&lt;/figure&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Writing the OKR&#xA;    &lt;div id=&#34;writing-the-okr&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#writing-the-okr&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;An OKR made of autonomy numbers alone will produce the wrong behavior before the quarter ends.&#xA;The version that works pairs one key result that moves against another that holds, and adds one that steers.&lt;/p&gt;&#xA;&lt;p&gt;One workable objective, with the numbers as placeholders until you have your own baseline, is &lt;strong&gt;shift more of the software lifecycle into trustworthy unattended work.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;Move:&lt;/strong&gt; touchless share of accepted changes rises from 20% to 50%.&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Move:&lt;/strong&gt; human minutes per accepted change fall from 45 to 20.&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Hold:&lt;/strong&gt; defect escape rate and change failure rate stay at or below the pre-factory baseline.&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Steer:&lt;/strong&gt; the share of work items carrying a machine-checkable outcome rises from 30% to 90%.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;The touchless-share result is the most visible and the easiest to raise without doing the work, which is why it is not alone.&#xA;The human-minutes result is the one that matters most, and it is hard to lower without removing the work, because it measures my minutes.&#xA;The stability result keeps the factory from trading trust for autonomy, and it cannot be met by shipping worse work.&#xA;The spec-coverage result is the input I can change this week, and it makes the other three possible.&lt;/p&gt;&#xA;&lt;p&gt;Write the baseline down before the first factory run.&#xA;Without a baseline, anyone who preferred the old process can dispute the comparison, and the argument becomes about the numbers rather than the result.&lt;/p&gt;&#xA;&lt;p&gt;One metric sits above the factory, and it is not a factory metric: the product outcome.&#xA;If the product outcome is flat while autonomy climbs, the factory is producing the wrong thing faster.&#xA;&lt;strong&gt;The factory metrics multiply a product metric, they never replace it.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;The Goodhart problem, and how to limit it&#xA;    &lt;div id=&#34;the-goodhart-problem-and-how-to-limit-it&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#the-goodhart-problem-and-how-to-limit-it&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Autonomy rate is now a target, so the system and the people running it will find the cheapest work that satisfies it.&#xA;Trivial changes, tiny diffs, safe files, and a narrow definition of intervention all raise the number without improving the factory.&#xA;Five rules are worth adding from the start.&lt;/p&gt;&#xA;&lt;ol&gt;&#xA;&lt;li&gt;Report autonomy by change class, because a touchless migration and a touchless typo fix are not the same result.&lt;/li&gt;&#xA;&lt;li&gt;Measure on production traffic rather than a curated set, because a curated set misses the rare cases where the factory fails.&lt;/li&gt;&#xA;&lt;li&gt;Set a cap on trivial work, because an autonomy score built from work that never needed verifying is not a factory score.&lt;/li&gt;&#xA;&lt;li&gt;Keep a human-owned roadmap and a written invariant list, since an autonomous loop optimizes whatever its signals measure and only a human-edited direction corrects the drift, the argument in &lt;a href=&#34;https://tomrochette.com/the-self-evolving-repository/&#34; &gt;The Self-Evolving Repository&lt;/a&gt;.&lt;/li&gt;&#xA;&lt;li&gt;Audit the gates themselves, because a gate whose value was never measured only adds time, and &lt;a href=&#34;https://tomrochette.com/the-merge-gate/&#34; &gt;The Merge Gate&lt;/a&gt; asks whether a gate is worth its cost.&lt;/li&gt;&#xA;&lt;/ol&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What not to measure&#xA;    &lt;div id=&#34;what-not-to-measure&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-not-to-measure&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Number of agents, sessions, or concurrent workers, since more agents producing the same output is a more expensive factory.&lt;/li&gt;&#xA;&lt;li&gt;Tokens consumed, since it is an input and the factory&amp;rsquo;s job is to spend fewer per accepted change.&lt;/li&gt;&#xA;&lt;li&gt;Pull requests merged and lines written, since volume is what an agent gets for free.&lt;/li&gt;&#xA;&lt;li&gt;Story points and velocity, since they measure effort from an era when effort was scarce.&lt;/li&gt;&#xA;&lt;li&gt;AI adoption percentage, since using a tool is not an outcome.&lt;/li&gt;&#xA;&lt;li&gt;Self-reported speedup, since it contradicted the measured time in the METR trial.&lt;/li&gt;&#xA;&lt;li&gt;Benchmark scores, since they measure the model rather than the factory, and your repository&amp;rsquo;s cost and yield are the benchmark that matters.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What to Do Next&#xA;    &lt;div id=&#34;what-to-do-next&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-to-do-next&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ol&gt;&#xA;&lt;li&gt;Instrument the change lifecycle with an id, a start, an end, and a human-intervened flag on every change, because nothing else works until this exists.&lt;/li&gt;&#xA;&lt;li&gt;Compute two numbers this week, the touchless share of accepted changes and the defect escape rate, by hand from the last fifty changes if you have to.&lt;/li&gt;&#xA;&lt;li&gt;Add the two economic numbers once ids exist: cost per accepted change, and human minutes per accepted change.&lt;/li&gt;&#xA;&lt;li&gt;Write one paired OKR: move autonomy, hold stability, steer spec coverage.&lt;/li&gt;&#xA;&lt;li&gt;Record the baseline before the next factory run.&lt;/li&gt;&#xA;&lt;li&gt;Track queue depth, and treat a growing review backlog as a stop signal.&lt;/li&gt;&#xA;&lt;li&gt;Read the result plainly: autonomy up with stability flat means the factory is working, and autonomy up with stability down means you fix the gates before adding agents.&lt;/li&gt;&#xA;&lt;/ol&gt;&#xA;&lt;p&gt;&lt;strong&gt;A factory is not fast because it produces a lot.&#xA;It is fast because it finishes work you did not have to touch.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;../outcome-driven-development/index.md&#34; &gt;Outcome-Driven Development&lt;/a&gt; - why effort metrics stop being useful once agents execute, and how to write an outcome a machine can check&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/zero-touch-engineering/&#34; &gt;Zero Touch Engineering&lt;/a&gt; - the limit case where autonomy reaches 100% and the human authors the system instead of the change&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/verifying-code-without-reading-it/&#34; &gt;Verifying Code Without Reading It&lt;/a&gt; - the checks that make yield measurable without reading diffs&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/the-foundation-predicts-the-house-of-cards/&#34; &gt;Team Maturity Explains the Friction, the Foundation Predicts the House of Cards&lt;/a&gt; - the DORA two-axis argument and why verification is the highest-return investment&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/the-self-evolving-repository/&#34; &gt;The Self-Evolving Repository&lt;/a&gt; - the human-owned roadmap and invariants that keep an autonomous loop from drifting&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/nine-months-of-llm-agents-on-large-projects/&#34; &gt;Nine Months of LLM Agents on Large Projects&lt;/a&gt; - the context-readiness exit condition and the steering-time cost of under-provisioning&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/the-shifting-bottleneck/&#34; &gt;The Shifting Bottleneck&lt;/a&gt; - why automating production moves the constraint to verification, the layer a factory&amp;rsquo;s yield measures&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/the-merge-gate/&#34; &gt;The Merge Gate&lt;/a&gt; - the cost a gate has to justify, the test a factory&amp;rsquo;s automated checks have to pass&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/software-factory/&#34; &gt;Software factory&lt;/a&gt; - the agent-maintained research section that tracks working factories and their designs&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://dora.dev/research/2025/dora-report/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=dora.dev&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;DORA, &amp;ldquo;State of AI-assisted Software Development 2025&amp;rdquo;&lt;/a&gt; - the 90% adoption, 80% perceived-productivity, and 30% distrust figures, and the finding that AI amplifies the system it is used in&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://dora.dev/guides/dora-metrics-four-keys/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=dora.dev&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;DORA, &amp;ldquo;DORA&amp;rsquo;s software delivery metrics: the four keys&amp;rdquo;&lt;/a&gt; - the throughput and stability axes a factory inherits and reads as speed and guardrails&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=metr.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;METR, &amp;ldquo;Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity&amp;rdquo;&lt;/a&gt; - the 19% slowdown and the 20% perceived speedup that make self-reported productivity unusable as a metric&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2507.09089&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Becker, Rush, Barnes, and Rein, &amp;ldquo;Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity&amp;rdquo;&lt;/a&gt; - the paper behind the METR trial&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://lucumr.pocoo.org/2026/9/7/astra-why/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=lucumr.pocoo.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Armin Ronacher, &amp;ldquo;Astra for Coding: Why Are We Doing This Again?&amp;rdquo;&lt;/a&gt; - the documented factory run: 79 commits, 75,000 lines, about $1,200, nothing of value, and the per-commit price that motivated &amp;ldquo;cost per accepted change&amp;rdquo;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://addyosmani.com/blog/software-factories/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=addyosmani.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Addy Osmani, &amp;ldquo;Software Factories, Light and Dark&amp;rdquo;&lt;/a&gt; - the review gate as the bottleneck, back pressure, and the dark factory that ran four months before its comprehension debt appeared&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://addyosmani.com/blog/agentic-autonomy-levels&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=addyosmani.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Addy Osmani, &amp;ldquo;Agentic Autonomy Levels&amp;rdquo;&lt;/a&gt; - the metric set this article&amp;rsquo;s autonomy and yield definitions overlap with, including token cost and defect escape rate per accepted change&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://vercel.com/blog/building-a-software-factory-for-ai-sdk&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=vercel.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Vercel, &amp;ldquo;Building a Software Factory for the AI SDK&amp;rdquo;&lt;/a&gt; - a lit factory with mandatory human review, where agents authored 25 to 35% of merged pull requests within four weeks&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://en.wikipedia.org/wiki/Goodhart%27s_law&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=en.wikipedia.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Wikipedia, &amp;ldquo;Goodhart&amp;rsquo;s law&amp;rdquo;&lt;/a&gt; - why every autonomy and output target degrades once the system can optimize it directly&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://en.wikipedia.org/wiki/First_pass_yield&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=en.wikipedia.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Wikipedia, &amp;ldquo;First pass yield&amp;rdquo;&lt;/a&gt; - the manufacturing definition of yield as the share that passes without rework, the framing a software factory borrows&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://en.wikipedia.org/wiki/OKR&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=en.wikipedia.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;Wikipedia, &amp;ldquo;OKR&amp;rdquo;&lt;/a&gt; - the objective and key results format the paired OKR above follows&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
