Skip to content

2026 05 16 agent operational cost vs gap closure cost

github-actions[bot] edited this page May 16, 2026 · 1 revision

Agent Operational Cost vs Gap Closure Cost

Research Question

What is the fully loaded operational cost of a production Artificial Intelligence (AI) agent used as a workaround for a missing system capability, relative to the cost of closing the underlying systems capability gap through software delivery across three-year and five-year horizons, and under what conditions does recurring agent operation produce positive return relative to closing the gap in software?

Scope

In scope:

  • Cost model components for production agent workarounds: build, inference, harness operations, governance, and failure recovery.
  • Cost model components for software-gap closure across missing integration capability, missing application functionality, and missing governed data access.
  • Breakeven analysis using plausible AI-assisted software engineering productivity multipliers over three-year and five-year horizons.

Out of scope:

  • Vendor selection for a specific model provider or orchestration platform.
  • Organisation-specific budgeting recommendations without comparable benchmark data.

Constraints:

  • Prioritise empirical cost datasets and benchmark studies over anecdotal case studies.
  • Keep assumptions explicit when source data is incomplete.

Context

Recurring agent-workaround operation only outperforms system change when the underlying capability gap is too expensive, too uncertain, or too unstable to close within the planning horizon. [inference; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md]

Prior repository work shows that stable, repeatable side effects become more reliable when they move from probabilistic agent behavior into deterministic system capability, which makes long-horizon cost comparison the relevant decision test. [inference; source: https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md]

Approach

  1. Compile production cost evidence for agent workarounds by complexity and task type, including inference, governance, and recovery overhead.
  2. Compile delivery-cost evidence for closing underlying systems capability gaps with AI-assisted software engineering.
  3. Model breakeven points and sensitivity to productivity multipliers for missing integration capability, missing application functionality, and missing governed data access.

Sources


Research Skill Output

§0 Initialise

§1 Question Decomposition

What is the cost of operating a recurring agent workaround versus closing the gap in software?
|
|-- Q1. What costs are unavoidable for a production agent workaround?
|   |-- Q1a. What do model tokens, runtime, and search actually cost?
|   |-- Q1b. What governance, review, and risk-management work remains human?
|   `-- Q1c. What failure-recovery or retry burden remains after deployment?
|
|-- Q2. What delivery-cost assumptions are defensible for software-gap closure?
|   |-- Q2a. What is a reasonable loaded software-delivery labor cost anchor?
|   |-- Q2b. What productivity multiplier range is supported by current evidence?
|   `-- Q2c. What ongoing maintenance cost remains after the gap is closed?
|
|-- Q3. How do common capability-gap classes differ?
|   |-- Q3a. What is a representative missing-integration closure effort?
|   |-- Q3b. What is a representative missing-functionality closure effort?
|   `-- Q3c. What is a representative missing-governed-data-access closure effort?
|
`-- Q4. Under what conditions does recurring agent operation produce positive return?
    |-- Q4a. What are the three-year breakeven thresholds?
    |-- Q4b. What are the five-year breakeven thresholds?
    `-- Q4c. Which contextual factors favor recurring agent operation as a bridge rather than an end state?

§2 Investigation

Prior completed-item cross-reference

Q1a - Direct model, runtime, and search costs

Q1b - Governance, review, and oversight burden

Q1c - Failure recovery and reliability burden

Q2a - Loaded labor cost anchor for software-gap-closure delivery

Q2b - Productivity multiplier range for AI-assisted software delivery

Q2c - Maintenance after software-gap closure

  • [fact] DORA's ROI report says leaders should budget for an initial productivity dip, reinvest reclaimed capacity, and connect engineering improvements to sustainable financial outcomes instead of assuming one-time deployment instantly translates into durable value. Source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development. Source class: primary.
  • [assumption] A 10% annual maintenance rate on the initial build cost is a reasonable generic planning assumption for a closed capability gap, because the delivered integration, function, or data-access layer still requires upkeep but no longer needs the full recurring operational workaround. Justification: the question asks for breakeven planning horizons, and DORA's ROI framing is about sustained but reduced post-rollout investment rather than zero maintenance. Source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development.

Q3 - Archetype capability-gap classes and total-cost model

Q4 - Breakeven thresholds and decision conditions

§3 Reasoning

§4 Consistency Check

§5 Depth and Breadth Expansion

§6 Synthesis

Executive Summary

Closing the capability gap in software usually beats recurring agent operation over three-year and five-year horizons when the underlying missing capability can be built in roughly 14 to 21 delivery-months in a lean-to-standard operating model, and in roughly 20 to 30 delivery-months over five years. [inference; source: https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development]

Recurring human review, governance, and recovery labor dominate production agent-workaround cost, while published model prices remain comparatively low and risk-management duties remain persistent. [inference; source: https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing; https://dora.dev/research/2024/ai-preview/; https://doi.org/10.6028/NIST.AI.100-1]

Recurring agent operation earns positive return mainly as a bridge, when demand is uncertain, the workflow is changing too quickly to encode safely, or the build effort exceeds about two years of AI-assisted delivery. [inference; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://dora.dev/research/2024/dora-report/; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md]

Stable, high-frequency, compliance-relevant work should usually migrate out of recurring agent operation and into deterministic system capability because recurring oversight cost compounds while software-gap closure converts the same need into a lower-maintenance asset. [inference; source: https://doi.org/10.6028/NIST.AI.100-1; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md]

Key Findings

  1. Under the modeled medium-volume workload used in this item, published frontier-model prices imply that annual token, runtime, and search spend stays in the low-thousands of dollars, which leaves labor as the dominant operational cost driver. ([inference]; medium confidence; source: https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing)
  2. NIST governance requirements, low trust in generated code, and GitHub's documented review flow show that production agents retain recurring human oversight, approval, and monitoring cost even when the technical execution path is automated. ([inference]; high confidence; source: https://doi.org/10.6028/NIST.AI.100-1; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf; https://dora.dev/research/2024/ai-preview/; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review)
  3. Current software-engineering productivity evidence supports a planning multiplier band from about 1.02x at the organizational level to about 1.56x on bounded tasks, with about 1.26x as the strongest central estimate for real deployment economics. ([inference]; medium confidence; source: https://cloud.google.com/resources/content/dora-impact-of-gen-ai-software-development; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://arxiv.org/abs/2302.06590)
  4. Using the central 1.26x multiplier, representative three-year software-gap-closure totals are about $97,000 for a missing integration capability, $145,000 for missing application functionality, and $193,000 for missing governed data access, all materially below the standard $336,000 three-year recurring agent-workaround cost. ([inference]; medium confidence; source: https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development)
  5. Under the modeled staffing range, software-gap closure breaks even against recurring agent workarounds at about 14.2 to 40.6 delivery-months over three years and about 20.2 to 58.0 delivery-months over five years, with the standard case centered at about 20.9 and 29.8 months. ([inference]; medium confidence; source: https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://doi.org/10.6028/NIST.AI.100-1; https://dora.dev/research/2024/ai-preview/)
  6. Recurring agent operation keeps a positive return profile mainly when the closure effort is unusually large, the demand pattern is intermittent, or the operating design is explicitly temporary while the organization learns the workflow and validates the target state. ([inference]; medium confidence; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://dora.dev/research/2024/dora-report/; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md)
  7. For regulated or safety-relevant workflows, the required review and evidence burden shifts the economics further toward software-gap closure because stricter oversight raises recurring agent-workaround cost faster than it raises post-build maintenance cost. ([inference]; medium confidence; source: https://doi.org/10.6028/NIST.AI.100-1; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md)
  8. The strategic role of recurring agent operation is therefore best understood as bridge capital for uncertain or rapidly changing work, not as the default steady-state operating model for stable high-frequency system gaps. ([inference]; medium confidence; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md)

Evidence Map

Claim Source Confidence Notes
[inference] Under the modeled medium-volume workload, token and runtime spend remains small relative to labor. https://www.anthropic.com/pricing#api ; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing medium Depends on explicit workload assumption, but the order of magnitude is stable across providers.
[inference] Oversight and governance remain recurring costs in production. https://doi.org/10.6028/NIST.AI.100-1 ; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf ; https://dora.dev/research/2024/ai-preview/ ; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task ; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review high Multiple primary sources agree that review and governance are not optional.
[inference] The defensible productivity band is about 1.02x to 1.56x, centered around 1.26x. https://cloud.google.com/resources/content/dora-impact-of-gen-ai-software-development ; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/ ; https://arxiv.org/abs/2302.06590 medium Uses organization, field, and bounded-task evidence together.
[inference] Three-year software-gap-closure totals for missing integration capability, missing application functionality, and missing governed data access remain below the standard three-year recurring agent-workaround total. https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/ ; https://www.bls.gov/oes/2023/may/oes151252.htm ; https://www.bls.gov/news.release/ecec.t01.htm ; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development medium Relies on explicit archetype effort assumptions.
[inference] Breakeven spans about 14.2 to 40.6 delivery-months over three years and about 20.2 to 58.0 over five years across the modeled staffing range, with the standard case at about 20.9 and 29.8 months. https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/ ; https://www.bls.gov/oes/2023/may/oes151252.htm ; https://www.bls.gov/news.release/ecec.t01.htm ; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development ; https://doi.org/10.6028/NIST.AI.100-1 ; https://dora.dev/research/2024/ai-preview/ medium Threshold is sensitive to staffing assumption but direction is robust.
[inference] Recurring agent operation is strongest as a temporary bridge for uncertain or fast-changing work. https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development ; https://dora.dev/research/2024/dora-report/ ; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md medium Supported by DORA's J-curve and prior repository architecture findings.
[inference] Stricter control environments move the economics further toward software-gap closure. https://doi.org/10.6028/NIST.AI.100-1 ; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf ; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md medium Governance intensity increases recurring agent-workaround labor more than post-build maintenance.
[inference] Stable, high-frequency work should migrate from recurring agent operation into deterministic system capability. https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md ; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md ; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development medium This is the combined economic and control-boundary conclusion.

Identified but not consulted

Assumptions

Analysis

This item relies mainly on published pricing schedules, governance frameworks, and software-engineering productivity studies rather than anecdotal case studies. [fact; source: https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing; https://doi.org/10.6028/NIST.AI.100-1; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://arxiv.org/abs/2302.06590]

Token spend remains small in the modeled workload, while recurring human review, governance, and failure-recovery work account for most annual operating cost in the agent-workaround scenarios used here. [inference; source: https://dora.dev/research/2024/ai-preview/; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review; https://doi.org/10.6028/NIST.AI.100-1; https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing]

The software-gap-closure side is more assumption-sensitive, so the analysis keeps the logic transparent: loaded labor, productivity multiplier, capability-gap effort, and maintenance rate are all visible, and changing them shifts the breakeven month threshold rather than reversing the direction for common stable gaps. [inference; source: https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development]

The main competing interpretation is that model quality will keep rising fast enough to erase oversight cost, but the reviewed DORA and GitHub evidence does not support that today because organizations still route consequential changes through human review and still report reliability trade-offs from adoption. [inference; source: https://dora.dev/research/2024/dora-report/; https://dora.dev/research/2024/ai-preview/; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review]

The second competing interpretation is to keep recurring agent operation indefinitely because it avoids the up-front project, but DORA's return-on-investment framing and prior repository architecture work both indicate that durable value comes from converting repeated workaround effort into a governed capability, not from paying the same workaround tax forever. [inference; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md]

Risks, Gaps, and Uncertainties

Open Questions

  • What is the empirical maintenance ratio for production agent workarounds after twelve months in stable enterprise use, broken down by review, operations, and incident recovery?
  • How different are the breakeven thresholds for low-frequency but high-value work compared with high-frequency transactional work?
  • What project-history datasets could replace the six-month, nine-month, and twelve-month capability-gap archetypes with observed delivery distributions?
  • How much option value does a deliberate recurring-agent pilot create when it is used to discover the right target-state system design rather than as an indefinite workaround?

Output

§7 Recursive Review

review_status: self-review completed
acronym_audit: complete
domain_term_audit: rewritten into plain language where no authoritative external definition was used
claim_audit: complete
findings_parity: maintained
remaining_uncertainty: threshold values depend on effort and staffing assumptions

Findings

Executive Summary

Closing the capability gap in software usually beats recurring agent operation over three-year and five-year horizons when the underlying missing capability can be built in roughly 14 to 21 delivery-months in a lean-to-standard operating model, and in roughly 20 to 30 delivery-months over five years. [inference; source: https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development]

Recurring human review, governance, and recovery labor dominate production agent-workaround cost, while published model prices remain comparatively low and risk-management duties remain persistent. [inference; source: https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing; https://dora.dev/research/2024/ai-preview/; https://doi.org/10.6028/NIST.AI.100-1]

Recurring agent operation earns positive return mainly as a bridge, when demand is uncertain, the workflow is changing too quickly to encode safely, or the build effort exceeds about two years of AI-assisted delivery. [inference; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://dora.dev/research/2024/dora-report/; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md]

Stable, high-frequency, compliance-relevant work should usually migrate out of recurring agent operation and into deterministic system capability because recurring oversight cost compounds while software-gap closure converts the same need into a lower-maintenance asset. [inference; source: https://doi.org/10.6028/NIST.AI.100-1; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md]

Key Findings

  1. Under the modeled medium-volume workload used in this item, published frontier-model prices imply that annual token, runtime, and search spend stays in the low-thousands of dollars, which leaves labor as the dominant operational cost driver. ([inference]; medium confidence; source: https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing)
  2. NIST governance requirements, low trust in generated code, and GitHub's documented review flow show that production agents retain recurring human oversight, approval, and monitoring cost even when the technical execution path is automated. ([inference]; high confidence; source: https://doi.org/10.6028/NIST.AI.100-1; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf; https://dora.dev/research/2024/ai-preview/; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review)
  3. Current software-engineering productivity evidence supports a planning multiplier band from about 1.02x at the organizational level to about 1.56x on bounded tasks, with about 1.26x as the strongest central estimate for real deployment economics. ([inference]; medium confidence; source: https://cloud.google.com/resources/content/dora-impact-of-gen-ai-software-development; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://arxiv.org/abs/2302.06590)
  4. Using the central 1.26x multiplier, representative three-year software-gap-closure totals are about $97,000 for a missing integration capability, $145,000 for missing application functionality, and $193,000 for missing governed data access, all materially below the standard $336,000 three-year recurring agent-workaround cost. ([inference]; medium confidence; source: https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development)
  5. Under the modeled staffing range, software-gap closure breaks even against recurring agent workarounds at about 14.2 to 40.6 delivery-months over three years and about 20.2 to 58.0 delivery-months over five years, with the standard case centered at about 20.9 and 29.8 months. ([inference]; medium confidence; source: https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://doi.org/10.6028/NIST.AI.100-1; https://dora.dev/research/2024/ai-preview/)
  6. Recurring agent operation keeps a positive return profile mainly when the closure effort is unusually large, the demand pattern is intermittent, or the operating design is explicitly temporary while the organization learns the workflow and validates the target state. ([inference]; medium confidence; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://dora.dev/research/2024/dora-report/; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md)
  7. For regulated or safety-relevant workflows, the required review and evidence burden shifts the economics further toward software-gap closure because stricter oversight raises recurring agent-workaround cost faster than it raises post-build maintenance cost. ([inference]; medium confidence; source: https://doi.org/10.6028/NIST.AI.100-1; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md)
  8. The strategic role of recurring agent operation is therefore best understood as bridge capital for uncertain or rapidly changing work, not as the default steady-state operating model for stable high-frequency system gaps. ([inference]; medium confidence; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md)

Evidence Map

Claim Source Confidence Notes
[inference] Under the modeled medium-volume workload, token and runtime spend remains small relative to labor. https://www.anthropic.com/pricing#api ; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing medium Depends on explicit workload assumption, but the order of magnitude is stable across providers.
[inference] Oversight and governance remain recurring costs in production. https://doi.org/10.6028/NIST.AI.100-1 ; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf ; https://dora.dev/research/2024/ai-preview/ ; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task ; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review high Multiple primary sources agree that review and governance are not optional.
[inference] The defensible productivity band is about 1.02x to 1.56x, centered around 1.26x. https://cloud.google.com/resources/content/dora-impact-of-gen-ai-software-development ; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/ ; https://arxiv.org/abs/2302.06590 medium Uses organization, field, and bounded-task evidence together.
[inference] Three-year software-gap-closure totals for missing integration capability, missing application functionality, and missing governed data access remain below the standard three-year recurring agent-workaround total. https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/ ; https://www.bls.gov/oes/2023/may/oes151252.htm ; https://www.bls.gov/news.release/ecec.t01.htm ; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development medium Relies on explicit archetype effort assumptions.
[inference] Breakeven spans about 14.2 to 40.6 delivery-months over three years and about 20.2 to 58.0 over five years across the modeled staffing range, with the standard case at about 20.9 and 29.8 months. https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/ ; https://www.bls.gov/oes/2023/may/oes151252.htm ; https://www.bls.gov/news.release/ecec.t01.htm ; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development ; https://doi.org/10.6028/NIST.AI.100-1 ; https://dora.dev/research/2024/ai-preview/ medium Threshold is sensitive to staffing assumption but direction is robust.
[inference] Recurring agent operation is strongest as a temporary bridge for uncertain or fast-changing work. https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development ; https://dora.dev/research/2024/dora-report/ ; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md medium Supported by DORA's J-curve and prior repository architecture findings.
[inference] Stricter control environments move the economics further toward software-gap closure. https://doi.org/10.6028/NIST.AI.100-1 ; https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf ; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md medium Governance intensity increases recurring agent-workaround labor more than post-build maintenance.
[inference] Stable, high-frequency work should migrate from recurring agent operation into deterministic system capability. https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md ; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md ; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development medium This is the combined economic and control-boundary conclusion.

Identified but not consulted

Assumptions

Analysis

This item relies mainly on published pricing schedules, governance frameworks, and software-engineering productivity studies rather than anecdotal case studies. [fact; source: https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing; https://doi.org/10.6028/NIST.AI.100-1; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://arxiv.org/abs/2302.06590]

Token spend remains small in the modeled workload, while recurring human review, governance, and failure-recovery work account for most annual operating cost in the agent-workaround scenarios used here. [inference; source: https://dora.dev/research/2024/ai-preview/; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review; https://doi.org/10.6028/NIST.AI.100-1; https://www.anthropic.com/pricing#api; https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing]

The software-gap-closure side is more assumption-sensitive, so the analysis keeps the logic transparent: loaded labor, productivity multiplier, capability-gap effort, and maintenance rate are all visible, and changing them shifts the breakeven month threshold rather than reversing the direction for common stable gaps. [inference; source: https://www.bls.gov/oes/2023/may/oes151252.htm; https://www.bls.gov/news.release/ecec.t01.htm; https://www.microsoft.com/en-us/research/publication/the-effects-of-generative-ai-on-high-skilled-work-evidence-from-three-field-experiments-with-software-developers/; https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development]

The main competing interpretation is that model quality will keep rising fast enough to erase oversight cost, but the reviewed DORA and GitHub evidence does not support that today because organizations still route consequential changes through human review and still report reliability trade-offs from adoption. [inference; source: https://dora.dev/research/2024/dora-report/; https://dora.dev/research/2024/ai-preview/; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/kick-off-a-task; https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/copilot-code-review]

The second competing interpretation is to keep recurring agent operation indefinitely because it avoids the up-front project, but DORA's return-on-investment framing and prior repository architecture work both indicate that durable value comes from converting repeated workaround effort into a governed capability, not from paying the same workaround tax forever. [inference; source: https://cloud.google.com/resources/content/dora-roi-of-ai-assisted-software-development; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-13-agent-process-reliability-architecture.md; https://github.com/davidamitchell/Research/blob/main/Research/completed/2026-05-09-hybrid-architecture-probabilistic-llm-deterministic-governance.md]

Risks, Gaps, and Uncertainties

Open Questions

  • What is the empirical maintenance ratio for production agent workarounds after twelve months in stable enterprise use, broken down by review, operations, and incident recovery?
  • How different are the breakeven thresholds for low-frequency but high-value work compared with high-frequency transactional work?
  • What project-history datasets could replace the six-month, nine-month, and twelve-month debt-class archetypes with observed delivery distributions?
  • How much option value does a deliberate recurring-agent pilot create when it is used to discover the right target-state system design rather than as an indefinite workaround?

Output


Output

Navigation

Home

By Tag

bureaucracy

change-management

coase

constraint-analysis

control-model

decision-rights

delegation

delivery-risk

demand-segmentation

enterprise

exception-handling

execution

flow

flow-design

flow-metrics

governance

governance-patterns

incentives

instability

institutional-economics

leading-indicators

operating-model

organisation

organisational-design

queue-design

queueing

regulated-enterprise

routing

throughput

throughput-risk

transaction-costs

triage

williamson

Clone this wiki locally