Skip to content

2026 04 30 se fundamentals ai code synthesis

github-actions[bot] edited this page May 1, 2026 · 1 revision

Software Engineering fundamentals and AI code generation: a synthesis of evidence, proposed insights, and follow-up research directions

Research Question

Drawing on the planned seven-item research programme on Software Engineering (SE) fundamentals in Artificial Intelligence (AI)-augmented development, six completed primary items plus external anchors for the missing Ubiquitous Language (UL) dimension, covering structured alignment (Grill-Me), code entropy and quality metrics, deep modules and architectural design, UL, Test-Driven Development (TDD) and feedback loops, strategic versus tactical roles, and empirical comparisons of fundamentals-first versus specs-to-code workflows, what is the overall relationship between traditional SE fundamentals and the effectiveness, reliability, and long-term maintainability of AI-generated code, and what are the key proposed insights and priority follow-up research directions?

Scope

In scope:

  • Cross-item synthesis: integrating findings from the six completed primary items into a coherent overall model of how SE fundamentals interact with AI coding workflows, while using external anchors to qualify the missing UL dimension
  • Proposed insights: articulating the key claims that emerge from the combined evidence; clearly labelled by confidence level and evidence base
  • Relationship mapping: identifying how the seven practice areas reinforce, conflict with, or are independent of each other; which practices are prerequisites for others; which combinations produce super-additive effects
  • Comparison with "specs-to-code" baseline: characterising the overall gap between fundamentals-first and prompt-only approaches across all dimensions studied
  • Practical recommendations: what should a team or individual developer adopt first, second, and third, given different project sizes, timelines, and domains?
  • Follow-up research directions: gaps in the evidence base that warrant new primary research; hypotheses generated by the synthesis that could be tested empirically
  • Connection to related completed research in the corpus: any completed items on AI coding, software quality, or SE practices that add relevant context

Out of scope:

  • New primary research, this item synthesises existing completed items and a small external anchor set; it does not run a new primary experiment
  • Domains outside software development, Artificial Intelligence (AI) in medicine, law, finance, unless analogous patterns are directly instructive
  • Detailed mechanics of individual practices, each primary item covers its own domain; this item works at the level of cross-cutting insights

Constraints:

  • The original design expected all seven primary items to be completed, but this synthesis proceeds with six completed items because 2026-04-30-ubiquitous-language-ai-code-consistency remains backlog and its dimension is therefore downgraded to an external-anchor synthesis
  • All claims in this synthesis must be traceable to specific findings in the cited primary items or to additional external evidence gathered during synthesis
  • item_type: synthesis, governed by the synthesis item protocols in ADR-0013

Context

This synthesis is the capstone of a planned seven-item research programme on how traditional Software Engineering (SE) fundamentals interact with AI-augmented code generation. At synthesis time, six primary items are completed and one UL-focused primary item remains backlog, so the shared-vocabulary dimension is qualified with external theory and workflow guidance rather than a completed corpus item. [fact; source: https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md; https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]

The synthesis is designed to integrate the completed findings into one overall model, surface cross-cutting insights that do not appear in the individual items, distinguish stronger evidence from more speculative synthesis claims, and identify the most decision-changing follow-up research directions. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html]

The synthesis item item_type: synthesis follows the frontmatter and workflow rules defined in ADR-0013, which distinguishes synthesis items from primary research and adds the cites, related, confidence, and versions fields used here. [fact; source: https://github.com/davidamitchell/Research/blob/main/docs-adr/0013-research-item-frontmatter-schema-extension.md]

Approach

This synthesis is conducted by the research agent after the planned primary corpus is sufficiently mature to support cross-item conclusions. Six direct primary items are completed; the missing UL item is treated as an explicit evidence gap. The approach is:

  1. Load and summarise the completed primary items: Extract the Key Findings, Evidence Map, and confidence levels from each completed primary item. Identify where evidence is high, medium, or low confidence.

  2. Map relationships: Identify dependencies, reinforcing combinations, and potential conflicts between the practice areas. Is TDD a prerequisite for deep modules to work well with AI? Does Grill-Me alignment produce better inputs for deep-module design? Construct a relationship map.

  3. Identify cross-cutting insights: What patterns emerge across multiple items? What does the combined evidence suggest that no single item could establish alone?

  4. Compare with specs-to-code baseline: Synthesise the overall gap between fundamentals-first and prompt-only workflows across all dimensions. What is the net evidence on whether fundamentals-first produces materially better outcomes?

  5. Produce practical recommendations: Given evidence strength and adoption cost, what is the priority order for a team or individual adopting fundamentals-first practices? What is the minimum viable fundamentals stack?

  6. Identify follow-up research: What are the most important unanswered questions? Which hypotheses generated by the primary items could be tested empirically? What would change the key conclusions if the evidence were updated?

  7. Write and validate the synthesis: Apply the research skill §§0-7 process to the synthesis task. Flag every cross-cutting claim with its evidence provenance. Recursive review at §7.

Sources

(Sources for this synthesis are the completed primary items and any additional external evidence gathered during the synthesis phase.)


Research Skill Output

(Full output from running the research skill, retained verbatim in the completed item. §§0-5 are the investigation; §6 seeds the Findings section below.)

§0 Initialise

  • Restated question: what does the combined evidence say about whether traditional Software Engineering (SE) fundamentals improve the effectiveness, reliability, and long-term maintainability of Artificial Intelligence (AI)-generated code, and which adoption priorities and follow-up studies matter most? [fact; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html]
  • Scope confirmed: this item is a synthesis of the completed SE-fundamentals corpus plus a small external anchor set for theory and workflow guidance, not a new primary experiment. [fact; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://github.com/davidamitchell/Research/blob/main/docs-adr/0013-research-item-frontmatter-schema-extension.md]
  • Constraint confirmed: six planned primary inputs are completed and directly usable, while the dedicated Ubiquitous Language (UL) primary item remains in backlog, so the shared-vocabulary dimension is synthesized from external sources and companion-item references rather than from a completed seventh primary item. [fact; source: https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html]
  • Output format: structured synthesis with executive summary, key findings, evidence map, assumptions, analysis, risks, and open questions, then a mirrored Findings section. [fact; source: https://github.com/davidamitchell/Research/blob/main/docs-adr/0013-research-item-frontmatter-schema-extension.md]

§1 Question Decomposition

  • Q1. What overall relationship does the evidence support between SE fundamentals and AI-generated code outcomes?
    • Q1.1. Do fundamentals help mostly through ambiguity reduction before generation?
    • Q1.2. Do fundamentals help mostly through verification and control after generation?
    • Q1.3. Are repository-scale maintainability effects different from bounded task effects?
  • Q2. How do the practice areas interact?
    • Q2.1. How do clarification and shared vocabulary affect prompt quality?
    • Q2.2. How do explicit interfaces and deep modules affect delegation scope?
    • Q2.3. How do Test-Driven Development (TDD) and fast feedback affect self-correction and review load?
  • Q3. What does the evidence imply about human versus AI role division?
  • Q4. What is the minimum viable fundamentals stack, and in what adoption order should teams add it?
  • Q5. Which claims are well-supported, which remain inference-level, and what follow-up research would most change the conclusion?

§2 Investigation

Prior completed-item cross-reference

  • [fact] The six completed primary items converge on the same control surfaces, clarification before generation, explicit constraints around generation, and verifier strength after generation, even though each item studies a different mechanism and evidence base. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]
  • [fact] Two adjacent completed items sharpen the same governance surface from outside the seven-item bundle, namely that AI output volume is already outpacing human verification capacity and that code is comparatively safer than many other Large Language Model (LLM) outputs because it can be externally verified before deployment. [source: https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://davidamitchell.github.io/Research/research/2026-04-26-llm-verifiability-asymmetry-code-world-action.html]

Overall relationship

  • [fact] The fundamentals-first versus specs-to-code item concludes that fundamentals-first workflows outperform pure prompt-to-code once generated code must survive review, debugging, and ongoing change, even though prompt-only workflows often win on initial prototype speed. [source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]
  • [fact] The Grill-Me item finds the strongest direct benefit from clarification-before-code on ambiguous tasks, with measured first-pass correctness gains in ClarifyGPT and supporting evidence from ClariGen and requirements-discovery analogues. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html]
  • [fact] The TDD item finds that explicit execution feedback, failing tests, and small validated increments give AI coding a stronger self-correction loop than bulk generation with delayed verification. [source: https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html]
  • [fact] The deep-modules item argues that explicit interfaces and bounded module surfaces improve delegation safety by localizing the amount of design knowledge each change requires. [source: https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://web.stanford.edu/~ouster/cgi-bin/book.php]
  • [fact] The entropy item distinguishes bounded task gains from repository-scale deterioration, showing that local code quality can improve while clone share, refactoring decline, hotspot weakness, and change-coupling risk worsen over time. [source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]
  • [fact] The strategic-versus-tactical item concludes that the best-supported operating model keeps humans responsible for architecture, context, clarification, interface definition, and verification policy while Artificial Intelligence (AI) executes bounded implementation tasks. [source: https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html]
  • [inference] Across the corpus, SE fundamentals improve AI-generated code primarily by narrowing the search space before generation and by strengthening verifiers and boundaries after generation, rather than by changing model fluency directly. [source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]

Relationship map

  • [inference] Grill-Me and shared vocabulary operate on the same pre-generation control surface, because both reduce ambiguity and hidden synonym drift before the model writes code. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/]
  • [inference] Shared vocabulary and deep modules reinforce each other, because a codebase with stable domain names and small explicit interfaces leaks less hidden context across files and sessions. [source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://web.stanford.edu/~ouster/cgi-bin/book.php; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]
  • [inference] TDD and fast feedback are the main post-generation control surface, because they convert AI coding from bulk emission into verifier-paced search and keep reviewable increments small. [source: https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]
  • [inference] Strategic versus tactical role division is not a separate practice from the others; it is the organisational consequence of them, because clarification, naming, interfaces, and verification are exactly the artifacts humans retain ownership of when AI is delegated bounded implementation. [source: https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/]
  • [inference] The entropy findings explain why the other practices compound, because weak clarification, weak naming, weak interfaces, and weak verification all manifest downstream as review overload, duplication, and rising cost of change. [source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-14-reliable-software-llm-era.html]

Shared-vocabulary and Ubiquitous Language dimension

  • [fact] Eric Evans defines Ubiquitous Language as a shared domain vocabulary used consistently across conversation, models, and code so that the software structure reflects the evolving understanding of the problem domain. [source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/]
  • [fact] GitHub and Anthropic both recommend precise language, repository rules, and explicit context scoping as prerequisites for reliable agentic coding, which provides direct workflow support for the shared-vocabulary mechanism even though neither source isolates domain glossaries experimentally. [source: https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices]
  • [fact] Matt Pocock's public workflow material treats shared repository instructions, explicit return types, and repeated skill patterns as aids to future AI understanding, which is directionally consistent with a living glossary practice. [source: https://www.aihero.dev/5-agent-skills-i-use-every-day; https://github.com/mattpocock/skills]
  • [inference] The UL dimension is probably beneficial for AI coding because domain-stable names should reduce synonymous prompt phrasing and make repository context more compressible, but the claim stays below high confidence because the dedicated primary item was not completed and no strong longitudinal naming-drift study was available in this session. [source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md]
  • [fact] The seeded source "Trebing et al. (2023) Evaluating the Consistency of LLMs in Code Generation" could not be verified in this session, because arXiv 2307.10680 resolves to an unrelated recommender-systems paper and title or author searches did not locate an accessible primary source for the claimed study. [source: https://arxiv.org/abs/2307.10680; https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md]

Practical synthesis and follow-up directions

  • [inference] The minimum viable fundamentals stack is clarification-first discovery, then shared vocabulary or glossary discipline, then executable verification, then interface and architecture hardening, because each later layer becomes cheaper when the earlier ambiguity has already been reduced. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]
  • [inference] The most decision-changing follow-up research would be longitudinal, controlled comparisons of full fundamentals-first and prompt-only teams, because that is the missing evidence needed to convert several current medium-confidence synthesis claims into higher-confidence operational guidance. [source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html]

§3 Reasoning

  • [inference] I weighted the corpus's direct outcome evidence above practitioner rhetoric, then used theory and workflow guidance mainly to explain mechanisms that the completed items had already surfaced. [source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/]
  • [inference] I treated bounded-task performance and repository-scale maintainability as distinct layers, because several apparent contradictions in the corpus disappear once timescale is separated from task-local correctness. [source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]
  • [inference] I capped the shared-vocabulary contribution below high confidence because the mechanism is plausible and well-grounded in Domain-Driven Design (DDD) and context-engineering theory, but the planned primary item for that dimension was not completed. [source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md]
  • [inference] I treated adjacent completed items on volume versus correctness and verifiability asymmetry as sharpening evidence rather than as substitutes for the six direct primary items, because they qualify the same control surface without answering the exact same question. [source: https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://davidamitchell.github.io/Research/research/2026-04-26-llm-verifiability-asymmetry-code-world-action.html]

§4 Consistency Check

  • [fact] The corpus is internally consistent on the central pattern that ambiguity reduction and external verification improve outcomes, while weak structure and weak review cause downstream cost, even though the items differ on whether the strongest evidence is direct or indirect. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html]
  • [inference] The seeming contradiction between "AI improves local quality" and "AI increases system entropy" is resolvable because the first claim is about bounded tasks under explicit evaluation and the second is about repository accumulation under weaker structural controls. [source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]
  • [fact] No reviewed source in this session showed that prompt-only workflows dominate fundamentals-first workflows on persistent codebases once review load, debugging, and future change are counted together. [source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]
  • [fact] Confidence remains medium overall because several key findings are cross-item syntheses supported by companion evidence and mechanism logic rather than by a single primary longitudinal field study. [source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html]

§5 Depth and Breadth Expansion

  • [inference] Technical lens: the strongest role of fundamentals is control of information flow, namely clarifying intent, stabilizing names, shrinking context surfaces, and creating executable acceptance criteria. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://web.stanford.edu/~ouster/cgi-bin/book.php; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html]
  • [inference] Economic lens: fundamentals shift cost left, because they add up-front clarification and structure to reduce later review debt, rescue work, and change friction that otherwise arrive after AI has already amplified output volume. [source: https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://www.oreilly.com/library/view/tidy-first/9781098151232/]
  • [inference] Behavioural lens: the practices matter partly because human reviewers over-accept locally plausible generated code, so fundamentals create deliberate pause points where hidden assumptions can be challenged before more code accumulates. [source: https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]
  • [inference] Organisational lens: the strategist-builder split is a governance design, not only an efficiency choice, because it preserves human ownership of the control artifacts that make AI output safe to integrate. [source: https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-04-26-llm-verifiability-asymmetry-code-world-action.html]
  • [inference] Historical lens: the fundamentals gaining leverage in the AI era are mostly old SE disciplines, not novel model-era inventions, which implies that the technology changed the throughput profile faster than it changed the laws of maintainable system design. [source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://web.stanford.edu/~ouster/cgi-bin/book.php; https://www.oreilly.com/library/view/tidy-first/9781098151232/]

§6 Synthesis

(This section seeds the Findings below.)

Executive summary:

  • Traditional Software Engineering (SE) fundamentals improve Artificial Intelligence (AI)-generated code primarily by reducing ambiguity before generation and by adding external verification and boundary structure after generation, so they change the workflow's control system more than the model's raw fluency. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]
  • Prompt-only workflows remain faster for disposable prototypes, but the combined evidence favors fundamentals-first once generated code must survive review, debugging, and repeated change inside a maintained repository. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]
  • A practical stack suggested by the evidence combines clarification-first discovery, shared vocabulary discipline, executable verification, and explicit interfaces or deep modules, while the exact rollout sequence still depends on project context and existing weaknesses. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]
  • Confidence is medium because the six completed primary items are mutually reinforcing, but the dedicated UL primary item was not completed and the strongest remaining gaps are longitudinal, whole-project comparisons of full fundamentals-first and prompt-only teams. [inference; source: https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]

Key findings:

  1. Traditional Software Engineering (SE) fundamentals help AI-generated code mainly by reducing ambiguity and constraining search around the model, which is why the strongest gains appear when clarification, interfaces, and tests are all present together. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html)
  2. Prompt-only or specs-to-code workflows retain a real speed advantage on bounded prototyping tasks, but the evidence no longer supports them as the best default for persistent codebases once review cost, debugging burden, and future change are included. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html)
  3. Clarification-first discovery is one of the best-supported first controls under ambiguity, because direct evidence shows that targeted questioning before code generation improves first-pass correctness and reduces later correction rounds. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html)
  4. Shared vocabulary or glossary discipline is a plausible pre-generation control, because stable domain names should reduce synonymous prompt phrasing and session-to-session naming drift, although this conclusion remains partly inferential because the planned UL primary item is still backlog. ([inference]; low confidence; source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md)
  5. Executable verification through Test-Driven Development (TDD), fast tests, and runtime feedback is the strongest post-generation control in the corpus, because it turns AI coding into verifier-paced search and materially improves self-correction compared with one-shot generation. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-26-llm-verifiability-asymmetry-code-world-action.html)
  6. Explicit interfaces and deep modules make delegation safer by localizing the context each change requires, which limits hidden design leakage and reduces the chance that locally plausible code creates repository-scale entropy later. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://web.stanford.edu/~ouster/cgi-bin/book.php)
  7. The strongest current team operating model keeps humans responsible for architecture, context curation, vocabulary, interfaces, and verification policy, while AI performs bounded implementation inside those constraints, because that is where human attention still has the highest leverage. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/)
  8. The dominant system-level failure mode in the corpus is generation volume outpacing human verification and structural discipline, which is why the downstream signal appears first as review overload, duplication, and rising change cost instead of immediate total failure. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-14-reliable-software-llm-era.html)

Evidence map:

Claim Source Confidence Notes
[inference] Fundamentals help primarily by reducing ambiguity and constraining search around generation. https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html medium Cross-item mechanism convergence
[inference] Prompt-only workflows lose relative advantage once maintenance work begins. https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html medium Prototype speed differs from repository economics
[inference] Clarification-first discovery is one of the best-supported first controls under ambiguity because it improves first-pass correctness on ambiguous tasks. https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html medium Strongest direct causal evidence in bundle
[inference] Shared vocabulary likely reduces naming drift and ambiguity, but direct corpus evidence is incomplete. https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md low Missing completed primary item limits confidence
[inference] Executable verification is the strongest post-generation control surface. https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-26-llm-verifiability-asymmetry-code-world-action.html medium Direct item plus deployment-boundary companion
[inference] Deep modules and explicit interfaces lower delegation risk by localizing context. https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://web.stanford.edu/~ouster/cgi-bin/book.php medium Theory plus repository-scale consequence
[inference] Humans should own strategic control artifacts while AI handles bounded implementation. https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/ medium Team-operating-model synthesis
[inference] The central scaling risk is volume outrunning verification and structural discipline. https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-14-reliable-software-llm-era.html medium Cross-item governance and telemetry surface

Assumptions:

  • [assumption] The missing dedicated UL primary item would probably sharpen, not reverse, the shared-vocabulary conclusion, because the remaining external and companion evidence is directionally aligned. Justification: the mechanism already appears in Domain-Driven Design (DDD), context-engineering guidance, and multiple completed companion items. [source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices]
  • [assumption] The six completed primary items are sufficiently representative of the fundamentals-first bundle to support an overall synthesis even though one planned dimension is incomplete. Justification: the same core control surfaces recur across the six completed items and the adjacent companion items. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]
  • [assumption] Combining ambiguity reduction, verifier hardening, and architecture hardening likely compounds benefits, because downstream controls are cheaper when upstream ambiguity is already reduced. Justification: each completed item describes costs that rise when earlier control surfaces are weak. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]

Analysis:

  • The synthesis weighs direct empirical outcome evidence most heavily where available, which is why clarification gains and verifier-feedback gains sit at the center of the final model rather than at the margin. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html]
  • Repository-scale evidence matters even though it is more observational, because the research question explicitly asks about long-term maintainability and that outcome cannot be inferred from bounded benchmark wins alone. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]
  • The most important trade-off is front-loaded discipline versus downstream rework, because fundamentals-first practices slow the first move but reduce the volume of ambiguous, weakly verified, or weakly structured code that later has to be understood and repaired. [inference; source: https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://www.oreilly.com/library/view/tidy-first/9781098151232/; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html]
  • A rival explanation is that fundamentals-first teams may simply be more mature, use stronger tools, or work in more disciplined codebases than prompt-only teams, and the current evidence does not fully isolate those factors from the workflow effects described here. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html]
  • The evidence supports a layered recommendation rather than a single silver bullet, because ambiguity reduction, shared vocabulary, verification, and interface design each solve different failure mechanisms and reinforce one another when combined. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]

Risks, gaps, uncertainties:

  • [fact; source: https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md] The planned dedicated UL primary item was not completed, so the shared-vocabulary contribution is less directly evidenced than the other six dimensions.
  • [fact; source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html] Repository-scale maintainability evidence remains partly observational and does not cleanly randomize entire codebases into alternative AI workflow conditions.
  • [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html] The strongest team-operating-model conclusions are still synthesis-level rather than bundle-level experimental findings.
  • [inference; source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://www.anthropic.com/engineering/claude-code-best-practices] Shared-vocabulary and context-engineering sources provide a plausible mechanism, but they do not yet provide a strong longitudinal naming-drift benchmark for AI-assisted repositories.

Open questions:

  • What does a twelve-month controlled comparison show for defect escape rate, review time, revert rate, and code-health decline in fundamentals-first versus prompt-only AI teams?
  • What is the minimum viable glossary artifact that captures most of the shared-vocabulary benefit without creating heavy maintenance overhead?
  • Which telemetry bundle best signals that a team should move from prototype-speed mode into fundamentals-first discipline, review time, hotspot decline, duplication, or change coupling?
  • How much of the observed advantage comes from the bundle effect, clarification plus vocabulary plus verification plus interfaces, versus from any single practice in isolation?

§7 Recursive Review

  • [fact] Every key claim in §6 is either directly grounded in completed-item findings or explicitly marked as inference or assumption. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]
  • [inference] Overall confidence remains medium because the synthesis is coherent and well-supported across six completed primary items, but the missing UL primary item and the absence of whole-project randomized workflow studies keep the strongest conclusions below high confidence. [source: https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]

Findings

(Populated from §6 Synthesis above.)

Executive Summary

Traditional Software Engineering (SE) fundamentals improve Artificial Intelligence (AI)-generated code primarily by reducing ambiguity before generation and by adding external verification and boundary structure after generation, so they change the workflow's control system more than the model's raw fluency. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]

Prompt-only workflows remain faster for disposable prototypes, but the combined evidence favors fundamentals-first once generated code must survive review, debugging, and repeated change inside a maintained repository. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]

A practical stack suggested by the evidence combines clarification-first discovery, shared vocabulary discipline, executable verification, and explicit interfaces or deep modules, while the exact rollout sequence still depends on project context and existing weaknesses. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]

Confidence is medium because the six completed primary items are mutually reinforcing, but the dedicated UL primary item was not completed and the strongest remaining gaps are longitudinal, whole-project comparisons of full fundamentals-first and prompt-only teams. [inference; source: https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]

Key Findings

  1. Traditional Software Engineering (SE) fundamentals help AI-generated code mainly by reducing ambiguity and constraining search around the model, which is why the strongest gains appear when clarification, interfaces, and tests are all present together. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html)
  2. Prompt-only or specs-to-code workflows retain a real speed advantage on bounded prototyping tasks, but the evidence no longer supports them as the best default for persistent codebases once review cost, debugging burden, and future change are included. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html)
  3. Clarification-first discovery is one of the best-supported first controls under ambiguity, because direct evidence shows that targeted questioning before code generation improves first-pass correctness and reduces later correction rounds. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html)
  4. Shared vocabulary or glossary discipline is a plausible pre-generation control, because stable domain names should reduce synonymous prompt phrasing and session-to-session naming drift, although this conclusion remains partly inferential because the planned UL primary item is still backlog. ([inference]; low confidence; source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md)
  5. Executable verification through Test-Driven Development (TDD), fast tests, and runtime feedback is the strongest post-generation control in the corpus, because it turns AI coding into verifier-paced search and materially improves self-correction compared with one-shot generation. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-26-llm-verifiability-asymmetry-code-world-action.html)
  6. Explicit interfaces and deep modules make delegation safer by localizing the context each change requires, which limits hidden design leakage and reduces the chance that locally plausible code creates repository-scale entropy later. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://web.stanford.edu/~ouster/cgi-bin/book.php)
  7. The strongest current team operating model keeps humans responsible for architecture, context curation, vocabulary, interfaces, and verification policy, while AI performs bounded implementation inside those constraints, because that is where human attention still has the highest leverage. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/)
  8. The dominant system-level failure mode in the corpus is generation volume outpacing human verification and structural discipline, which is why the downstream signal appears first as review overload, duplication, and rising change cost instead of immediate total failure. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-14-reliable-software-llm-era.html)

Evidence Map

Claim Source Confidence Notes
[inference] Fundamentals help primarily by reducing ambiguity and constraining search around generation. https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html medium Cross-item mechanism convergence
[inference] Prompt-only workflows lose relative advantage once maintenance work begins. https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html medium Prototype speed differs from repository economics
[inference] Clarification-first discovery is one of the best-supported first controls under ambiguity because it improves first-pass correctness on ambiguous tasks. https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html medium Strongest direct causal evidence in bundle
[inference] Shared vocabulary likely reduces naming drift and ambiguity, but direct corpus evidence is incomplete. https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md low Missing completed primary item limits confidence
[inference] Executable verification is the strongest post-generation control surface. https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-26-llm-verifiability-asymmetry-code-world-action.html medium Direct item plus deployment-boundary companion
[inference] Deep modules and explicit interfaces lower delegation risk by localizing context. https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://web.stanford.edu/~ouster/cgi-bin/book.php medium Theory plus repository-scale consequence
[inference] Humans should own strategic control artifacts while AI handles bounded implementation. https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://www.anthropic.com/engineering/claude-code-best-practices; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/ medium Team-operating-model synthesis
[inference] The central scaling risk is volume outrunning verification and structural discipline. https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-14-reliable-software-llm-era.html medium Cross-item governance and telemetry surface

Assumptions

  • [assumption] The missing dedicated UL primary item would probably sharpen, not reverse, the shared-vocabulary conclusion, because the remaining external and companion evidence is directionally aligned. Justification: the mechanism already appears in Domain-Driven Design (DDD), context-engineering guidance, and multiple completed companion items. [source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://github.blog/ai-and-ml/github-copilot/how-to-build-reliable-ai-workflows-with-agentic-primitives-and-context-engineering/; https://www.anthropic.com/engineering/claude-code-best-practices]
  • [assumption] The six completed primary items are sufficiently representative of the fundamentals-first bundle to support an overall synthesis even though one planned dimension is incomplete. Justification: the same core control surfaces recur across the six completed items and the adjacent companion items. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html]
  • [assumption] Combining ambiguity reduction, verifier hardening, and architecture hardening likely compounds benefits, because downstream controls are cheaper when upstream ambiguity is already reduced. Justification: each completed item describes costs that rise when earlier control surfaces are weak. [source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]

Analysis

The synthesis weighs direct empirical outcome evidence most heavily where available, which is why clarification gains and verifier-feedback gains sit at the center of the final model rather than at the margin. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html]

Repository-scale evidence matters even though it is more observational, because the research question explicitly asks about long-term maintainability and that outcome cannot be inferred from bounded benchmark wins alone. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html]

The most important trade-off is front-loaded discipline versus downstream rework, because fundamentals-first practices slow the first move but reduce the volume of ambiguous, weakly verified, or weakly structured code that later has to be understood and repaired. [inference; source: https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html; https://www.oreilly.com/library/view/tidy-first/9781098151232/; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html]

A rival explanation is that fundamentals-first teams may simply be more mature, use stronger tools, or work in more disciplined codebases than prompt-only teams, and the current evidence does not fully isolate those factors from the workflow effects described here. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html; https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html]

The evidence supports a layered recommendation rather than a single silver bullet, because ambiguity reduction, shared vocabulary, verification, and interface design each solve different failure mechanisms and reinforce one another when combined. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-grill-me-ai-alignment-shared-design.html; https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://davidamitchell.github.io/Research/research/2026-04-30-tdd-feedback-loops-ai-augmented-dev.html; https://davidamitchell.github.io/Research/research/2026-04-30-deep-modules-ai-augmented-codebases.html]

Risks, Gaps, and Uncertainties

  • [fact; source: https://github.com/davidamitchell/Research/blob/main/Research/backlog/2026-04-30-ubiquitous-language-ai-code-consistency.md] The planned dedicated UL primary item was not completed, so the shared-vocabulary contribution is less directly evidenced than the other six dimensions.
  • [fact; source: https://davidamitchell.github.io/Research/research/2026-04-30-ai-code-entropy-quality-metrics.html; https://davidamitchell.github.io/Research/research/2026-03-12-volume-vs-correctness-ai-era.html] Repository-scale maintainability evidence remains partly observational and does not cleanly randomize entire codebases into alternative AI workflow conditions.
  • [inference; source: https://davidamitchell.github.io/Research/research/2026-04-30-fundamentals-first-vs-specs-to-code.html; https://davidamitchell.github.io/Research/research/2026-04-30-strategic-tactical-division-ai-teams.html] The strongest team-operating-model conclusions are still synthesis-level rather than bundle-level experimental findings.
  • [inference; source: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/; https://www.anthropic.com/engineering/claude-code-best-practices] Shared-vocabulary and context-engineering sources provide a plausible mechanism, but they do not yet provide a strong longitudinal naming-drift benchmark for AI-assisted repositories.

Open Questions

  • What does a twelve-month controlled comparison show for defect escape rate, review time, revert rate, and code-health decline in fundamentals-first versus prompt-only AI teams?
  • What is the minimum viable glossary artifact that captures most of the shared-vocabulary benefit without creating heavy maintenance overhead?
  • Which telemetry bundle best signals that a team should move from prototype-speed mode into fundamentals-first discipline, review time, hotspot decline, duplication, or change coupling?
  • How much of the observed advantage comes from the bundle effect, clarification plus vocabulary plus verification plus interfaces, versus from any single practice in isolation?

Output

(Filled in on completion, what was produced as a result of this research?)

Navigation

Home

By Tag

bureaucracy

change-management

coase

constraint-analysis

control-model

decision-rights

delegation

delivery-risk

demand-segmentation

enterprise

exception-handling

execution

flow

flow-design

flow-metrics

governance

governance-patterns

incentives

instability

institutional-economics

leading-indicators

operating-model

organisation

organisational-design

queue-design

queueing

regulated-enterprise

routing

throughput

throughput-risk

transaction-costs

triage

williamson

Clone this wiki locally