-
Notifications
You must be signed in to change notification settings - Fork 0
2026 05 08 ai skill decay deskilling measurement interventions
To what degree does over-reliance on AI tools accelerate measurable skill decay in practitioners, and what interventions best preserve human capability without sacrificing productivity gains?
To what degree and through what mechanisms does over-reliance on Artificial Intelligence (AI) tools, particularly tools that can plan or act across multi-step workflows, accelerate measurable skill decay in verification, judgment, and domain expertise among practitioners? How does the "AI for speed" paradigm affect junior versus senior practitioners differently, and what interventions, including deliberate practice protocols, hybrid apprenticeship models, mandatory human challenge thresholds, and skill audits, best preserve human oversight competence in AI-assisted environments without sacrificing short-term efficiency gains?
In scope:
- Empirical evidence and longitudinal studies on AI-induced skill decay or deskilling in knowledge work, software engineering, and regulated professional domains
- Proxy metrics and measurement approaches for skill decay where longitudinal data is unavailable, for example fallback performance during AI outages and error-detection rates without AI assistance
- Differential impact of AI tool adoption on junior versus senior practitioners, including apprenticeship and knowledge-transfer implications
- Intervention designs that demonstrably preserve human oversight competence, including deliberate practice, mandatory challenge thresholds, competency audits, and hybrid training
- Analogous evidence from automation-induced deskilling in adjacent domains, including aviation and medicine
Out of scope:
- AI model capability improvement or training methodology
- Consumer-tier AI use outside enterprise or professional contexts
- Workforce planning or headcount optimisation
- Macro-level labour-market displacement effects
Constraints:
- Flag inferences clearly; longitudinal causal evidence is sparse and must be distinguished from cross-sectional correlation
- Expand all acronyms on first use
- Include analogous automation-deskilling evidence where direct AI evidence is limited
Existing repository research has already identified skill decay as one of four interacting enterprise failure modes in AI governance, and it has established that human oversight quality degrades under volume pressure. [fact; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html]
The remaining gap is specific empirical evidence on the rate and mechanisms of AI-induced skill decay, measurable proxies that organisations can track, the differential effect across seniority levels, and which intervention designs have been validated rather than merely proposed. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html]
-
Sub-question 1 - Degree and mechanisms of skill decay: What is the strongest empirical evidence for AI-induced skill decay in knowledge work and software engineering? Are there longitudinal studies, controlled experiments, or credible proxy metrics such as fallback performance, error-detection rates, or competency-audit scores?
-
Sub-question 2 - Differential impact by seniority: How does the "AI for speed" paradigm affect junior practitioners differently from senior ones? What do cognitive apprenticeship, the learning model in which novices observe expert reasoning, practice with scaffolding, and then assume progressively independent responsibility, and analogous automation transitions in aviation and medicine say about differential decay?
-
Sub-question 3 - Effective interventions: What deliberate practice protocols, periodic unaided drills, mandatory challenge thresholds, meaning workflows that require a human preliminary judgment before seeing or accepting AI advice, hybrid apprenticeship models, or skill-audit designs have been empirically tested for preserving human competence in AI-assisted environments? What is the evidence for each?
Starting points:
- Bainbridge (1983) Ironies of Automation - foundational automation argument; the DOI redirected to an access-controlled ScienceDirect page in this session
- Parasuraman and Riley (1997) Humans and Automation: Use, Misuse, Disuse, Abuse - foundational misuse taxonomy; the DOI returned 403 in this session, so downstream claims rely on accessible later reviews when quoted
- Carr (2014) The Glass Cage: Automation and Us - contextual book-length treatment; not used for core factual support
- Cabitza et al. (2017) Unintended Consequences of Machine Learning in Medicine - official article page was access-restricted in this session; downstream claims rely on accessible summaries when quoted
- Peng et al. (2023) The Impact of AI on Developer Productivity
- Macnamara et al. (2024) Does using artificial intelligence assistance accelerate skill decay and hinder skill development without performers' awareness?
- Shen and Tamkin (2026) How AI Impacts Skill Formation
- Goddard et al. (2012) Automation bias: a systematic review of frequency, effect mediators, and mitigators
- Schubert et al. (2023) Strategies to reduce automation bias in AI-based personnel preselection
- Casner et al. (2014) The Retention of Manual Flying Skills in the Automated Cockpit
- Collins, Brown, and Holum (1991) Cognitive Apprenticeship: Making Thinking Visible
- Zhu et al. (2026) Augmenting Clinical Decision-Making with an Interactive and Interpretable AI Copilot
- Williams et al. (2026) Flight rules for clinical AI: lessons from aviation for human-AI collaboration
- Mitchell (2026) What capability and control design is needed to mitigate incentive misalignment, shadow AI, rail bypass, and skill decay at enterprise scale?
- Mitchell (2026) What is the evidence for human oversight as an effective quality gate in AI-assisted software development?
- Mitchell (2026) How should human-in-the-loop design be adapted when AI review volume makes human reviewers a bottleneck or causes rubber-stamping?
(Full output from running the research skill - retained verbatim in the completed item. Sections 0-5 are the investigation; Section 6 seeds the Findings section below.)
- Question: To what degree does over-reliance on AI tools measurably erode practitioner capability, by what mechanisms does that happen, how does the effect differ for juniors and seniors, and which interventions preserve capability without giving up most productivity gains?
- Scope: Knowledge work, software engineering, and regulated professional settings, with aviation and medicine used as analogical evidence where direct AI evidence is limited.
- Constraints: Publicly accessible sources only in this session, explicit [fact] and [inference] labeling, acronym expansion on first use, and direct separation between assisted-task productivity evidence and unaided-competence evidence.
- Output: knowledge.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html] Prior completed items already establish that skill decay is an enterprise governance risk, that human review is still a meaningful quality mechanism, and that high review volume degrades verification quality.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html] This item therefore focuses on the missing empirical layer: what can actually be measured today, how strongly the evidence supports differential junior-senior risk, and which interventions have evidence beyond general managerial intuition.
Q: To what degree does over-reliance on AI accelerate measurable skill decay, and what preserves human capability?
├── Q1: What direct AI-era evidence shows measurable capability loss or hindered skill formation?
│ ├── Q1a: What do theory and review papers say about hidden performance loss and unrecognised deskilling?
│ ├── Q1b: What controlled experiments show weaker learning or lower unaided performance after AI use?
│ └── Q1c: What field evidence shows worse non-AI performance after routine AI exposure?
├── Q2: Which skills decay first, and through what mechanisms?
│ ├── Q2a: Are verification, debugging, anomaly detection, and situational awareness more vulnerable than basic execution?
│ ├── Q2b: What role do trust, workload, time pressure, and cognitive offloading play?
│ └── Q2c: What measurable proxies can organisations use when long longitudinal datasets are absent?
├── Q3: How do juniors and seniors differ?
│ ├── Q3a: What evidence shows novices use AI as scaffolding rather than challenge material?
│ ├── Q3b: What evidence shows experts challenge AI differently, yet still remain vulnerable to atrophy?
│ └── Q3c: What do apprenticeship and deliberate-practice models imply about "never-skilling" risk?
├── Q4: Which interventions have direct empirical support?
│ ├── Q4a: Do error briefings, less aggregated evidence, or forced verification reduce automation bias?
│ ├── Q4b: Do periodic manual drills or AI-off practice preserve fallback competence?
│ └── Q4c: Which intervention elements preserve productivity while limiting deskilling?
└── Q5: What operating model follows for organisations using AI for speed?
├── Q5a: Which tasks can be safely accelerated?
├── Q5b: Which roles require apprenticeship plus fading rather than full delegation?
└── Q5c: Which metrics can expose hidden deskilling before failure occurs?
- [assumption; source: https://doi.org/10.1016/0005-1098(83)90046-8] Access note: the Bainbridge DOI redirected to an access-controlled ScienceDirect page in this session, so it was checked as a seeded source but not used as stand-alone support for downstream factual claims.
- [assumption; source: https://doi.org/10.1177/001872089703900402] Access note: the Parasuraman and Riley DOI returned 403 in this session, so downstream claims use accessible later reviews that summarize the misuse and automation-bias literature.
- [assumption; source: https://jamanetwork.com/journals/jama/fullarticle/2645749] Access note: the official Journal of the American Medical Association page for Cabitza et al. returned 403 in this session, so downstream claims rely on accessible secondary summaries and later reviews when that argument is used.
- [assumption; source: https://www.amazon.com/Glass-Cage-Automation-Nicholas-Carr/dp/0393240762] Access note: Nicholas Carr's book page was accessible only as a commercial listing in this session, so it was treated as contextual background rather than evidentiary support.
- [assumption; source: https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/fulltext] Failed primary-source search record: query
"Budzyn endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy"identified the official Lancet Gastroenterology and Hepatology landing page, but direct full-text retrieval was blocked in this session, so the downstream claim uses an accessible Nature Digital Medicine perspective that explicitly summarizes the study. - [assumption; source: https://doi.org/10.1177/001872089703900402] Failed primary-source search record: query
"Parasuraman Riley 1997 humans and automation use misuse disuse abuse pdf"did not locate an accessible full text in this session, so the item keeps Parasuraman and Riley as a checked seed but anchors factual downstream claims to accessible reviews instead.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html] The nearest completed items already connect skill decay to enterprise control failure, argue that human oversight remains a scarce verification asset, and show that high-volume review collapses into nominal sign-off without stronger control design.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html] The gap left by prior repository work is empirical specificity: organisations still need measurable indicators, stronger evidence on who is most exposed, and a defensible list of interventions that preserve oversight competence rather than only recommend caution.
- [fact; source: https://par.nsf.gov/biblio/10545734-does-using-artificial-intelligence-assistance-accelerate-skill-decay-hinder-skill-development-without-performers-awareness] Macnamara et al. argue that AI assistants can accelerate skill decay among experts, hinder skill acquisition among learners, and hide those effects from the users themselves, producing a form of hidden performance loss rather than an immediately visible productivity-cost trade-off.
- [fact; source: https://arxiv.org/html/2601.20245v1] Shen and Tamkin report randomized experiments in software engineering showing that AI assistance reduced evaluation scores on conceptual understanding, code reading, and debugging by 17% when participants learned a new asynchronous Python library.
- [fact; source: https://arxiv.org/html/2601.20245v1] The same study did not find a statistically significant improvement in completion time on average, which means the skill-loss result cannot be dismissed as a simple trade of learning for large universal speed gains.
- [fact; source: https://arxiv.org/html/2601.20245v1] Shen and Tamkin also found that participants who fully delegated coding tasks to the AI showed some local productivity improvements but learned less about the new library, while cognitively engaged usage patterns, such as asking for explanations or conceptual guidance, preserved more learning.
- [fact; source: https://www.nature.com/articles/s41746-026-02410-1] Williams et al. summarize a 2025 multicentre colonoscopy study showing that endoscopists who regularly used AI-assisted detection later performed worse when the AI was removed, with adenoma detection falling by 6.0 percentage points in standard non-AI colonoscopies.
- [inference; source: https://par.nsf.gov/biblio/10545734-does-using-artificial-intelligence-assistance-accelerate-skill-decay-hinder-skill-development-without-performers-awareness; https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1] The strongest current direct evidence therefore supports medium confidence that over-reliance on AI can already produce measurable deterioration in unaided competence, but the evidence base is still concentrated in a small number of studies rather than in broad multi-year enterprise panels.
- [fact; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/] Goddard et al.'s systematic review defines automation bias as over-reliance on automated recommendations and finds across 74 studies that workload, task complexity, time constraint, trust, confidence, and experience all mediate that over-reliance.
- [fact; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/] The same review identifies two recurrent failure modes, commission errors that follow bad advice and omission errors that arise when users fail to act because the system did not prompt them.
- [fact; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf] Casner et al. found that airline pilots' basic instrument-scanning and manual-control skills remained mostly intact, but the cognitive tasks needed for manual flight, including navigation planning, tracking position without map automation, and recognizing instrument failures, degraded more noticeably under prolonged automation use.
- [fact; source: https://www.nature.com/articles/s41746-026-02410-1] Williams et al. argue that medicine is encountering the same pattern aviation saw earlier: routine automation can reduce vigilance and situational awareness, while rare failures expose that the most vulnerable capabilities are anomaly detection, contextual judgment, and fallback reasoning rather than rote execution alone.
- [inference; source: https://arxiv.org/html/2601.20245v1; https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf] Across software engineering, medicine, and aviation, the first capabilities at risk are the ones humans use to supervise automation, namely debugging, evidence checking, error recognition, and mental-model maintenance, rather than every low-level execution skill equally.
- [fact; source: https://arxiv.org/abs/2602.00726] Zhu et al. report that trust in a clinical AI copilot was actively constructed through verification, and that junior clinicians used the system as cognitive scaffolding while expert clinicians engaged in adversarial verification to challenge the AI's logic.
- [fact; source: https://arxiv.org/html/2601.20245v1] Shen and Tamkin frame software engineering as a setting where novice workers may compromise skill acquisition if they rely heavily on AI to complete unfamiliar tasks, and they note that less-experienced and lower-skilled workers often benefit most in short-run productivity studies even while long-run competence remains under-measured.
- [fact; source: https://www.aft.org/ae/winter1991/collins_brown_holum] Collins, Brown, and Holum describe cognitive apprenticeship as a learning model built on modeling, scaffolding, fading, coaching, and increasingly independent practice, with observation and guided struggle needed to build self-monitoring and correction skills.
- [inference; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/html/2601.20245v1; https://arxiv.org/abs/2602.00726] Junior practitioners are therefore more exposed to "never-skilling" risk because heavy delegation can bypass the independent error encounters and fading process that apprenticeship models require for durable skill formation.
- [inference; source: https://arxiv.org/abs/2602.00726; https://www.nature.com/articles/s41746-026-02410-1; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf] Senior practitioners are better positioned to challenge AI because they already possess stronger mental models, but the field and analogical evidence shows they can still lose fallback capability when routine practice of anomaly detection and manual recovery disappears.
- [fact; source: https://arxiv.org/abs/2302.06590] Peng et al. show that developers using GitHub Copilot completed a bounded JavaScript server task 55.8% faster than the control group in a controlled experiment.
- [fact; source: https://arxiv.org/abs/2302.06590] The same paper does not measure long-term skill retention, unaided competence, or fallback performance after the AI is removed.
- [inference; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1] Productivity studies and skill-decay studies answer different questions, so short-run speed gains cannot by themselves rebut evidence that heavy delegation weakens later debugging, conceptual understanding, or non-AI performance.
- [fact; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/] Goddard et al. identify training, explicit user accountability, updated confidence cues, advice placement, and providing information instead of bare recommendation as recurring mitigators of automation bias.
- [fact; source: https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full] Schubert et al. found that informing users about possible system errors increased verification-intensity indicators and improved objective decision quality, while merely reminding them of responsibility did not have the same effect.
- [fact; source: https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full] The same experiment found that less aggregated evidence presentation encouraged deeper checking, which supports design patterns that expose supporting evidence rather than only surface a dashboard recommendation.
- [fact; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf] Casner et al. conclude that retaining cognitive skills for manual flying depends on pilots remaining actively engaged while supervising automation, which supports periodic unaided drills and active cross-checking rather than passive monitoring.
- [fact; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/html/2601.20245v1] Apprenticeship and software-learning evidence converge on the need for scaffolding plus fading: support can accelerate progress, but durable competence requires progressively independent practice rather than indefinite delegation.
- [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://arxiv.org/html/2601.20245v1] The best-supported intervention bundle is therefore not "use less AI" in the abstract, but a combination of challenge-before-accept workflows, evidence-rich interfaces, periodic AI-off drills, and apprenticeship structures that fade support as competence increases.
- [inference; source: https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html] When direct longitudinal skill-retention data is absent, the most defensible current proxies are fallback performance when AI is unavailable, error-detection rates on intentionally seeded or naturally occurring system mistakes, verification-intensity measures, and performance on periodic unaided competency checks.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html] For software and enterprise workflows, that proxy set maps cleanly to code-review challenge rate, independent test or debugging performance, recovery speed during AI outages, and audit scores from no-AI exercises on critical tasks.
- [inference; source: https://arxiv.org/html/2601.20245v1; https://arxiv.org/abs/2302.06590] The core reasoning error to avoid is collapsing assisted productivity and retained competence into the same construct, because the strongest accessible software studies measure them separately and reach different conclusions.
- [inference; source: https://par.nsf.gov/biblio/10545734-does-using-artificial-intelligence-assistance-accelerate-skill-decay-hinder-skill-development-without-performers-awareness; https://www.nature.com/articles/s41746-026-02410-1; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf] A second reasoning constraint is that direct enterprise field evidence remains limited, so the item has to separate what is directly shown in AI-specific studies from what is inferred by mechanism transfer from aviation and medicine.
- [inference; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/abs/2602.00726] The junior-senior conclusion is also partly mechanistic rather than fully measured, because the best direct evidence is about how novices and experts verify differently, not a large panel proving exact differential decay rates over time.
- [fact; source: https://arxiv.org/html/2601.20245v1; https://arxiv.org/abs/2302.06590] There is no contradiction between the Shen and Tamkin result and the Copilot productivity experiment, because one studies learning a new library and later unaided evaluation while the other studies speed on a bounded implementation task.
- [fact; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full] There is no contradiction between recommending accountability and noting that responsibility reminders alone were weak, because the broader review literature treats accountability as one element in a larger design bundle, while Schubert et al. specifically show that error salience and richer evidence were stronger direct levers than responsibility reminders by themselves.
- [inference; source: https://www.nature.com/articles/s41746-026-02410-1; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf] The analogical evidence remains internally coherent: aviation and medicine both point to a pattern in which routine automation preserves some low-level execution but weakens the higher-order skills humans need when automation fails.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html] From a governance lens, skill decay is not only a training problem; it is a control-surface problem because organisations lose the human capacity needed to challenge, override, and recover from AI failure.
- [inference; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/html/2601.20245v1] From a learning-science lens, the critical loss is not just lower output quality but removal of productive struggle, because learners who never practice independent problem framing do not build the internal models needed for later oversight.
- [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full] From a human-factors lens, the most actionable intervention surface is interface and workflow design, because blind trust rises when systems compress evidence, hide uncertainty, and encourage approval faster than active verification.
- [inference; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1] From a software-delivery lens, the practical trade-off is not between speed and capability in the abstract, but between where to use AI for bounded acceleration and where to preserve independent practice because the task is still part of how engineers learn to supervise future automation.
Executive summary:
Over-reliance on AI already shows measurable capability loss in a small but credible set of direct studies, and the loss is concentrated in verification, debugging, anomaly detection, and fallback reasoning rather than in every low-level execution skill equally. [inference; source: https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf]
Short-run productivity gains do not refute that risk, because the main software productivity experiments measure assisted completion speed, while the strongest skill-formation evidence measures later unaided competence and finds weaker independent performance after heavy AI use. [inference; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1]
Junior practitioners face the larger risk because they are still building mental models and self-correction habits, while senior practitioners more often challenge AI adversarially but can still lose fallback competence when manual or diagnostic recovery is rarely practiced. [inference; source: https://arxiv.org/abs/2602.00726; https://www.aft.org/ae/winter1991/collins_brown_holum; https://www.nature.com/articles/s41746-026-02410-1]
The best-supported interventions are challenge-before-accept workflows, evidence-rich interfaces, periodic AI-off drills, and apprenticeship models that deliberately fade support as competence grows, rather than generic calls for "human oversight" without changes to workflow design. [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum]
Key findings:
- The strongest direct AI-era evidence shows that heavy reliance on AI can reduce later unaided competence, because randomized software experiments and field evidence from clinical AI both report weaker non-AI performance after routine AI assistance. ([inference]; medium confidence; source: https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1)
- The capabilities most exposed to decay are verification, debugging, anomaly detection, situational awareness, and fallback reasoning, because those are the skills humans use when automation fails and the skills that become less practiced under routine delegated execution. ([inference]; medium confidence; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1)
- Short-run AI productivity studies do not settle the deskilling question on their own, because they measure assisted completion speed on bounded tasks rather than retained competence, fallback performance, or independent problem-solving after the tool is removed. ([inference]; high confidence; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1)
- Junior practitioners are more vulnerable to never-skilling than senior practitioners, because apprenticeship and human-AI collaboration evidence show that novices use AI as scaffolding during skill formation while experts are more likely to challenge outputs against richer prior mental models. ([inference]; medium confidence; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/abs/2602.00726; https://arxiv.org/html/2601.20245v1)
- Senior practitioners are not immune to skill erosion, because the aviation and medical evidence shows that experienced operators can still lose manual or diagnostic recovery capability when routine automated support removes the need for active cross-checking. ([inference]; medium confidence; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1)
- A central mechanism is automation bias under trust, workload, and time pressure, because over-reliance rises when the system is usually right, evidence is compressed, and users do not need to generate or defend an independent judgment. ([inference]; medium confidence; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full)
- The most evidence-backed intervention bundle combines error-salience briefings, less aggregated evidence views, challenge-before-accept workflow steps, and periodic AI-off practice, because those are the interventions with direct experimental or operational support across the retrieved literature. ([inference]; medium confidence; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf)
- A cautious enterprise response is to pair bounded AI acceleration with skill audits such as fallback drills, seeded-error reviews, and periodic unaided assessments, while reserving apprenticeship tasks for progressive independence rather than full delegation. ([inference]; low confidence; source: https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html)
Evidence map:
| Claim | Source | Confidence | Notes |
|---|---|---|---|
| [inference] Heavy AI reliance can reduce later unaided competence in both software learning and clinical field performance. | https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1 | medium | small direct-study base |
| [inference] Verification, debugging, anomaly detection, and fallback reasoning decay before every low-level execution skill does. | https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1 | medium | cross-domain pattern |
| [inference] Productivity studies do not settle deskilling on their own because they measure assisted speed rather than retained unaided competence. | https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1 | high | scope-bound evidence |
| [inference] Juniors face larger never-skilling risk because AI can replace the guided struggle needed for apprenticeship. | https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/abs/2602.00726; https://arxiv.org/html/2601.20245v1 | medium | mechanism-led inference |
| [inference] Seniors retain stronger challenge capacity but can still lose fallback competence under routine automation. | https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1 | medium | fallback analogue |
| [inference] Automation bias under workload, trust, and compressed evidence presentation is a central mechanism behind observed over-reliance. | https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full | medium | review-plus-experiment base |
| [inference] Error briefings, evidence-rich views, challenge-before-accept steps, and AI-off drills are the strongest current intervention bundle. | https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf | medium | bundled intervention synthesis |
| [inference] A cautious enterprise response is to pair bounded AI acceleration with skill audits and progressive independence for apprenticeship tasks. | https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html | low | conservative operating synthesis |
Assumptions:
- [assumption; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1] Aviation and medicine transfer usefully to enterprise AI oversight because the shared mechanism is supervisory work under usually reliable automation with rare but consequential failure.
- [assumption; source: https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum] Skill formation during software-library learning is a reasonable analogue for other knowledge-work domains where users must build new concepts before they can verify AI-generated output independently.
- [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html] Same-repository completed items sharpen enterprise implications but are treated as supporting synthesis rather than as independent external evidence.
Analysis:
The evidence supports a narrower claim than "AI always deskills people." [inference; source: https://arxiv.org/html/2601.20245v1; https://arxiv.org/abs/2302.06590] The more defensible conclusion is that deskilling risk rises when AI replaces the exact cognitive work users still need later for supervision, debugging, or recovery, especially on unfamiliar tasks. [inference; source: https://arxiv.org/html/2601.20245v1; https://par.nsf.gov/biblio/10545734-does-using-artificial-intelligence-assistance-accelerate-skill-decay-hinder-skill-development-without-performers-awareness]
The junior-senior split is also more specific than a blanket statement that juniors always suffer and seniors always cope. [inference; source: https://arxiv.org/abs/2602.00726; https://www.aft.org/ae/winter1991/collins_brown_holum] Juniors are more exposed because AI can bypass the independent struggle that builds internal models, while seniors are less exposed on routine tasks but still vulnerable on rarely practiced fallback work. [inference; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1]
One competing interpretation says the real issue is not skill decay but simply poor workflow design. [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full] The retrieved evidence partly supports that view, which is why the recommended interventions focus on interface design, error salience, and deliberate practice rather than on banning AI use. [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://arxiv.org/html/2601.20245v1]
Another rival remedy is to rely on stronger models so that human capability matters less. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html] The current evidence does not justify that move, because the same studies that show higher speed also show bounded-task framing and leave fallback competence unresolved. [inference; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1]
Risks, gaps, uncertainties:
- [fact; source: https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1] Direct AI-era studies of measurable skill decay remain few, so confidence stays at medium even though the available findings point in a consistent direction.
- [fact; source: https://doi.org/10.1177/001872089703900402; https://doi.org/10.1016/0005-1098(83)90046-8] Two foundational automation papers were checked but not retrievable in full text here, which means their role in this item is contextual rather than claim-bearing.
- [fact; source: https://jamanetwork.com/journals/jama/fullarticle/2645749; https://www.nature.com/articles/s41746-026-02410-1] The medicine-specific deskilling argument partly depends on accessible summaries because the official 2017 Journal of the American Medical Association page was access-restricted in this session.
- [fact; source: https://arxiv.org/abs/2602.00726] The strongest junior-versus-senior AI-verification evidence is still small-sample qualitative or mixed-method work, so exact effect sizes by seniority remain uncertain.
- [fact; source: https://arxiv.org/html/2601.20245v1; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html] No retrieved source provides a universal enterprise metric set for capability loss, so the proposed proxy metrics remain a pragmatic synthesis rather than a published standard.
Open questions:
- [inference; source: https://arxiv.org/html/2601.20245v1; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html] Which software-engineering review metrics best distinguish healthy augmentation from hidden loss of debugging and verification skill at team scale?
- [inference; source: https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/] Which interface design preserves verification intensity best in enterprise AI tooling: preliminary answer capture, richer evidence packs, disagreement prompts, or peer-review pairing?
- [inference; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/abs/2602.00726] What is the best apprenticeship schedule for introducing AI to juniors without sacrificing the independent practice needed for durable expertise?
Output:
- Type: knowledge
- Description: Organisations should treat AI skill decay as a measurable control problem and preserve human capability through bounded delegation, challenge-before-accept workflow design, and periodic no-AI audits rather than through generic policy language about oversight alone. [inference; source: https://arxiv.org/html/2601.20245v1; https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html]
- Links:
- Claim-label audit: complete
- Acronym-expansion audit: complete
- Findings and §6 parity: aligned
- Adjacent completed-item sweep: repeated before synthesis
Over-reliance on AI already shows measurable capability loss in a small but credible set of direct studies, and the loss is concentrated in verification, debugging, anomaly detection, and fallback reasoning rather than in every low-level execution skill equally. [inference; source: https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf]
Short-run productivity gains do not refute that risk, because the main software productivity experiments measure assisted completion speed, while the strongest skill-formation evidence measures later unaided competence and finds weaker independent performance after heavy AI use. [inference; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1]
Junior practitioners face the larger risk because they are still building mental models and self-correction habits, while senior practitioners more often challenge AI adversarially but can still lose fallback competence when manual or diagnostic recovery is rarely practiced. [inference; source: https://arxiv.org/abs/2602.00726; https://www.aft.org/ae/winter1991/collins_brown_holum; https://www.nature.com/articles/s41746-026-02410-1]
The best-supported interventions are challenge-before-accept workflows, evidence-rich interfaces, periodic AI-off drills, and apprenticeship models that deliberately fade support as competence grows, rather than generic calls for human oversight without changes to workflow design. [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum]
- The strongest direct AI-era evidence shows that heavy reliance on AI can reduce later unaided competence, because randomized software experiments and field evidence from clinical AI both report weaker non-AI performance after routine AI assistance. ([inference]; medium confidence; source: https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1)
- The capabilities most exposed to decay are verification, debugging, anomaly detection, situational awareness, and fallback reasoning, because those are the skills humans use when automation fails and the skills that become less practiced under routine delegated execution. ([inference]; medium confidence; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1)
- Short-run AI productivity studies do not settle the deskilling question on their own, because they measure assisted completion speed on bounded tasks rather than retained competence, fallback performance, or independent problem-solving after the tool is removed. ([inference]; high confidence; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1)
- Junior practitioners are more vulnerable to never-skilling than senior practitioners, because apprenticeship and human-AI collaboration evidence show that novices use AI as scaffolding during skill formation while experts are more likely to challenge outputs against richer prior mental models. ([inference]; medium confidence; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/abs/2602.00726; https://arxiv.org/html/2601.20245v1)
- Senior practitioners are not immune to skill erosion, because the aviation and medical evidence shows that experienced operators can still lose manual or diagnostic recovery capability when routine automated support removes the need for active cross-checking. ([inference]; medium confidence; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1)
- A central mechanism is automation bias under trust, workload, and time pressure, because over-reliance rises when the system is usually right, evidence is compressed, and users do not need to generate or defend an independent judgment. ([inference]; medium confidence; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full)
- The most evidence-backed intervention bundle combines error-salience briefings, less aggregated evidence views, challenge-before-accept workflow steps, and periodic AI-off practice, because those are the interventions with direct experimental or operational support across the retrieved literature. ([inference]; medium confidence; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf)
- A cautious enterprise response is to pair bounded AI acceleration with skill audits such as fallback drills, seeded-error reviews, and periodic unaided assessments, while reserving apprenticeship tasks for progressive independence rather than full delegation. ([inference]; low confidence; source: https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html)
| Claim | Source | Confidence | Notes |
|---|---|---|---|
| [inference] Heavy AI reliance can reduce later unaided competence in both software learning and clinical field performance. | https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1 | medium | small direct-study base |
| [inference] Verification, debugging, anomaly detection, and fallback reasoning decay before every low-level execution skill does. | https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1 | medium | cross-domain pattern |
| [inference] Productivity studies do not settle deskilling on their own because they measure assisted speed rather than retained unaided competence. | https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1 | high | scope-bound evidence |
| [inference] Juniors face larger never-skilling risk because AI can replace the guided struggle needed for apprenticeship. | https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/abs/2602.00726; https://arxiv.org/html/2601.20245v1 | medium | mechanism-led inference |
| [inference] Seniors retain stronger challenge capacity but can still lose fallback competence under routine automation. | https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1 | medium | fallback analogue |
| [inference] Automation bias under workload, trust, and compressed evidence presentation is a central mechanism behind observed over-reliance. | https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full | medium | review-plus-experiment base |
| [inference] Error briefings, evidence-rich views, challenge-before-accept steps, and AI-off drills are the strongest current intervention bundle. | https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf | medium | bundled intervention synthesis |
| [inference] A cautious enterprise response is to pair bounded AI acceleration with skill audits and progressive independence for apprenticeship tasks. | https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html | low | conservative operating synthesis |
- Aviation and medicine transfer usefully to enterprise AI oversight because the shared mechanism is supervisory work under usually reliable automation with rare but consequential failure. [assumption; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1]
- Skill formation during software-library learning is a reasonable analogue for other knowledge-work domains where users must build new concepts before they can verify AI-generated output independently. [assumption; source: https://arxiv.org/html/2601.20245v1; https://www.aft.org/ae/winter1991/collins_brown_holum]
- Same-repository completed items sharpen enterprise implications but are treated as supporting synthesis rather than as independent external evidence. [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html]
The evidence supports a narrower claim than "AI always deskills people." [inference; source: https://arxiv.org/html/2601.20245v1; https://arxiv.org/abs/2302.06590] The more defensible conclusion is that deskilling risk rises when AI replaces the exact cognitive work users still need later for supervision, debugging, or recovery, especially on unfamiliar tasks. [inference; source: https://arxiv.org/html/2601.20245v1; https://par.nsf.gov/biblio/10545734-does-using-artificial-intelligence-assistance-accelerate-skill-decay-hinder-skill-development-without-performers-awareness]
The junior-senior split is also more specific than a blanket statement that juniors always suffer and seniors always cope. [inference; source: https://arxiv.org/abs/2602.00726; https://www.aft.org/ae/winter1991/collins_brown_holum] Juniors are more exposed because AI can bypass the independent struggle that builds internal models, while seniors are less exposed on routine tasks but still vulnerable on rarely practiced fallback work. [inference; source: https://www.faasafety.gov/files/events/SO/SO15/2025/SO15138466/the_retention_of_manual_flying_skills_in_the_automated_cockpit-NASA.pdf; https://www.nature.com/articles/s41746-026-02410-1]
One competing interpretation says the real issue is not skill decay but simply poor workflow design. [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full] The retrieved evidence partly supports that view, which is why the recommended interventions focus on interface design, error salience, and deliberate practice rather than on banning AI use. [inference; source: https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://arxiv.org/html/2601.20245v1]
Another rival remedy is to rely on stronger models so that human capability matters less. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html] The current evidence does not justify that move, because the same studies that show higher speed also show bounded-task framing and leave fallback competence unresolved. [inference; source: https://arxiv.org/abs/2302.06590; https://arxiv.org/html/2601.20245v1]
- Direct AI-era studies of measurable skill decay remain few, so confidence stays at medium even though the available findings point in a consistent direction. [fact; source: https://arxiv.org/html/2601.20245v1; https://www.nature.com/articles/s41746-026-02410-1]
- Two foundational automation papers were checked but not retrievable in full text here, which means their role in this item is contextual rather than claim-bearing. [fact; source: https://doi.org/10.1177/001872089703900402; https://doi.org/10.1016/0005-1098(83)90046-8]
- The medicine-specific deskilling argument partly depends on accessible summaries because the official 2017 Journal of the American Medical Association page was access-restricted in this session. [fact; source: https://jamanetwork.com/journals/jama/fullarticle/2645749; https://www.nature.com/articles/s41746-026-02410-1]
- The strongest junior-versus-senior AI-verification evidence is still small-sample qualitative or mixed-method work, so exact effect sizes by seniority remain uncertain. [fact; source: https://arxiv.org/abs/2602.00726]
- No retrieved source provides a universal enterprise metric set for capability loss, so the proposed proxy metrics remain a pragmatic synthesis rather than a published standard. [fact; source: https://arxiv.org/html/2601.20245v1; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html]
- Which software-engineering review metrics best distinguish healthy augmentation from hidden loss of debugging and verification skill at team scale? [inference; source: https://arxiv.org/html/2601.20245v1; https://davidamitchell.github.io/Research/research/2026-05-01-human-oversight-ai-software-development.html]
- Which interface design preserves verification intensity best in enterprise AI tooling: preliminary answer capture, richer evidence packs, disagreement prompts, or peer-review pairing? [inference; source: https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1118723/full; https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/]
- What is the best apprenticeship schedule for introducing AI to juniors without sacrificing the independent practice needed for durable expertise? [inference; source: https://www.aft.org/ae/winter1991/collins_brown_holum; https://arxiv.org/abs/2602.00726]
- Type: knowledge
- Description: Organisations should treat AI skill decay as a measurable control problem and preserve human capability through bounded delegation, challenge-before-accept workflow design, and periodic no-AI audits rather than through generic policy language about oversight alone. [inference; source: https://arxiv.org/html/2601.20245v1; https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/; https://davidamitchell.github.io/Research/research/2026-05-02-incentive-misalignment-shadow-ai-skill-decay-controls.html]
- Links:
Navigation
By Tag
bureaucracy
change-management
coase
constraint-analysis
control-model
decision-rights
delegation
- Q4: Decision rights that should move closer to execution
- Q5: Control model for the best throughput-risk trade-off
delivery-risk
- Operating model synthesis for split-authority delivery systems
- Q6: Leading indicators of instability in split-authority flow systems
demand-segmentation
enterprise
exception-handling
execution
flow
flow-design
flow-metrics
governance
- Operating model synthesis for split-authority delivery systems
- Q1: Dominant flow constraint in split-authority delivery systems
- Q2: Demand segmentation for fast-path vs controlled-path flow
- Q4: Decision rights that should move closer to execution
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
governance-patterns
incentives
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
instability
institutional-economics
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
leading-indicators
operating-model
organisation
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
organisational-design
queue-design
queueing
regulated-enterprise
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
routing
throughput
throughput-risk
transaction-costs
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
triage
- Q2: Demand segmentation for fast-path vs controlled-path flow
- Q3: Routing design that isolates exceptions from routine flow
williamson