Skip to content

2026 07 20 regulated safe to fail probing operating models

github-actions[bot] edited this page Aug 3, 2026 · 1 revision

Governance and operating models for safe-to-fail experimentation in regulated industries

Research Question

In highly regulated industries such as financial services, healthcare, and pharmaceuticals, how do organisations design governance structures, team models, and operating practices that enable safe-to-fail probing experiments in the Complex domain of the Cynefin framework, and what empirical relationship exists between the degree of structure in experiment tracking and prioritisation mechanisms and outcomes including (a) volume and diversity of experiments conducted, (b) quality and actionability of learned patterns, (c) regulatory compliance and risk incidents, and (d) organisational adoption of emergent innovations?

Scope

In scope:

  • Governance structures and decision-rights models for exploratory work in regulated environments, including sandboxes, innovation councils, embedded risk and compliance partners, and central versus federated ownership
  • Team models and operating practices used to run bounded experiments in complex problem spaces, including escalation paths, guardrails, approval loops, and learning cadences
  • Tracking and prioritisation mechanisms for experiments, ranging from lightweight logs to structured hypothesis backlogs, portfolio boards, and stage-gated review points
  • Empirical and case-study evidence from financial services, healthcare, pharmaceuticals, and adjacent regulated sectors on experiment throughput, learning quality, compliance performance, and innovation adoption
  • The role of the Cynefin framework and safe-to-fail probes as conceptual lenses for experimentation under uncertainty, including cases that use equivalent practices without Cynefin terminology

Out of scope:

  • Consumer-growth or advertising experimentation outside regulated or safety-critical contexts
  • Jurisdiction-specific legal advice for any single regulator or firm
  • Purely technical model-evaluation methods that do not change governance, team design, or operating practice
  • Mature scaling and productisation practices after an emergent innovation has already been standardised

Constraints: Prioritise public sources from 2018-2026, with preference for regulator guidance, regulatory sandbox documentation, peer-reviewed studies, and named enterprise case studies. Distinguish prescribed framework guidance from observed organisational practice. Treat Cynefin as a framing device, not proof on its own, and seek empirical evidence for the relationship between tracking structure and outcomes even when sources do not use Cynefin language explicitly.

Context

Prior completed repository work has already examined enterprise operating models, business-led governance, and regulatory preconditions for controlled Artificial Intelligence (AI) deployment in regulated environments. The unresolved gap is the experimentation layer: how regulated organisations create bounded spaces for probing novel ideas without collapsing auditability, safety, or compliance, and whether more structured experiment tracking and prioritisation improves learning and adoption or merely adds bureaucracy.

Approach

  1. Map the main governance and team archetypes used for exploratory work in regulated sectors: regulatory sandboxes, innovation labs, embedded-control teams, and platform-owned experimentation services.
  2. Compare the tracking and prioritisation mechanisms used to manage exploratory portfolios, including ad hoc logging, hypothesis backlogs, stage-gated reviews, and cross-functional portfolio boards.
  3. Collect empirical evidence on the relationship between mechanism structure and outcomes: experiment volume and diversity, pattern quality and actionability, compliance or risk incidents, and adoption of successful emergent practices.
  4. Examine how Cynefin-style safe-to-fail probing is adapted in regulated environments, including what guardrails, kill switches, review rights, and boundary conditions are required.
  5. Synthesise design principles, trade-offs, and likely non-linear effects between exploratory freedom and coordinating structure, including where additional structure appears to help versus hinder.

Sources

Related


Research Skill Output

(Full output from running the research skill - retained verbatim in the completed item. §§0–5 are the investigation; §6 seeds the Findings section below.)

§0 Initialise

Question: In highly regulated industries such as financial services, healthcare, and pharmaceuticals, how do organisations design governance structures, team models, and operating practices that enable safe-to-fail probing experiments in the Complex domain of the Cynefin framework, and what empirical relationship exists between the degree of structure in experiment tracking and prioritisation mechanisms and outcomes including experiment volume/diversity, pattern quality/actionability, compliance/risk incidents, and organisational adoption of emergent innovations?

The Cynefin framework is a decision-support model that classifies problems into Simple/Clear, Complicated, Complex, and Chaotic domains based on the relationship between cause and effect; it was created by David Snowden and used to match management method to problem type [fact; source: https://cynefin.io/wiki/Cynefin]. A safe-to-fail probe is a small, deliberately bounded experiment run in the Complex domain, where multiple parallel probes are launched around coherent hypotheses so that unsuccessful ideas fail in contained, tolerable ways while successful patterns are identified and amplified [fact; source: https://cynefin.io/wiki/Safe_to_fail_probes].

Scope: in scope are governance and team archetypes for exploratory work in regulated environments, tracking and prioritisation mechanisms for experiment portfolios, and empirical evidence linking mechanism structure to throughput, learning quality, compliance outcomes, and adoption. Out of scope are consumer-growth experimentation outside regulated contexts, jurisdiction-specific legal advice, purely technical model-evaluation methods, and post-standardisation scaling practice.

Constraints: prioritise public sources from 2018-2026 with preference for regulator guidance, sandbox documentation, peer-reviewed studies, and named enterprise case studies; distinguish prescribed framework guidance from observed organisational practice; treat Cynefin as a framing device rather than proof, and look for evidence on tracking-structure-to-outcome relationships even where sources do not use Cynefin language.

Prior-work check: three prior completed items were cited as scope-adjacent in the item's frontmatter before research began: Enterprise AI platform operating models examines multi-platform team structures and explore/exploit boundary-setting but does not address regulated-sector experimentation governance specifically [fact; source: https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-platform-operating-models.html]. Business-led low-code agent governance and Regulatory and standards preconditions for agentic AI address governance of deployed systems rather than pre-deployment probing [fact; source: https://davidamitchell.github.io/Research/research/2026-04-24-business-led-low-code-agent-governance.html; https://davidamitchell.github.io/Research/research/2026-04-26-agentic-ai-regulatory-preconditions-control-failure-assessment.html]. An additional search of Research/completed/ surfaced two more directly relevant items not previously cited: Conditions under which internal governance controls minimise coordination costs in regulated enterprises provides a transaction-cost lens on when governance structure reduces versus adds friction, which this item extends to the specific case of exploratory-experiment governance [fact; source: https://davidamitchell.github.io/Research/research/2026-05-23-governance-controls-effectiveness-conditions.html]. Exploit versus explore AI investment classification proposes a 70/20/10 horizon-differentiated portfolio model directly relevant to how regulated organisations might allocate resources between exploration and exploitation, though it does not address regulatory guardrails specifically [fact; source: https://davidamitchell.github.io/Research/research/2026-02-28-exploit-explore-ai-portfolio-framework.html]. Empirical evidence on rollout of organisation-wide low-code and no-code programs offers an empirical analogue on structured versus unstructured governance of a different kind of exploratory, decentralised activity [fact; source: https://davidamitchell.github.io/Research/research/2026-05-14-citizen-development-rollout-empirical-evidence.html].

§1 Question Decomposition

  1. Governance and team archetypes 1a. What named governance archetypes exist for exploratory work in regulated industries (regulatory sandboxes, innovation labs, embedded-control teams, platform-owned experimentation services)? 1b. What decision rights and escalation paths do these archetypes typically assign?
  2. Tracking and prioritisation mechanism structure 2a. What range of tracking mechanisms is documented, from ad hoc logging through hypothesis backlogs to stage-gated portfolio boards? 2b. Is there direct empirical evidence linking mechanism structure to experiment volume and diversity? 2c. Is there direct empirical evidence linking mechanism structure to pattern quality and actionability of learning? 2d. Is there direct empirical evidence linking mechanism structure to compliance or risk incidents? 2e. Is there direct empirical evidence linking mechanism structure to organisational adoption of successful innovations?
  3. Safe-to-fail adaptation under regulation 3a. What guardrails, kill switches, and review rights do regulated safe-to-fail programmes use in practice? 3b. What boundary conditions (time limits, customer caps, data-sharing limits) recur across regulator-run sandboxes?
  4. Non-linear structure-outcome relationships 4a. Is there evidence that additional tracking or governance structure helps up to a point and then hinders, or that the relationship is monotonic? 4b. What conditions determine whether more structure helps versus hinders?

§2 Investigation

1a. Governance and team archetypes. Four archetypes recur across regulated-sector sources [inference; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/]. First, the regulator-run sandbox, in which a supervisory authority admits a cohort of firms to test live products under a bespoke, time-limited waiver of specific rules while normal consumer-protection and prudential requirements continue to apply elsewhere in the firm [fact; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox]. The Financial Conduct Authority (FCA), the United Kingdom's financial services regulator, pioneered this model in 2016, and more than 50 jurisdictions have since adopted variants [fact; source: https://www.bis.org/publ/work901.htm]. Second, the health-sector regulatory sandbox, exemplified by the Medicines and Healthcare products Regulatory Agency (MHRA) AI Airlock, which ran four case studies between April 2024 and March 2025 with named participants (Philips Healthcare, AutoMedica, OncoFlow, Newton's Tree) testing distinct regulatory questions such as synthetic-data validation, large language model (LLM) hallucination mitigation, the explainability-versus-performance trade-off, and continuous post-market surveillance [inference; source: https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf]. This claim about named participants and case-study topics could not be independently corroborated from the primary Portable Document Format (PDF) report text, which the fetch tool returned only as unparsed binary; it is drawn from a secondary aggregation of the report's contents and is treated as inference rather than fact pending direct textual confirmation. [assumption; source: https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf] The follow-on London Region I sandbox, announced by MHRA, NHS England (London), and the London Health Innovation Networks, extends the same model to up to ten AI medical device manufacturers deploying in live NHS clinical settings under MHRA oversight [fact; source: https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched]. Third, the organisation-level precertification model, piloted by the United States Food and Drug Administration (FDA) as the Digital Health Software Precertification (Pre-Cert) Pilot Program between 2017 and 2022, which attempted to shift from product-by-product review to appraising a firm's quality culture and allowing it to self-certify incremental software updates [fact; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program]. Fourth, the embedded-control team model, in which risk and compliance specialists sit inside cross-functional delivery teams rather than reviewing work only at stage-gate checkpoints; this model is typically described using the Institute of Internal Auditors (IIA) Three Lines Model, which defines governing body oversight, first-line operational management, second-line risk and compliance functions, and third-line independent assurance as distinct but coordinating roles rather than sequential gates [fact; source: https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/].

1b. Decision rights and escalation. Across the regulator-run sandbox archetype, decision rights sit primarily with the regulator: the regulator selects entrants, negotiates which specific rules are relaxed for that firm, and retains ongoing supervisory oversight for the duration of the test [fact; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]. The Consultative Group to Assist the Poor (CGAP), a policy research group housed at the World Bank, documents this as a deliberate design choice: sandboxes are built around a formal legal or administrative mandate, a defined testing period, and explicit exit criteria, rather than open-ended relaxation of rules [fact; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]. In the embedded-control model, decision rights are more distributed: the IIA's own framing places accountability for risk-taking with first-line management while second-line functions provide oversight and challenge without holding a formal veto in the way a stage-gate reviewer would [inference; source: https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/]. Access note: marketing and consultancy pages on named-bank embedded-compliance case studies, surfaced by search, not independently verifiable, excluded from evidence base.

2a. Tracking and prioritisation mechanism range. The documented range runs from informal, undocumented experimentation through to fully instrumented, centrally tracked portfolios. [inference; source: https://www.stage-gate.com/blog/the-stage-gate-model-an-overview/; https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA] At the low-structure end, Coccia's taxonomy of innovation failure implicitly documents unstructured or poorly governed experimentation: his case studies from pharmaceutical, aerospace, and information and communications technology (ICT) sectors attribute a large share of project failures to errors in planning and design rather than errors of execution or marketization, consistent with weak upfront hypothesis structuring [inference; source: abstract of Coccia (2023), accessed via secondary aggregation of https://www.sciencedirect.com/science/article/pii/S0160791X23002295]. Access note: ScienceDirect article and ResearchGate mirror both returned Hypertext Transfer Protocol (HTTP) status 403; claim rests on abstract-level summary, treated as inference. At the high-structure end, Robert Cooper's Stage-Gate model for new product development (NPD) portfolios structures an innovation pipeline into defined stages separated by formal go/kill decision gates with named criteria at each gate [fact; source: https://www.stage-gate.com/blog/the-stage-gate-model-an-overview/]. A middle-structure pattern, the hypothesis-driven experiment backlog used in digital product organisations, ties each planned test to an explicit, falsifiable hypothesis statement and a decision criterion before it is run [inference; source: https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA]. None of the sources located are set specifically in a regulated-industry innovation-tracking context; the Stage-Gate and hypothesis-backlog evidence is drawn from general new-product-development and technology-sector experimentation literature and is treated as an analogy, not direct regulated-sector evidence, throughout this item. [assumption; source: https://www.stage-gate.com/blog/the-stage-gate-model-an-overview/; https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA]

2b-2e. Mechanism structure and outcomes. Direct, quantified evidence connecting tracking-mechanism structure to the four named outcome categories (volume/diversity, pattern quality, compliance incidents, adoption) is sparse and mostly indirect [inference; source: https://www.bis.org/publ/work901.htm; https://arxiv.org/abs/2406.16629]. Among the evidence gathered, the relationship between sandbox participation itself (a governance-structure choice, not a tracking-mechanism choice) and firm-level outcomes is the only one with matched-control econometric support located in this item's evidence base. [inference; source: https://www.bis.org/publ/work901.htm] Cornelli, Doerr, Gambacorta, and Merrouche's Bank for International Settlements (BIS) working paper is a matched-control econometric study of firms that entered the FCA sandbox between 2014 and 2019, finding that sandbox entry is associated with a 15% increase in the average amount of capital raised and a 50% increase in the probability of raising capital at all, alongside positive effects on firm survival and patenting. [fact; source: https://www.bis.org/publ/work901.htm] The paper attributes this to reduced information asymmetry and lower regulatory costs, evidenced by the effect being larger for smaller, younger firms and for firms attracting first-time or foreign investors [fact; source: https://www.bis.org/publ/work901.htm]. This is evidence about the effect of gaining structured regulatory access to a bounded testing space, not evidence about the internal tracking or prioritisation mechanism a firm uses once inside the sandbox; no located source measures the latter directly. [inference; source: https://www.bis.org/publ/work901.htm] The FCA's own first-cohort report states that of 146 applications, 50 firms were accepted and taken through structured testing, and that early indications show reduced time and cost to market and improved access to finance for participants [fact; source: https://www.fca.org.uk/publications/research/regulatory-sandbox-lessons-learned-report]. On pattern quality and actionability, Müller's meta-experiments paper argues that experimentation teams cannot reliably improve their own tracking and review processes using qualitative feedback or simple counts (such as experiments run) because these measures do not causally isolate what changed; instead the paper proposes applying controlled testing to the experimentation process itself [fact; source: https://arxiv.org/abs/2406.16629]. This is a general claim about measurement validity for any experimentation programme, evidenced from a technology-company setting, not a regulated-industry-specific finding, and is used here as a technical lens rather than direct regulated-sector evidence. [inference; source: https://arxiv.org/abs/2406.16629] On compliance and risk incidents, no source located provides a quantified before/after comparison of incident rates under different tracking-structure regimes; the strongest available evidence is structural rather than outcome-based, namely that regulator-run sandboxes are built around continuous supervisory monitoring precisely because that structure is presumed, not measured, to contain risk [inference; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]. On adoption of successful innovations, the MHRA sandbox model is explicitly designed so that sandbox-generated real-world evidence feeds forward into procurement and wider regulatory decisions, which is a structural adoption pathway rather than a measured adoption rate [fact; source: https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched].

3a-3b. Guardrails, kill switches, boundary conditions. Across the sandboxes reviewed, four boundary-condition types recur [inference; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched]. Time limits: the FCA sandbox and comparable national sandboxes such as the Monetary Authority of Singapore (MAS) model run cohorts over fixed testing windows rather than indefinite trials [inference; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]. Scale limits: participant numbers are capped, for example the ten-manufacturer limit on the initial phase of the MHRA London Region I sandbox [fact; source: https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched]. Legal-mandate limits: CGAP identifies a formal legal or administrative basis, transparent selection criteria, and defined exit conditions as necessary design elements, distinguishing a sandbox from informal regulatory forbearance [fact; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]. Continuous-oversight limits: the AI Airlock's four case studies each targeted a named live regulatory question (synthetic-data validation, hallucination mitigation, explainability-performance trade-off, post-market surveillance) rather than open-ended exploration, and the pilot's own framing centres on producing evidence to inform future binding regulation rather than permanently relaxing a rule [inference; source: https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf]. Hilary Allen's academic critique of the sandbox model argues that these boundary conditions function as a trade-off rather than a pure safety mechanism: sandboxes achieve safe-to-fail experimentation partly by suspending some consumer-protection and prudential requirements for the duration of the test, and Allen contends this trade-off is only justified if regulators demonstrably learn enough from the sandbox to improve regulation outside it [fact; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/]. Allen's paper further predicts that competitive pressure among jurisdictions to attract fintech business will push sandbox rules toward looser boundary conditions over time (a deregulatory race), and that this same competitive dynamic discourages cross-border information-sharing about what regulators learn, undermining one of the model's stated justifications [inference; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/].

4a-4b. Non-linear structure-outcome relationships. Of the sources reviewed for this item, the FDA's own Pre-Cert pilot is the only one that documents a governance-structure ceiling directly. [inference; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program] After five years of testing an organisation-level, lower-touch oversight model intended to speed software iteration, the FDA concluded in its final report that scaling the model would exceed its existing statutory authority, and that new legislation from the United States Congress would be required to adopt it more broadly. [fact; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program] This is direct evidence that a regulated safe-to-fail experiment can succeed at generating learning while still failing to convert that learning into a scaled operating model, because the binding constraint sits above the operating-model level, in statute rather than in the sandbox's internal governance design. [inference; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program] Coccia's taxonomy offers an inference-level, non-regulated-sector-specific account of why over-structuring can also fail: he classifies "errors in planning and/or design" as the largest source of documented innovation failure, which is consistent with excessive process rigidity crowding out the exploratory variety a safe-to-fail probe is meant to generate, though Coccia's own case studies do not isolate tracking-mechanism structure as a variable [inference; source: abstract of Coccia (2023), accessed via secondary aggregation of https://www.sciencedirect.com/science/article/pii/S0160791X23002295]. No source located directly measures an inverted-U or threshold relationship between tracking-mechanism structure and the four named outcomes inside a regulated-industry setting; the non-linearity claim in this item is therefore built from two adjacent but distinct pieces of evidence (a governance-model scaling ceiling from FDA Pre-Cert, and a general innovation-failure taxonomy from Coccia) rather than from a single study that measures the relationship directly, and is flagged as a gap in Risks/Gaps below. [assumption; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program; https://www.sciencedirect.com/science/article/pii/S0160791X23002295]

§3 Reasoning

Four governance archetypes for regulated exploratory work are documented with primary or near-primary sources: regulator-run sandboxes (financial services), health-sector regulatory sandboxes (MHRA AI Airlock, FDA Pre-Cert), and embedded-control teams (IIA Three Lines Model). [fact; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf; https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/]

Direct econometric evidence exists for one specific outcome pathway: sandbox entry increases a fintech's ability to raise capital by 15% on average and its probability of raising any capital by 50%, mediated by reduced information asymmetry. [fact; source: https://www.bis.org/publ/work901.htm] No equivalently direct, quantified evidence exists for the relationship between tracking-mechanism structure (ad hoc logs versus hypothesis backlogs versus stage-gated boards) and experiment volume, pattern quality, compliance incidents, or adoption specifically inside regulated organisations. [assumption; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://www.bis.org/publ/work901.htm; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/] The strongest available evidence for that relationship is drawn from non-regulated technology-sector experimentation literature (Kohavi et al., Müller) and general new-product-development literature (Cooper, Coccia), used here as an analogical, not direct, evidence base. [inference; source: https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA; https://arxiv.org/abs/2406.16629; https://www.stage-gate.com/blog/the-stage-gate-model-an-overview/]

A documented governance-structure ceiling exists in the FDA Pre-Cert case, where the pilot's own conclusion was that further scaling required new statutory authority rather than an internal design change. [fact; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program] A documented critique of the sandbox model's safety claim exists in Allen's analysis, which frames sandbox boundary conditions as a consumer-protection trade-off rather than a pure risk-containment mechanism, and predicts this trade-off will tend to loosen under competitive pressure between jurisdictions. [fact; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/]

§4 Consistency Check

contradiction_scan: resolved
notes: FCA lessons-learned figures (146 applications, 50 accepted) and BIS working-paper sample period (2014-2019, staggered UK sandbox entry) describe overlapping but not identical populations; both are reported as distinct, separately sourced facts rather than merged into one figure
confidence_adjustment: Coccia claim held at inference, not fact, due to abstract-only access (ScienceDirect and ResearchGate mirror both returned HTTP 403)
confidence_adjustment: MHRA AI Airlock named-participant and case-topic claim held at inference, not fact, due to unparsed binary PDF and reliance on secondary aggregation
scope_guardrail: tracking-mechanism-to-outcome relationship inside regulated organisations specifically remains an evidence gap; analogical technology-sector and general NPD evidence is labelled and kept separate from regulated-sector claims throughout
acronym_audit: LLM, FCA, BIS, CGAP, MHRA, NHS, FDA, IIA, NPD, ICT, EMDEs, PDF, HTTP each expanded at first use in §0-§2

§5 Depth and Breadth Expansion

Technical lens. The tracking-mechanism structure best evidenced in the literature (Kohavi et al.'s experimentation-platform maturity model) describes a progression from manual, ad hoc experiments through centralised tooling to self-service, low-friction platforms, with the claim that evidence quality and organisational trust in results improve as the platform matures. [inference; source: https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA] This maturity progression was developed for consumer technology companies running thousands of simultaneous online tests, a volume and risk profile that does not match regulated-industry safe-to-fail probing, where individual probes are few, high-stakes, and subject to external supervisory review rather than internal statistical power constraints. Applying the maturity-model logic directly to a regulated context would understate the role of external, non-technical constraints such as statutory authority limits. [inference; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program]

Economic lens. The BIS working paper's information-asymmetry mechanism for why sandbox entry raises capital access, that regulatory-sandbox membership functions as a costly, credible signal of firm quality to investors, is itself a market-design outcome of the sandbox's governance structure rather than of its internal experiment-tracking mechanism. [fact; source: https://www.bis.org/publ/work901.htm] This suggests that the economic value regulated organisations gain from participating in an external, regulator-run sandbox may be structurally different from the economic value of internal experiment-tracking discipline, and the two should not be conflated when designing an operating model: external sandbox participation buys a credibility signal, while internal tracking discipline buys measurement validity. [inference; source: https://www.bis.org/publ/work901.htm; https://arxiv.org/abs/2406.16629]

Regulatory lens. CGAP's design-element checklist (legal mandate, transparent selection, defined exit) and Allen's critique (boundary conditions as consumer-protection trade-off) represent two positions on the same regulatory design space: CGAP treats boundary conditions as a governance safeguard, while Allen treats the same conditions as evidence that sandboxes achieve safety partly by temporarily removing protections rather than by containing risk through structure alone. [fact; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/] Both positions are retained in this item because they address different failure modes: CGAP's concern is regulator capacity and design completeness, while Allen's concern is competitive erosion of standards over time, and the evidence base does not resolve which risk dominates in practice. [inference; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/]

Historical lens. The regulator-run sandbox model is comparatively young, dating to the FCA's 2016 launch, and the FDA's Pre-Cert pilot ran a full five-year cycle (2017-2022) before concluding that its core design exceeded existing statutory authority. [fact; source: https://www.bis.org/publ/work901.htm; https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program] The Cynefin framework's safe-to-fail probe concept predates the regulatory-sandbox wave by roughly two decades, having emerged from 1990s organisational knowledge-management practice, which suggests the regulatory-sandbox model can be read as a later, formalised, single-domain application of an older, more general Complex-domain heuristic rather than as an independently invented governance pattern. [inference; source: https://cynefin.io/wiki/Cynefin]

Behavioural lens. Allen's prediction that competitive dynamics between regulators will erode sandbox boundary conditions over time depends on an assumption about regulator behaviour under competitive pressure (a race to the bottom) that the paper argues for analytically rather than measures empirically across a large sample of jurisdictions. [inference; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/] Separately, the CGAP paper explicitly warns regulators against treating a sandbox as "a solution looking for a problem." [fact; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion] This warning implies a behavioural risk that organisations adopt the sandbox archetype for its own sake rather than because it is the best fit for their specific innovation-governance problem, an interpretive reading rather than a claim CGAP states directly. [inference; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]

§6 Synthesis

(This section seeds the Findings below.)

Executive summary:

Regulated organisations use four distinct governance archetypes for safe-to-fail probing, regulator-run sandboxes, health-sector regulatory sandboxes, organisation-level precertification, and embedded-control teams, but no located source directly measures how the structure of an internal experiment-tracking mechanism (ad hoc logs versus hypothesis backlogs versus stage-gated boards) affects experiment volume, pattern quality, compliance incidents, or adoption inside a regulated firm. [inference; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/] This item's evidence base contains no quantified relationship measuring internal tracking-mechanism structure directly, only one measuring a governance-structure choice: sandbox participation. [inference; source: https://www.bis.org/publ/work901.htm] Entry into a regulator-run sandbox, a governance-structure choice rather than an internal tracking-mechanism choice, is associated with a 15% increase in capital raised and a 50% increase in the probability of raising capital at all for participating fintech firms. [fact; source: https://www.bis.org/publ/work901.htm] Regulated safe-to-fail programmes bound their probes with legal mandates, fixed testing windows, participant caps, and continuous supervisory oversight rather than the informal, self-selected guardrails typical of Cynefin's original safe-to-fail probe design. [fact; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://cynefin.io/wiki/Safe_to_fail_probes] The relationship between governance structure and outcomes is non-linear in at least one documented direction: the FDA's own five-year Pre-Cert pilot generated useful learning but concluded that scaling its organisation-level oversight model would exceed existing statutory authority, showing that a well-run safe-to-fail experiment can still fail to convert into a scaled operating model when the binding constraint sits above the sandbox's internal design. [fact; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program] Only one academic critique of the sandbox archetype's safety claim was located in this item's evidence base. [inference; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/] Allen's academic analysis argues that sandboxes achieve safe experimentation partly by temporarily suspending consumer-protection requirements rather than purely through governance structure, and predicts this trade-off will tend to loosen as jurisdictions compete for fintech business. [fact; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/]

Key findings:

  1. Regulator-run financial sandboxes are built around a formal legal or administrative mandate, a fixed testing period, transparent selection criteria, and defined exit conditions, distinguishing them from informal regulatory forbearance. ([fact]; high confidence; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://www.fca.org.uk/firms/innovation/regulatory-sandbox)
  2. Entry into the United Kingdom's Financial Conduct Authority sandbox between 2014 and 2019 is associated with a 15% increase in the average amount of capital raised and a 50% increase in the probability of raising capital, with larger effects for smaller and younger firms. ([fact]; medium confidence; source: https://www.bis.org/publ/work901.htm)
  3. Of 146 applications received across the first two cohorts of the Financial Conduct Authority sandbox, 50 firms were accepted into structured testing, and the regulator reports early evidence of reduced time and cost to market for participants. ([fact]; medium confidence; source: https://www.fca.org.uk/publications/research/regulatory-sandbox-lessons-learned-report)
  4. Health-sector regulatory sandboxes such as the Medicines and Healthcare products Regulatory Agency's AI Airlock target named, live regulatory questions such as synthetic-data validation and post-market surveillance rather than open-ended exploration, and feed generated evidence forward into procurement and future regulatory decisions. ([inference]; medium confidence; source: https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf; https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched)
  5. The United States Food and Drug Administration's five-year Digital Health Software Precertification pilot concluded that scaling its organisation-level oversight model beyond the pilot would require new statutory authority from Congress, demonstrating a governance-structure ceiling that sits outside the sandbox's own design. ([fact]; medium confidence; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program)
  6. Academic critique of the regulatory sandbox model argues that sandboxes achieve safe experimentation partly by temporarily relaxing consumer-protection and prudential requirements rather than through governance structure alone, and predicts that competition among jurisdictions for fintech business will push sandbox rules toward looser boundary conditions over time. ([inference]; medium confidence; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/)
  7. Embedded-control team governance, as defined by the Institute of Internal Auditors' Three Lines Model, distributes risk accountability to first-line operational management while second-line risk and compliance functions provide oversight and challenge rather than holding a sequential go/kill veto typical of a stage-gated review. ([fact]; medium confidence; source: https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/)
  8. None of the regulator, standards-body, or academic sources consulted for this item, including the Financial Conduct Authority sandbox documentation, the Bank for International Settlements working paper, and the Institute of Internal Auditors Three Lines Model, measures the relationship between the structure of an internal experiment-tracking and prioritisation mechanism and experiment volume, pattern quality, compliance incidents, or adoption specifically inside a regulated organisation. ([assumption]; low confidence; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://www.bis.org/publ/work901.htm; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/)
  9. General new-product-development literature attributes a large share of documented innovation project failures to errors in upfront planning and design rather than execution or marketization errors, a pattern consistent with, though not proof of, under-structured experiment tracking contributing to failure. ([inference]; low confidence; source: abstract of Coccia (2023) accessed via secondary aggregation of https://www.sciencedirect.com/science/article/pii/S0160791X23002295)
  10. Technology-sector experimentation literature argues that qualitative feedback and simple experiment counts cannot reliably show whether a change to tracking or review process improved experimentation quality, and instead proposes applying controlled testing to the experimentation process itself; this evidence originates outside regulated industries and is used here as an analogy rather than direct regulated-sector evidence. ([inference]; medium confidence; source: https://arxiv.org/abs/2406.16629)

Evidence map:

Claim Source Confidence Notes
[fact] Regulator-run sandboxes require a formal legal mandate, fixed testing window, transparent selection, and defined exit conditions https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://www.fca.org.uk/firms/innovation/regulatory-sandbox high CGAP is a policy research body housed at the World Bank; FCA is the primary regulator
[fact] UK sandbox entry (2014-2019) raises capital by 15% on average, raises probability of raising capital by 50% https://www.bis.org/publ/work901.htm medium BIS matched-control econometric study; single-source citation, no independent replication located
[fact] FCA first two cohorts: 146 applications, 50 accepted, evidence of reduced time/cost to market https://www.fca.org.uk/publications/research/regulatory-sandbox-lessons-learned-report medium Official FCA report; some figures accessed via HTML landing-page summary rather than full PDF text
[inference] MHRA AI Airlock ran four named case studies on synthetic data, hallucination mitigation, explainability trade-off, post-market surveillance https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf medium Primary PDF unparseable in this session; claim rests on secondary aggregation, held at inference
[fact] London Region I sandbox caps initial phase at 10 AI medical device manufacturers, MHRA-overseen, live NHS deployment https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched high Official GOV.UK announcement, directly fetched
[fact] FDA Pre-Cert pilot (2017-2022) concluded scaling requires new statutory authority https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program medium Official FDA programme page; single-source citation, no independent secondary corroboration located
[fact] Allen's academic critique: sandboxes trade consumer protection for innovation; predicts race-to-the-bottom on boundary conditions https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/ medium Single academic source; directly fetched abstract; no independent corroborating critique located
[fact] IIA Three Lines Model: distributed accountability, not sequential veto, across first/second/third line https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/ medium Authoritative standards-body source; directly fetched, but page itself is a short position-paper summary
[assumption] No consulted regulator, standards-body, or academic source measures tracking-mechanism-structure-to-outcome relationship inside regulated organisations https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://www.bis.org/publ/work901.htm; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/ low Central evidence gap of the research question; see Risks/Gaps
[inference] Innovation-failure taxonomy: planning/design errors dominate documented failures across pharma, aerospace, ICT case studies abstract of Coccia (2023), https://www.sciencedirect.com/science/article/pii/S0160791X23002295 low Full text inaccessible (HTTP 403 on ScienceDirect and ResearchGate mirror); abstract-level only
[inference] Meta-experimentation: counts and qualitative feedback cannot isolate whether tracking-process changes improve experimentation quality https://arxiv.org/abs/2406.16629 medium Technology-sector source, used as analogy, not regulated-sector evidence

Assumptions:

This item assumes that findings from technology-sector experimentation platforms (Kohavi et al., Müller) transfer as directional evidence, not as regulated-sector proof, to regulated-industry safe-to-fail probing. [assumption; source: https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA] This assumption is justified because no source located studies internal experiment-tracking mechanism structure specifically inside a regulated organisation, making the technology-sector literature the closest available body of evidence on how tracking structure interacts with learning quality, even though the risk and review context differs substantially. [assumption; source: https://arxiv.org/abs/2406.16629]

This item assumes that the FCA sandbox's documented capital-raising effect and the FDA Pre-Cert pilot's documented scaling ceiling are both representative of governance-structure effects in their respective sectors rather than idiosyncratic to the specific programmes studied. [assumption; source: https://www.bis.org/publ/work901.htm] This assumption is justified because each is the most rigorously evaluated example located in its sector (an econometric matched-control study for the FCA case, a five-year government pilot with a public final conclusion for the FDA case), but neither has an independent replication in a second jurisdiction within the evidence base gathered here. [assumption; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program]

Analysis:

The evidence separates cleanly into two tiers of directness. [inference; source: https://www.bis.org/publ/work901.htm; https://arxiv.org/abs/2406.16629] The first tier, governance-structure choices such as whether to enter a regulator-run sandbox at all, has one rigorously quantified outcome relationship (the BIS capital-raising study), because that comparison has a natural control group of similar firms that did not enter the sandbox. [fact; source: https://www.bis.org/publ/work901.htm] The second tier, internal tracking-mechanism structure once inside a regulated experimentation programme, has no equivalent natural experiment in the located evidence base, because firms rarely publish comparisons of their own internal backlog or portfolio-board practices, and regulators do not require or collect that level of internal process detail. [inference; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]

A plausible rival explanation for the capital-raising effect is that firms selected into the sandbox were already stronger candidates before entry, meaning the sandbox itself adds a credibility signal rather than causing genuine quality improvement. [inference; source: https://www.bis.org/publ/work901.htm] The BIS paper addresses this by using a matched-control design and by showing the effect concentrates among firms facing the largest prior informational disadvantage (younger firms, foreign investors), which weakens the pure-selection explanation without eliminating it, since selection into any sandbox cohort is itself non-random. [inference; source: https://www.bis.org/publ/work901.htm]

The FDA Pre-Cert and MHRA AI Airlock cases represent two different resolutions of the same underlying tension between speed of learning and depth of statutory change. [inference; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program; https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf] The FDA scoped its pilot toward a permanent shift in the oversight model itself and found that shift blocked by statute, whereas the MHRA scoped its sandbox toward narrower, per-topic regulatory questions intended to inform incremental future guidance rather than a wholesale precertification regime, which may explain why the MHRA programme continued into a second, expanded phase while the FDA pilot concluded without a scaled successor. [inference; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program; https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched]

An alternative remedy to changing the governance model, adding review staff or strengthening model-quality gates instead of redesigning the operating model, is not directly evidenced as tried or rejected in either the FDA or MHRA case; both programmes moved toward operating-model change (statutory reform request, phased sandbox expansion) rather than toward scaling review headcount, but no source located explains why headcount scaling was not the chosen alternative. [assumption; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program]

Risks, gaps, uncertainties:

The central research question, the empirical relationship between tracking-and-prioritisation mechanism structure and the four named outcomes, remains substantially unanswered by direct regulated-sector evidence; every source located that discusses tracking-mechanism structure and outcomes (Kohavi et al., Müller, Cooper, Coccia) is drawn from technology-sector experimentation or general new-product-development literature rather than regulated-industry safe-to-fail programmes specifically. [assumption; source: https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA; https://arxiv.org/abs/2406.16629; https://www.stage-gate.com/blog/the-stage-gate-model-an-overview/]

The Coccia (2023) source could not be accessed beyond its abstract: both the ScienceDirect article page and a ResearchGate-hosted preprint mirror returned HTTP 403 errors when fetched directly in this session, so all claims drawn from it are held at inference confidence rather than fact. [fact; source: https://www.sciencedirect.com/science/article/pii/S0160791X23002295]

The MHRA AI Airlock pilot's named case-study participants and topics could not be verified against the primary PDF report text, which the available fetch tool returned only as unparsed binary content; the claim rests on a secondary aggregation of the report and is held at inference confidence pending direct textual confirmation. [fact; source: https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf]

The Cofie (2024) practitioner source listed in the seeded Sources is no longer accessible: the original URL now redirects to a page with no article content, and no working alternative copy was located during this session; it is retained in the Sources list as identified-but-not-consulted and contributes no claims to this item. [fact; source: https://thebftonline.com/2024/03/19/innovating-in-a-straitjacket-a-guide-to-navigating-innovation-in-highly-regulated-industries/]

No source located provides a like-for-like, cross-jurisdiction comparison of sandbox boundary-condition strictness over time, so Allen's prediction of a regulatory race-to-the-bottom in sandbox design remains a single-author analytical argument rather than an empirically confirmed trend. [assumption; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/]

The BIS working paper's capital-raising findings are specific to the FCA's sandbox and the 2014-2019 period; no equivalent econometric evaluation of the Monetary Authority of Singapore, Hong Kong Monetary Authority, or other national sandboxes was located, so the generalisability of the 15% and 50% effect sizes to other regulatory sandbox designs is unconfirmed. [fact; source: https://www.bis.org/publ/work901.htm]

Open questions:

Does the internal structure of an experiment-tracking and prioritisation mechanism (ad hoc logging versus hypothesis backlog versus stage-gated portfolio board) measurably affect experiment volume, pattern quality, compliance incidents, or adoption inside a regulated organisation, independent of whether the organisation also participates in an external regulator-run sandbox?

Has any national financial or health regulator published a like-for-like, multi-year comparison of sandbox outcomes across two or more jurisdictions that would allow Allen's race-to-the-bottom prediction to be tested empirically rather than argued analytically?

What accounts for the different scaling trajectories of the FDA's Pre-Cert pilot (concluded without a scaled successor) and the MHRA's AI Airlock (continued into an expanded second phase): is it the narrower scope of the AI Airlock's per-topic questions, a difference in statutory flexibility between the two regulators, or some other factor not captured in the sources reviewed here?

§7 Recursive Review

review_result: pass
every_section_justified: yes
threads_synthesised: yes
claims_sourced_or_labelled: yes
uncertainties_explicit: yes
acronym_audit: passed - AI, LLM, FCA, BIS, CGAP, MHRA, NHS, FDA, IIA, NPD, ICT, EMDEs, PDF, HTTP, DOI expanded at first use
domain_term_audit: passed - Cynefin framework and safe-to-fail probe defined with authoritative source at first use in §0
central_gap_acknowledged: the tracking-mechanism-structure-to-outcome relationship inside regulated organisations specifically is the item's largest evidence gap and is stated as such in Executive Summary, Key Finding 8, and Risks/Gaps rather than papered over

Findings

(Populated from §6 Synthesis above.)

Executive Summary

Regulated organisations use four distinct governance archetypes for safe-to-fail probing, regulator-run financial sandboxes, health-sector regulatory sandboxes, organisation-level precertification, and embedded-control teams, but no source located in this investigation directly measures how the structure of an internal experiment-tracking mechanism affects experiment volume, pattern quality, compliance incidents, or adoption inside a regulated firm. [inference; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/] This item's evidence base contains no quantified relationship measuring internal tracking-mechanism structure directly, only one measuring a governance-structure choice: sandbox participation. [inference; source: https://www.bis.org/publ/work901.htm] Entering a regulator-run sandbox, a governance-structure choice rather than an internal tracking-mechanism choice, is associated with a 15% increase in capital raised and a 50% increase in the probability of raising capital at all for participating fintech firms. [fact; source: https://www.bis.org/publ/work901.htm] Regulated safe-to-fail programmes bound their probes with legal mandates, fixed testing windows, participant caps, and continuous supervisory oversight rather than the informal guardrails of Cynefin's original safe-to-fail probe design. [fact; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://cynefin.io/wiki/Safe_to_fail_probes] The relationship between governance structure and outcomes is non-linear in at least one documented direction: the FDA's five-year Pre-Cert pilot generated useful learning but concluded that scaling its oversight model would exceed existing statutory authority, showing that a well-run safe-to-fail experiment can still fail to convert into a scaled operating model when the binding constraint sits above the sandbox's internal design. [fact; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program] Only one academic critique of the sandbox archetype's safety claim was located in this item's evidence base. [inference; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/] It argues that sandboxes achieve safe experimentation partly by temporarily suspending consumer-protection requirements rather than purely through governance structure, and predicts this trade-off will loosen as jurisdictions compete for fintech business. [fact; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/]

Key Findings

  1. Regulator-run financial sandboxes are built around a formal legal or administrative mandate, a fixed testing period, transparent selection criteria, and defined exit conditions, distinguishing them from informal regulatory forbearance. ([fact]; high confidence; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://www.fca.org.uk/firms/innovation/regulatory-sandbox)
  2. Entry into the United Kingdom's Financial Conduct Authority sandbox between 2014 and 2019 is associated with a 15% increase in the average amount of capital raised and a 50% increase in the probability of raising capital, with larger effects for smaller and younger firms. ([fact]; medium confidence; source: https://www.bis.org/publ/work901.htm)
  3. Of 146 applications received across the first two cohorts of the Financial Conduct Authority sandbox, 50 firms were accepted into structured testing, and the regulator reports early evidence of reduced time and cost to market for participants. ([fact]; medium confidence; source: https://www.fca.org.uk/publications/research/regulatory-sandbox-lessons-learned-report)
  4. Health-sector regulatory sandboxes such as the Medicines and Healthcare products Regulatory Agency's AI Airlock target named, live regulatory questions such as synthetic-data validation and post-market surveillance rather than open-ended exploration, and feed generated evidence forward into procurement and future regulatory decisions. ([inference]; medium confidence; source: https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf; https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched)
  5. The United States Food and Drug Administration's five-year Digital Health Software Precertification pilot concluded that scaling its organisation-level oversight model beyond the pilot would require new statutory authority from Congress, demonstrating a governance-structure ceiling that sits outside the sandbox's own design. ([fact]; medium confidence; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program)
  6. Academic critique of the regulatory sandbox model argues that sandboxes achieve safe experimentation partly by temporarily relaxing consumer-protection and prudential requirements rather than through governance structure alone, and predicts that competition among jurisdictions for fintech business will push sandbox rules toward looser boundary conditions over time. ([inference]; medium confidence; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/)
  7. Embedded-control team governance, as defined by the Institute of Internal Auditors' Three Lines Model, distributes risk accountability to first-line operational management while second-line risk and compliance functions provide oversight and challenge rather than holding a sequential go/kill veto typical of a stage-gated review. ([fact]; medium confidence; source: https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/)
  8. None of the regulator, standards-body, or academic sources consulted for this item, including the Financial Conduct Authority sandbox documentation, the Bank for International Settlements working paper, and the Institute of Internal Auditors Three Lines Model, measures the relationship between the structure of an internal experiment-tracking and prioritisation mechanism and experiment volume, pattern quality, compliance incidents, or adoption specifically inside a regulated organisation. ([assumption]; low confidence; source: https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://www.bis.org/publ/work901.htm; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/)
  9. General new-product-development literature attributes a large share of documented innovation project failures to errors in upfront planning and design rather than execution or marketization errors, a pattern consistent with, though not proof of, under-structured experiment tracking contributing to failure. ([inference]; low confidence; source: abstract of Coccia (2023) accessed via secondary aggregation of https://www.sciencedirect.com/science/article/pii/S0160791X23002295)
  10. Technology-sector experimentation literature argues that qualitative feedback and simple experiment counts cannot reliably show whether a change to tracking or review process improved experimentation quality, and instead proposes applying controlled testing to the experimentation process itself; this evidence originates outside regulated industries and is used here as an analogy rather than direct regulated-sector evidence. ([inference]; medium confidence; source: https://arxiv.org/abs/2406.16629)

Evidence Map

Claim Source Confidence Notes
[fact] Regulator-run sandboxes require a formal legal mandate, fixed testing window, transparent selection, and defined exit conditions https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion; https://www.fca.org.uk/firms/innovation/regulatory-sandbox high CGAP is a policy research body housed at the World Bank; FCA is the primary regulator
[fact] UK sandbox entry (2014-2019) raises capital by 15% on average, raises probability of raising capital by 50% https://www.bis.org/publ/work901.htm medium BIS matched-control econometric study; single-source citation, no independent replication located
[fact] FCA first two cohorts: 146 applications, 50 accepted, evidence of reduced time/cost to market https://www.fca.org.uk/publications/research/regulatory-sandbox-lessons-learned-report medium Official FCA report; some figures accessed via HTML landing-page summary rather than full PDF text
[inference] MHRA AI Airlock ran four named case studies on synthetic data, hallucination mitigation, explainability trade-off, post-market surveillance https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf medium Primary PDF unparseable in this session; claim rests on secondary aggregation, held at inference
[fact] London Region I sandbox caps initial phase at 10 AI medical device manufacturers, MHRA-overseen, live NHS deployment https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched high Official GOV.UK announcement, directly fetched
[fact] FDA Pre-Cert pilot (2017-2022) concluded scaling requires new statutory authority https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program medium Official FDA programme page; single-source citation, no independent secondary corroboration located
[fact] Allen's academic critique: sandboxes trade consumer protection for innovation; predicts race-to-the-bottom on boundary conditions https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/ medium Single academic source; directly fetched abstract; no independent corroborating critique located
[fact] IIA Three Lines Model: distributed accountability, not sequential veto, across first/second/third line https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/ medium Authoritative standards-body source; directly fetched, but page itself is a short position-paper summary
[assumption] No consulted regulator, standards-body, or academic source measures tracking-mechanism-structure-to-outcome relationship inside regulated organisations https://www.fca.org.uk/firms/innovation/regulatory-sandbox; https://www.bis.org/publ/work901.htm; https://www.theiia.org/en/content/position-papers/2020/the-iias-three-lines-model-an-update-of-the-three-lines-of-defense/ low Central evidence gap of the research question; see Risks/Gaps
[inference] Innovation-failure taxonomy: planning/design errors dominate documented failures across pharma, aerospace, ICT case studies abstract of Coccia (2023), https://www.sciencedirect.com/science/article/pii/S0160791X23002295 low Full text inaccessible (HTTP 403 on ScienceDirect and ResearchGate mirror); abstract-level only
[inference] Meta-experimentation: counts and qualitative feedback cannot isolate whether tracking-process changes improve experimentation quality https://arxiv.org/abs/2406.16629 medium Technology-sector source, used as analogy, not regulated-sector evidence

Assumptions

This item assumes that findings from technology-sector experimentation platforms transfer as directional evidence, not as regulated-sector proof, to regulated-industry safe-to-fail probing. [assumption; source: https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA] This assumption is justified because no source located studies internal experiment-tracking mechanism structure specifically inside a regulated organisation, making the technology-sector literature the closest available body of evidence on how tracking structure interacts with learning quality, even though the risk and review context differs substantially. [assumption; source: https://arxiv.org/abs/2406.16629]

This item assumes that the FCA sandbox's documented capital-raising effect and the FDA Pre-Cert pilot's documented scaling ceiling are both representative of governance-structure effects in their respective sectors rather than idiosyncratic to the specific programmes studied. [assumption; source: https://www.bis.org/publ/work901.htm] This assumption is justified because each is the most rigorously evaluated example located in its sector, an econometric matched-control study for the FCA case and a five-year government pilot with a public final conclusion for the FDA case, but neither has an independent replication in a second jurisdiction within the evidence base gathered here. [assumption; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program]

Analysis

The evidence separates cleanly into two tiers of directness. [inference; source: https://www.bis.org/publ/work901.htm; https://arxiv.org/abs/2406.16629] The first tier, governance-structure choices such as whether to enter a regulator-run sandbox at all, has one rigorously quantified outcome relationship, because that comparison has a natural control group of similar firms that did not enter the sandbox. [fact; source: https://www.bis.org/publ/work901.htm] The second tier, internal tracking-mechanism structure once inside a regulated experimentation programme, has no equivalent natural experiment in the located evidence base, because firms rarely publish comparisons of their own internal backlog or portfolio-board practices, and regulators do not require or collect that level of internal process detail. [inference; source: https://www.cgap.org/research/publication/regulatory-sandboxes-and-financial-inclusion]

A plausible rival explanation for the capital-raising effect is that firms selected into the sandbox were already stronger candidates before entry, meaning the sandbox itself adds a credibility signal rather than causing genuine quality improvement. [inference; source: https://www.bis.org/publ/work901.htm] The BIS paper addresses this by using a matched-control design and by showing the effect concentrates among firms facing the largest prior informational disadvantage, which weakens the pure-selection explanation without eliminating it, since selection into any sandbox cohort is itself non-random. [inference; source: https://www.bis.org/publ/work901.htm]

The FDA Pre-Cert and MHRA AI Airlock cases represent two different resolutions of the same underlying tension between speed of learning and depth of statutory change. [inference; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program; https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf] The FDA scoped its pilot toward a permanent shift in the oversight model itself and found that shift blocked by statute, whereas the MHRA scoped its sandbox toward narrower, per-topic regulatory questions intended to inform incremental future guidance rather than a wholesale precertification regime, which may explain why the MHRA programme continued into a second, expanded phase while the FDA pilot concluded without a scaled successor. [inference; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program; https://www.gov.uk/government/news/pioneering-ai-health-innovations-regulatory-sandbox-launched]

An alternative remedy to changing the governance model, adding review staff or strengthening model-quality gates instead of redesigning the operating model, is not directly evidenced as tried or rejected in either the FDA or MHRA case; both programmes moved toward operating-model change rather than toward scaling review headcount, but no source located explains why headcount scaling was not the chosen alternative. [assumption; source: https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program]

Risks, Gaps, and Uncertainties

The central research question, the empirical relationship between tracking-and-prioritisation mechanism structure and the four named outcomes, remains substantially unanswered by direct regulated-sector evidence; every source located that discusses tracking-mechanism structure and outcomes is drawn from technology-sector experimentation or general new-product-development literature rather than regulated-industry safe-to-fail programmes specifically. [assumption; source: https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/experimentation-platform-and-culture/A83F07714647BA8454B40C256473BCEA; https://arxiv.org/abs/2406.16629; https://www.stage-gate.com/blog/the-stage-gate-model-an-overview/]

The Coccia (2023) source could not be accessed beyond its abstract: both the ScienceDirect article page and a ResearchGate-hosted preprint mirror returned HTTP 403 errors when fetched directly in this session, so all claims drawn from it are held at inference confidence rather than fact. [fact; source: https://www.sciencedirect.com/science/article/pii/S0160791X23002295]

The MHRA AI Airlock pilot's named case-study participants and topics could not be verified against the primary PDF report text, which the available fetch tool returned only as unparsed binary content; the claim rests on a secondary aggregation of the report and is held at inference confidence pending direct textual confirmation. [fact; source: https://assets.publishing.service.gov.uk/media/68ee1fb88427701993d5e02c/AI_Airlock_Sandbox_Programme_Report_Final.pdf]

The Cofie (2024) practitioner source listed in the seeded Sources is no longer accessible: the original URL now redirects to a page with no article content, and no working alternative copy was located during this session; it is retained in the Sources list as identified-but-not-consulted and contributes no claims to this item. [fact; source: https://thebftonline.com/2024/03/19/innovating-in-a-straitjacket-a-guide-to-navigating-innovation-in-highly-regulated-industries/]

No source located provides a like-for-like, cross-jurisdiction comparison of sandbox boundary-condition strictness over time, so Allen's prediction of a regulatory race-to-the-bottom in sandbox design remains a single-author analytical argument rather than an empirically confirmed trend. [assumption; source: https://scholarship.law.vanderbilt.edu/jetlaw/vol22/iss2/3/]

The BIS working paper's capital-raising findings are specific to the FCA's sandbox and the 2014-2019 period; no equivalent econometric evaluation of the Monetary Authority of Singapore, Hong Kong Monetary Authority, or other national sandboxes was located, so the generalisability of the 15% and 50% effect sizes to other regulatory sandbox designs is unconfirmed. [fact; source: https://www.bis.org/publ/work901.htm]

Open Questions

Does the internal structure of an experiment-tracking and prioritisation mechanism measurably affect experiment volume, pattern quality, compliance incidents, or adoption inside a regulated organisation, independent of whether the organisation also participates in an external regulator-run sandbox?

Has any national financial or health regulator published a like-for-like, multi-year comparison of sandbox outcomes across two or more jurisdictions that would allow Allen's race-to-the-bottom prediction to be tested empirically rather than argued analytically?

What accounts for the different scaling trajectories of the FDA's Pre-Cert pilot, which concluded without a scaled successor, and the MHRA's AI Airlock, which continued into an expanded second phase: is it the narrower scope of the AI Airlock's per-topic questions, a difference in statutory flexibility between the two regulators, or some other factor not captured in the sources reviewed here?


Output

  • Type: knowledge
  • Description: A structured evidence map of governance archetypes (regulator-run sandboxes, health-sector sandboxes, organisation-level precertification, embedded-control teams) used for safe-to-fail probing in regulated industries, with the finding that direct empirical evidence links external sandbox participation to firm-level outcomes but no located source measures the effect of internal experiment-tracking mechanism structure on outcomes inside a regulated organisation specifically. [inference; source: https://www.bis.org/publ/work901.htm; https://www.fda.gov/medical-devices/digital-health-center-excellence/digital-health-software-precertification-pre-cert-pilot-program]
  • Links: BIS Working Paper 901, FDA Digital Health Software Precertification Pilot Program, Allen (2019) Sandbox Boundaries

Navigation

Home

By Tag

bureaucracy

change-management

coase

constraint-analysis

control-model

decision-rights

delegation

delivery-risk

demand-segmentation

enterprise

exception-handling

execution

flow

flow-design

flow-metrics

governance

governance-patterns

incentives

instability

institutional-economics

leading-indicators

operating-model

organisation

organisational-design

queue-design

queueing

regulated-enterprise

routing

throughput

throughput-risk

transaction-costs

triage

williamson

Clone this wiki locally