-
Notifications
You must be signed in to change notification settings - Fork 0
2026 05 08 ai capability reference architecture second cycle update
Updating the enterprise Artificial Intelligence ecosystem capability reference architecture using second-cycle 2026-05 completed items
How should the enterprise Artificial Intelligence (AI) ecosystem capability reference architecture (as expressed in 2026-04-22-enterprise-ai-capability-model, 2026-05-05-enterprise-ai-capability-stack, and 2026-05-06-ai-capability-reference-architecture-security-supply-chain-update) be revised and extended to incorporate findings from the second-cycle 2026-05 completed items, covering AI Bill of Materials (AIBOM) declared construction practices, effectiveness and risk-mitigation limits, multi-agent identity attribution, platform observability controls, European Union (EU) AI Act regulatory intersection, and OpenTelemetry-based runtime capture; AI production incidents; regulatory guidance updates; Five Eyes AI risk posture; open-weight model safeguard policies; and Information Technology (IT) legibility and measurement frameworks, that were not incorporated in the first-cycle update?
In scope:
- Gap analysis: which capability domains in the first-cycle updated architecture are absent or underspecified relative to the second-cycle completed items listed in
cites - AIBOM operational practice integration: how AIBOM declared construction practices, effectiveness and risk limits, multi-agent identity attribution, platform observability comparison, European Union (EU) AI Act intersection, and OpenTelemetry (OTel) runtime capture translate to refined or new capability components in the architecture
- Production incident mapping: how failure modes and operational controls documented in
2026-05-07-ai-production-incidents-deep-divemap to existing or missing capability layers; whether the architecture adequately anticipates the failure classes observed - Regulatory guidance currency: what capability requirements arise from the regulatory guidance delta (
2026-05-07-ai-regulatory-guidance-update-gap-check) that are absent from the current architecture, particularly for financial-services regulators (Australian Prudential Regulation Authority (APRA), Reserve Bank of New Zealand (RBNZ), Financial Markets Authority (FMA), Office of the Superintendent of Financial Institutions (OSFI)) - Five Eyes security integration: how the Five Eyes AI risk posture (
2026-05-07-five-eyes-ai-risks-and-advice) translates to architectural control requirements not covered by the first-cycle security threat model - Open-weight model safeguard capabilities: where open-weight model safeguard and policy-enforcement approaches (
2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight) belong in the architecture - IT legibility and measurement: how IT legibility and measurement frameworks (
2026-05-06-it-system-legibility-measurement-frameworks) inform the observability and evidence-collection layers - Reconciliation of overlaps and conflicts across the second-cycle items before finalising architectural recommendations
- Updated capability ownership recommendations and a revised component list (additions, amendments, and explicit supersessions of first-cycle components)
Out of scope:
- Items already incorporated in the first-cycle update (
2026-05-06-ai-capability-reference-architecture-security-supply-chain-update) - Implementation guidance for specific vendor products
- Original primary research into any of the cited source domains (rely on already-completed items)
- Revision of the underlying five-layer baseline model structure (unless the second-cycle items give specific grounds to do so)
Constraints:
- Must cite only completed research items; no speculative claims about items still in backlog or in-progress
- Must produce a revised architecture component list that is directly traceable to evidence in the cited items
- Must use the same structural conventions as
2026-05-06-ai-capability-reference-architecture-security-supply-chain-updateto enable comparison
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html; https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md] The first-cycle architecture update retained the five-layer enterprise AI backbone and added two cross-cutting planes, one for supply-chain provenance and one for policy, evaluation, and evidence.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-declared-construction-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-effectiveness-risk-mitigation-limits.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html; https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] The second-cycle completed items add operational detail on declared construction, runtime capture, identity attribution, observability, regulation, incidents, open-weight safeguards, and estate legibility that was not folded into the first-cycle architecture.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] Without a second-cycle revision, the architecture would understate the operating-model, runtime-evidence, and regulatory capabilities needed to keep enterprise AI systems reviewable after deployment rather than only before release.
- AIBOM operational gap analysis: For each of the six second-cycle AIBOM items, declared construction, effectiveness limits, multi-agent attribution, platform observability, European Union (EU) AI Act intersection, and OpenTelemetry runtime capture, identify what capability is implied that is absent from or only partially present in the first-cycle architecture. Produce a gap table.
-
Production incident layer review: Map the failure modes and control categories from
2026-05-07-ai-production-incidents-deep-diveonto the current architecture layers. Determine whether existing capability domains cover the required controls or whether new domains are needed. -
Regulatory capability requirements: Extract capability-level requirements from
2026-05-07-ai-regulatory-guidance-update-gap-checkand evaluate coverage in the current architecture. Focus on financial-services-specific obligations not addressed in the first cycle. -
Five Eyes security posture mapping: Translate the AI risk and advisory posture from
2026-05-07-five-eyes-ai-risks-and-adviceinto architectural control surfaces. Identify gaps relative to the existing security capability domains. -
Open-weight model safeguard placement: Determine the correct architectural home for open-weight model safeguard policies and policy-enforcement gates (
2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight). -
Legibility and measurement layer: Evaluate whether the observability and evidence-collection layers sufficiently reflect IT legibility frameworks (
2026-05-06-it-system-legibility-measurement-frameworks). Propose additions where coverage is lacking. - Cross-item reconciliation: Identify overlaps or conflicts among the second-cycle items, for example AIBOM runtime capture versus production-incident telemetry requirements, and state explicitly how they are resolved in the revised architecture.
- Revised architecture output: Produce an updated capability component list and ownership table that incorporates findings from steps 1 through 7, explicitly noting which first-cycle components are amended and which new components are added.
- Mitchell (2026) Enterprise AI capability model for use-case maturity decisions: baseline five-layer reference architecture
- Mitchell (2026) Enterprise capability stack for sustainable multi-provider Artificial Intelligence: shared-control-core synthesis that the architecture extends
- Mitchell (2026) Integrating 2026-05 security and supply chain findings into the enterprise AI capability reference architecture: first-cycle update; establishes the immediate baseline for this second-cycle revision
- Mitchell (2026) AIBOM declared construction practice: operational practice for constructing declared AIBOM artifacts
- Mitchell (2026) AIBOM effectiveness and risk-mitigation limits: where AIBOM controls are and are not effective
- Mitchell (2026) AIBOM identity attribution in multi-agent systems: practice: multi-agent identity attribution practices and architectural implications
- Mitchell (2026) AIBOM platform observability and control comparison: comparative view of platform observability controls for AIBOM
- Mitchell (2026) AIBOM and EU AI Act regulatory intersection: compliance obligations and intersection between AIBOM and the European Union (EU) AI Act
- Mitchell (2026) AIBOM runtime capture via OpenTelemetry: practice: runtime AIBOM generation using OpenTelemetry instrumentation
- Mitchell (2026) Production incidents linked to AI systems: documented AI production incident failure modes and mitigations
- Mitchell (2026) AI regulatory guidance delta check: new regulatory guidance and coverage gaps since prior review
- Mitchell (2026) Five Eyes stance on AI risk and policy advice: Five Eyes AI risk posture and architectural security implications
- Mitchell (2026) Open-weight model safeguard policy enforcement: safeguard and policy enforcement for open-weight models
- Mitchell (2026) IT system legibility and measurement frameworks: legibility and measurement frameworks for IT systems; informs observability layer
- Mitchell (2026) AI agent control-plane architecture in the enterprise: prior control-surface architecture used to qualify placement decisions
- Mitchell (2026) AI agent identity and access management in the enterprise: prior identity and delegation architecture relevant to second-cycle identity findings
- Mitchell (2026) Permission-safe Retrieval-Augmented Generation enterprise information architecture: prior retrieval and authorization architecture relevant to incident and observability conclusions
- Mitchell (2026) Knowledge curation governance for regulated Artificial Intelligence: prior authoritative-source governance item that qualifies incident and evidence recommendations
- Open Worldwide Application Security Project AIBOM: authoritative public definition surface for the AI bill-of-materials concept
- CycloneDX AI/ML-BOM: standards-aligned definition and scope for AI and machine-learning bill-of-materials work
- Backstage Software Catalog: authoritative catalog-coverage and ownership surface used in the legibility definition
- ServiceNow CMDB Health: authoritative configuration-data quality surface used in the legibility definition
- Dynatrace Smartscape: authoritative runtime-topology surface used in the legibility definition
(Full output from running the research skill, retained verbatim in the completed item. §§0–5 are the investigation; §6 seeds the Findings section below.)
- Question: how should the existing enterprise AI capability reference architecture be revised so that second-cycle findings become explicit architectural services, control surfaces, and ownership responsibilities rather than remaining implicit extensions of the first-cycle update?
- Scope: extend the five-layer architecture and first-cycle cross-cutting planes; identify missing or weakly specified capability components; produce a revised component list and ownership recommendations; avoid product-specific implementation playbooks.
- Constraints: use completed repository items and adjacent completed prior work only; keep claims traceable to URL-backed completed items; preserve the baseline architecture unless the second-cycle evidence gives strong reasons to restructure it.
- Output: knowledge, specifically a second-cycle architecture update that adds capability deltas, ownership recommendations, and explicit control-surface revisions.
- Prior completed items reviewed before investigation: https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html ; https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md ; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html ; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html ; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-identity-access-management-enterprise.html ; https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html ; https://davidamitchell.github.io/Research/research/2026-04-22-knowledge-curation-governance-for-regulated-ai.html
- Root question: what concrete architecture changes are now required to absorb the second-cycle evidence without losing the baseline five-layer shape?
-
A. Baseline and gap test
- A1. Which first-cycle capabilities already exist and remain valid?
- A2. Which second-cycle findings introduce new control surfaces rather than merely richer examples of existing ones?
-
B. Design-time supply-chain capability
- B1. Which declared-construction capabilities must exist before promotion?
- B2. Which release-time records must connect approved design to later runtime evidence?
-
C. Runtime evidence and observability
- C1. Which runtime-observed capabilities are required to compare declared design with observed execution?
- C2. Which observability gaps are themselves governance failures?
-
D. Identity, authority, and action control
- D1. Which identity and delegation controls belong in orchestration rather than in static identity registries alone?
- D2. Which actions need checkpoint, stop-authority, or exception pathways?
-
E. Incident and regulatory control surfaces
- E1. Which production-incident mitigations belong in the architecture as first-class controls?
- E2. Which regulatory and Five Eyes requirements now require named capabilities or named owners?
-
F. Legibility and evidence
- F1. How should estate legibility be represented, as one score or as a composite evidence surface?
- F2. Where should architecture blueprints, catalogs, runtime topology, and structural-drift evidence live?
-
G. Policy reasoning and ownership
- G1. Where do open-weight safeguard models fit, if at all, inside the architecture?
- G2. Which capabilities should remain centrally owned and which should stay domain-owned?
- [fact; source: https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html; https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md] The baseline architecture remains a five-layer model with a shared control core around policy, identity, observability, evaluation, and vendor-governance functions.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html] The first-cycle update already promoted provenance, delegated identity, runtime evidence, and evaluation into explicit cross-cutting enterprise services rather than leaving them implicit.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-identity-access-management-enterprise.html; https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html] Adjacent prior work already established that policy translation, multi-hop identity, and permission-safe retrieval are layer-specific trust decisions rather than generic gateway functions.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] The second-cycle task is therefore not to replace the architecture, but to widen the first-cycle planes so they can absorb runtime legibility, operational assurance, and regulator-facing evidence obligations.
- [fact; source: https://owaspaibom.org/; https://cyclonedx.org/capabilities/mlbom/; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-declared-construction-practice.html] AIBOM refers here to an AI bill-of-materials artifact that extends bill-of-materials style inventory to models, data, configuration, prompts, tools, and related execution context, and declared construction is operationally feasible from both managed control-plane resources and source-controlled orchestration artifacts.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html] Runtime-observed AIBOM capture is feasible only when telemetry emission, collector design, retention, and query backends are treated as part of the control surface rather than as optional debugging tooling.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html] Vendor observability surfaces differ materially across configuration export, runtime trace depth, registry visibility, and policy hooks, so portable assurance still requires an enterprise-side normalization layer.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html] Multi-agent identity attribution breaks when systems switch credential type, cross trust boundaries, or fall back to shared runtime credentials, which means delegation-chain evidence must survive beyond the presenting workload identity.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-effectiveness-risk-mitigation-limits.html] AIBOM is structurally useful for versions, dependencies, guardrails, retrieval stores, delegation edges, and divergence checks, but it does not by itself prove semantic safety against prompt injection, poisoned retrieval content, memory poisoning, or harmful use of declared tools.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html] The European Union AI Act and adjacent governance frameworks make AIBOM most useful as a traceability and dependency artifact that links upward to provider documentation and outward to runtime evidence, not as a complete compliance dossier on its own.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-declared-construction-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html] The architecture therefore needs three named AIBOM services that the first-cycle version only implied: declared-construction generation, runtime-evidence normalization, and divergence classification between approved design and observed execution.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-identity-access-management-enterprise.html] The first-cycle identity plane is also too weakly specified for second-cycle evidence, because it needs an explicit delegation-receipt service rather than generic identity and access management alone.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html] The validated production-incident set clusters around four failure classes: authoritative but wrong generated guidance, discriminatory automated decision logic, infrastructure and privacy defects around AI services, and accumulated control gaps created when deployment outruns validation and control design.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-04-22-knowledge-curation-governance-for-regulated-ai.html; https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html] Prior repository work already showed that authoritative-source governance and permission-safe retrieval are decisive deployment controls when systems produce consequential answers from enterprise or public knowledge.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-04-22-knowledge-curation-governance-for-regulated-ai.html; https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html] The architecture needs a stronger operating-assurance layer for authoritative-source binding, fairness validation, rollback authority, and external challenge pathways, because disclaimers and pilot labels did not prevent harm once systems looked authoritative to users.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html] Production evidence also widens the runtime-monitoring requirement from generic observability to control-oriented observability, because the decisive signals are trigger conditions, approved source sets, fairness thresholds, escalation paths, and rollback readiness rather than only latency or availability.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html] The main post-baseline supervisory change is APRA's April 2026 letter, which adds concrete expectations on board literacy, supplier transparency, concentration risk, continuous validation, and control over AI-enabled workflows that can take multi-step actions with limited human intervention.
- [fact; source: https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html] Five Eyes guidance converges on secure-by-design development, least privilege, threat modelling, supply-chain scrutiny, strong logging and monitoring, rollback readiness, and clear human accountability.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html] The architecture now needs named operating-model capabilities for board and executive AI literacy, supplier and concentration-risk management, regulator-facing evidence assembly, and action-governance over systems that can take multi-step actions with limited human intervention.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html] Five Eyes advice and incident evidence reinforce each other rather than competing: secure-by-design control surfaces matter because public failures repeatedly showed that user harm surfaced after deployment where runtime logging, input control, source governance, and rollback should have been stronger.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://docs.github.com/en/actions/reference/runners/github-hosted-runners] Open-weight safeguard models are best treated as policy-conditioned safety classifiers that are useful for context-sensitive review, but are not practical inline dependencies for standard GitHub-hosted runners because the documented model footprints exceed default runner hardware.
- [fact; source: https://backstage.io/docs/features/software-catalog/; https://www.servicenow.com/community/cmdb-forum/cmdb-health-dashboard-completeness-compliance-amp-correctness/m-p/3488492; https://docs.dynatrace.com/docs/analyze-explore-automate/smartscape; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] Estate legibility, meaning architecture-model completeness, catalog discoverability, configuration-data quality, runtime dependency visibility, drift detection, and shared team understanding, has no single universal benchmark.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html] Open-weight safeguard capability therefore belongs as an optional second-stage policy-reasoning service inside the governance and evaluation plane rather than as a default platform primitive in every runtime path.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html] The first-cycle architecture also understated legibility as a capability family, because it now needs an explicit composite made of architecture blueprints, ownership-bearing catalogs, runtime topology, and structural-drift evidence rather than generic observability alone.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] The least disruptive revision is to keep the five vertical layers and the two cross-cutting planes, but widen the provenance plane into a provenance-and-legibility plane and widen the governance plane into a governance, evaluation, and operational-assurance plane.
- [inference; source: https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md; https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html] Central ownership remains strongest for shared semantics, evidence retention, regulatory reporting, and policy baselines, while domain ownership remains strongest for authoritative content, workflow composition, and threshold tuning tied to local risk and operating context.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] The second-cycle evidence supports architectural widening rather than architectural replacement, because the new material mostly sharpens runtime-evidence, operational-assurance, and legibility requirements inside the existing five-layer frame.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-effectiveness-risk-mitigation-limits.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html] A complete-looking inventory is not a complete control system, so the architecture must keep structural traceability and semantic or behavioral assurance as separate but linked capability families.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html] The incident and Five Eyes evidence are more decision-useful for architecture than for model ranking, because they point to where control surfaces failed in deployment rather than to which model family was intrinsically strongest.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html] Open-weight safeguard evidence changes placement, not platform shape, because it suggests an optional policy-reasoning stage rather than a reason to collapse governance into one heavyweight moderation service.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-effectiveness-risk-mitigation-limits.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html] Apparent contradiction resolved: AIBOM is structurally valuable and still insufficient for semantic safety, so the architecture must combine declared inventory, runtime evidence, and semantic-control capabilities rather than choosing one.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html] Apparent contradiction resolved: a shared enterprise control layer is still required, but trust decisions stay distributed across data, model, orchestration, and delivery surfaces rather than collapsing into one gateway.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://docs.github.com/en/actions/reference/runners/github-hosted-runners] Apparent contradiction resolved: open-weight safeguard models are architecturally relevant even though they are not practical on default hosted runners, because the architecture can represent them as optional external services rather than default inline dependencies.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] Apparent contradiction resolved: legibility does not need one universal score to become a capability family, because the best-supported design is a composite of complementary evidence surfaces.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] Technical lens: the decisive new technical capability is not another model primitive, but a correlated evidence loop that can join declared design, runtime traces, catalogs, topology maps, and divergence results.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html] Regulatory lens: the architecture now needs stronger document-assembly and reporting capabilities because regulators are asking for governance proof, supplier visibility, validation evidence, and implementation timing discipline, not only high-level principles.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html] Economic lens: the highest-leverage investments remain platform and evidence controls, because incidents and safeguard constraints both show that weak review capacity and weak runtime evidence destroy more value than slower generation speed.
- [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html] Behavioral lens: users and operators trust systems that look authoritative, so the architecture must assume that if a system can act or answer confidently, then legibility, escalation, and rollback have to exist before harm rather than after complaint.
The enterprise AI capability reference architecture should keep the five-layer backbone and the two first-cycle cross-cutting planes, but both planes now need explicit second-cycle services for legibility, runtime evidence, operational assurance, and regulator-facing governance. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html]
The first-cycle update correctly introduced provenance and evidence as shared concerns, but the second-cycle items show that design-time inventory is not enough unless the architecture also captures runtime divergence, delegation receipts, incident-triggering conditions, and estate legibility across catalogs, topology, and drift evidence. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html]
Regulatory and Five Eyes guidance also turn operating-model capabilities into architecture components, because board literacy, supplier-risk management, secure-by-design logging, rollback authority, and evidence assembly now function as named control surfaces rather than background process assumptions. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
This revision widens the provenance plane into a provenance-and-legibility plane and widens the governance plane into a governance, evaluation, and operational-assurance plane, while assigning explicit central ownership for the shared semantics and retained evidence that domain teams cannot safely recreate on their own. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md]
- The five-layer enterprise AI architecture still holds, but its first-cycle provenance plane must expand into a provenance-and-legibility plane that explicitly covers declared construction, runtime normalization, divergence classification, catalog coverage, runtime topology, and structural-drift evidence. ([inference]; high confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html)
- Declared AIBOM generation must become a named delivery and platform-engineering capability, because second-cycle evidence shows that managed platforms and code-native orchestration frameworks both need deliberate extraction paths before approved design can be compared with later runtime behavior. ([inference]; high confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-declared-construction-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html)
- Runtime-observed evidence must become a first-class architecture service rather than a debugging by-product, because the second-cycle items show that prompts, retrieved context, tool order, authority handoffs, and missing observability are all governance-relevant facts that only appear after execution begins. ([inference]; high confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html)
- Delegation-aware identity and tool-action control now need explicit orchestration-layer components, because second-cycle identity findings show that portable attribution breaks when systems change credential type, cross trust boundaries, or execute under shared runtime identities without a surviving delegation receipt. ([inference]; high confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-identity-access-management-enterprise.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html)
- Production-incident evidence requires the architecture to add operating-assurance capabilities for authoritative-source binding, fairness validation, rollback authority, and external challenge escalation, because public harm repeatedly surfaced where systems looked trustworthy but source governance, deployment constraints, or oversight loops were too weak. ([inference]; high confidence; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-04-22-knowledge-curation-governance-for-regulated-ai.html; https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html)
- Regulatory and Five Eyes updates mean that board literacy, supplier and concentration-risk governance, secure-by-design logging, input control, and regulator-facing evidence assembly must now be named operating-model capabilities rather than assumed management practices outside the architecture. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html)
- Open-weight safeguard models fit best as shared governance and evaluation services that delivery gates or higher-risk runtime checkpoints can invoke, because they offer organization-specific policy judgment while remaining too infrastructure-heavy for standard hosted-runner paths and too narrow to replace deterministic or human review layers. ([inference]; low confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://docs.github.com/en/actions/reference/runners/github-hosted-runners; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html)
- The architecture’s ownership model should stay centralized for provenance schema, telemetry normalization, retained evidence, policy semantics, regulatory reporting, and evaluation standards, while domain teams keep responsibility for authoritative content, workflow composition, and risk-tuned thresholds that depend on local business context. ([inference]; high confidence; source: https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md; https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html)
- Assumption: Enterprises can emit enough runtime telemetry to reconcile approved design and observed execution at least for high-risk workflows. Justification: the runtime-capture and platform-observability items treat telemetry enablement as difficult but operationally feasible. [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html]
- Assumption: Financial-services and security-sensitive enterprises will continue to prefer centralized evidence retention and reporting even when execution remains federated. Justification: the regulatory-delta and Five Eyes items both point toward stronger centralized accountability rather than looser local proof models. [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
- Assumption: Open-weight safeguard services will stay optional because their infrastructure cost and policy-design overhead will not be justified for every workflow. Justification: the safeguard item supports selective, second-stage use rather than universal inline placement. [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html]
The second-cycle evidence makes the architecture more operational, not more abstract, because the missing pieces are evidence-handling and control-surface capabilities that only appear when systems are released, observed, challenged, and regulated in practice. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html]
The main trade-off is between central coherence and layer-local enforcement. [inference; source: https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html] Centralizing everything would hide where trust decisions really occur, while leaving every team to invent its own evidence model would destroy comparability, auditability, and incident reconstruction. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html]
The best-supported resolution is to keep control semantics and retained evidence centralized while leaving enforcement near the layer that owns the relevant trust decision, retrieval permissions in data and knowledge, semantic safeguards near inference, tool and delegation controls in orchestration, and release-time signing, extraction, and evaluation in delivery and platform engineering. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html]
Plausible alternative placements exist for open-weight safeguard capability, especially model-layer controls and pre-deployment review gates, but the evidence supports locating the canonical policy service in the governance and evaluation plane so that delivery gates and runtime checkpoints can invoke one shared policy logic rather than duplicate policy semantics in each layer. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html]
Revised component list:
- Data and knowledge layer: authoritative-source registry, permission-safe retrieval, retrieval-snapshot metadata, and upstream provider-disclosure links. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html]
- Model and inference layer: approved model and provider registry, semantic safeguards, policy-conditioned classification hooks, and model-facing runtime signal capture. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
- Orchestration and execution layer: tool allowlists, delegation-receipt capture, action checkpoints, recursion and stop authority, and per-run authority context. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html]
- Delivery and platform-engineering layer: declared AIBOM builder, signed artifact lineage, promotion-time evaluation gates, telemetry collector pipeline, and approved-versus-observed divergence classifier. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-declared-construction-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html]
- Operating-model layer: board and executive literacy, supplier and concentration-risk governance, fairness review, incident response, rollback authority, and regulator-facing evidence assembly. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html]
- Provenance-and-legibility plane: declared and observed supply-chain records, ownership-bearing catalogs, architecture blueprints, runtime topology, and structural-drift evidence. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html]
- Governance, evaluation, and operational-assurance plane: policy semantics, evaluation standards, exceptions, explainability artifacts, incident records, and optional second-stage safeguard review. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html]
Ownership recommendations:
- Central platform, security, and risk functions should own provenance schema, telemetry normalization, retained evidence, evaluation standards, and regulatory reporting because those assets lose value when fragmented by domain. [inference; source: https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html]
- Domain teams should own authoritative content, workflow composition, business-specific thresholds, and local exception context because those choices depend on business meaning and risk appetite that the central platform cannot infer. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html; https://davidamitchell.github.io/Research/research/2026-04-22-knowledge-curation-governance-for-regulated-ai.html]
- Shared services should expose signed interfaces and evidence contracts rather than monolithic approval queues, because second-cycle evidence favors strong shared semantics with layer-local enforcement and action-specific stop authority. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
- Runtime-evidence quality remains sensitive to platform configuration and adapter quality, so an architecture can still overstate observability if collectors, spans, or export paths are only partially enabled. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html]
- The architecture can reduce structural and operational blind spots without eliminating semantic failure modes, because AIBOM and telemetry remain weaker than adversarial content and authority misuse at proving behavior is safe. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-effectiveness-risk-mitigation-limits.html]
- The strongest financial-services regulatory evidence in this cycle comes from APRA and cross-jurisdiction synthesis rather than from a uniform new standard across all reviewed regulators, so some operating-model recommendations remain medium-confidence generalizations. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html]
- Legibility metrics are still composite rather than standardized, so estates may need local scorecards before cross-domain comparisons become reliably meaningful. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html]
- What minimum delegation-receipt schema would let orchestration runtimes preserve subject, actor, scope, target, and approval context across cross-platform tool calls?
- Which minimal composite legibility dashboard can compare catalog coverage, runtime topology coverage, divergence rates, and structural-drift findings without becoming another stale governance artifact?
- When does a policy-conditioned safeguard service justify its infrastructure cost compared with deterministic rules plus sampled human review?
- Which evidence thresholds should trigger automatic rollback versus manual escalation when approved design and observed runtime diverge in high-risk workflows?
- Completeness check: §§0-6 populated.
- Label and source audit: completed.
- Prior-work cross-reference sweep: repeated before Findings.
- Confidence setting: medium.
(Populated from §6 Synthesis above.)
The enterprise AI reference architecture should remain a five-layer model, but the shared provenance and governance planes now need explicit second-cycle services for legibility, meaning catalog, configuration, topology, and drift visibility, runtime evidence, operational assurance, and regulator-facing governance. [inference; source: https://backstage.io/docs/features/software-catalog/; https://www.servicenow.com/community/cmdb-forum/cmdb-health-dashboard-completeness-compliance-amp-correctness/m-p/3488492; https://docs.dynatrace.com/docs/analyze-explore-automate/smartscape; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html]
Second-cycle evidence shows that design-time inventory alone is insufficient, because the architecture also has to preserve runtime divergence, delegation receipts, incident-triggering conditions, and estate legibility across catalogs, topology, and structural-drift evidence. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html]
Regulatory and Five Eyes updates also move operating-model work into the architecture itself, because board literacy, supplier-risk management, secure logging, rollback authority, and evidence assembly are now observable control surfaces instead of background management assumptions. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
The practical revision is to widen the provenance plane into a provenance-and-legibility plane, widen the governance plane into a governance, evaluation, and operational-assurance plane, and keep shared semantics and retained evidence under explicit central ownership. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md]
- The existing five-layer architecture remains usable, but the provenance plane now has to cover legibility, meaning catalog, configuration, topology, and drift visibility, and runtime evidence in addition to design-time lineage if it is to explain what the system actually became in operation. ([inference]; medium confidence; source: https://backstage.io/docs/features/software-catalog/; https://www.servicenow.com/community/cmdb-forum/cmdb-health-dashboard-completeness-compliance-amp-correctness/m-p/3488492; https://docs.dynatrace.com/docs/analyze-explore-automate/smartscape; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html)
- Delivery and platform engineering now need an explicit declared AIBOM capability, where AIBOM is an AI bill-of-materials artifact for models, data, configuration, prompts, tools, and related execution context, because both managed platforms and code-native orchestration frameworks require deliberate extraction paths before approved design can be compared with later runtime behavior. ([inference]; medium confidence; source: https://owaspaibom.org/; https://cyclonedx.org/capabilities/mlbom/; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-declared-construction-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html)
- Runtime-observed evidence has to become a first-class architecture service, because prompts, retrieved context, tool order, authority handoffs, and missing observability are governance-relevant facts that only appear after execution begins. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html)
- The orchestration layer now needs explicit delegation-aware identity and tool-action components, because portable attribution breaks when systems change credential type, cross trust boundaries, or execute under shared runtime identities without a surviving delegation receipt. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-identity-access-management-enterprise.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html)
- Operating assurance now needs named architecture components for authoritative-source binding, fairness validation, rollback authority, and external challenge escalation, because public harm repeatedly surfaced where systems looked trustworthy but source governance, deployment constraints, or oversight loops were too weak. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-04-22-knowledge-curation-governance-for-regulated-ai.html; https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html)
- Regulatory and Five Eyes updates mean that board literacy, supplier and concentration-risk governance, secure-by-design logging, input control, and regulator-facing evidence assembly must now be named operating-model capabilities rather than assumed management practices outside the architecture. ([inference]; medium confidence; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html)
- Open-weight safeguard models fit best as shared governance and evaluation services that delivery gates or higher-risk runtime checkpoints can invoke, because they offer organization-specific policy judgment while remaining too infrastructure-heavy for standard hosted-runner paths and too narrow to replace deterministic or human review layers. ([inference]; low confidence; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://docs.github.com/en/actions/reference/runners/github-hosted-runners; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html)
- Central ownership should stay with provenance schema, telemetry normalization, retained evidence, policy semantics, regulatory reporting, and evaluation standards, while domain teams keep responsibility for authoritative content, workflow composition, and risk-tuned thresholds tied to local business context. ([inference]; medium confidence; source: https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md; https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html)
- Assumption: Enterprises can emit enough runtime telemetry to reconcile approved design and observed execution at least for high-risk workflows. Justification: the runtime-capture and platform-observability items treat telemetry enablement as difficult but operationally feasible. [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html]
- Assumption: Financial-services and security-sensitive enterprises will continue to prefer centralized evidence retention and reporting even when execution remains federated. Justification: the regulatory-delta and Five Eyes items both point toward stronger centralized accountability rather than looser local proof models. [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
- Assumption: Open-weight safeguard services will stay optional because their infrastructure cost and policy-design overhead will not be justified for every workflow. Justification: the safeguard item supports selective, second-stage use rather than universal inline placement. [assumption; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html]
The second-cycle evidence makes the architecture more operational, not more abstract, because the missing pieces are evidence-handling and control-surface capabilities that only appear when systems are released, observed, challenged, and regulated in practice. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html]
The main trade-off is between central coherence and layer-local enforcement. [inference; source: https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html] Centralizing everything would hide where trust decisions really occur, while leaving every team to invent its own evidence model would destroy comparability, auditability, and incident reconstruction. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html]
The best-supported resolution is to keep control semantics and retained evidence centralized while leaving enforcement near the layer that owns the relevant trust decision, retrieval permissions in data and knowledge, semantic safeguards near inference, tool and delegation controls in orchestration, and release-time signing, extraction, and evaluation in delivery and platform engineering. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html]
A major alternative would be to keep every second-cycle addition inside the existing layers without widening either cross-cutting plane, or to replace the baseline stack with a flatter capability mesh, but the evidence weighs against both moves because the new requirements cluster around shared evidence, legibility, and retained-governance services that cut across every layer rather than belonging cleanly to one local component family. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html; https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md]
Plausible alternative placements exist for open-weight safeguard capability, especially model-layer controls and pre-deployment review gates, but the evidence supports locating the canonical policy service in the governance and evaluation plane so that delivery gates and runtime checkpoints can invoke one shared policy logic rather than duplicate policy semantics in each layer. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-02-hitl-review-volume-bottleneck-rubber-stamp.html]
Revised component list:
- Data and knowledge layer: authoritative-source registry, permission-safe retrieval, retrieval-snapshot metadata, and upstream provider-disclosure links. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-26-permission-safe-rag-enterprise-information-architecture.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-regulatory-eu-ai-act-intersection.html]
- Model and inference layer: approved model and provider registry, semantic safeguards, policy-conditioned classification hooks, and model-facing runtime signal capture. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
- Orchestration and execution layer: tool allowlists, delegation-receipt capture, action checkpoints, recursion and stop authority, and per-run authority context. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-identity-attribution-multiagent-practice.html; https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html]
- Delivery and platform-engineering layer: declared AIBOM builder, signed artifact lineage, promotion-time evaluation gates, telemetry collector pipeline, and approved-versus-observed divergence classifier. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-declared-construction-practice.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html]
- Operating-model layer: board and executive literacy, supplier and concentration-risk governance, fairness review, incident response, rollback authority, and regulator-facing evidence assembly. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html]
- Provenance-and-legibility plane: declared and observed supply-chain records, ownership-bearing catalogs, architecture blueprints, runtime topology, and structural-drift evidence. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html; https://davidamitchell.github.io/Research/research/2026-05-06-aibom-platform-observability-control-comparison.html]
- Governance, evaluation, and operational-assurance plane: policy semantics, evaluation standards, exceptions, explainability artifacts, incident records, and optional second-stage safeguard review. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-gpt-oss-safeguard-policy-enforcement-open-weight.html; https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html]
Ownership recommendations:
- Central platform, security, and risk functions should own provenance schema, telemetry normalization, retained evidence, evaluation standards, and regulatory reporting because those assets lose value when fragmented by domain. [inference; source: https://github.com/davidamitchell/Research/blob/main/Knowledge/2026-05-05-enterprise-ai-capability-stack.md; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html]
- Domain teams should own authoritative content, workflow composition, business-specific thresholds, and local exception context because those choices depend on business meaning and risk appetite that the central platform cannot infer. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-22-enterprise-ai-capability-model.html; https://davidamitchell.github.io/Research/research/2026-04-22-knowledge-curation-governance-for-regulated-ai.html]
- Shared services should expose signed interfaces and evidence contracts rather than monolithic approval queues, because second-cycle evidence favors strong shared semantics with layer-local enforcement and action-specific stop authority. [inference; source: https://davidamitchell.github.io/Research/research/2026-04-26-ai-agent-control-plane-architecture-enterprise.html; https://davidamitchell.github.io/Research/research/2026-05-07-five-eyes-ai-risks-and-advice.html]
- Runtime-evidence quality remains sensitive to platform configuration and adapter quality, so an architecture can still overstate observability if collectors, spans, or export paths are only partially enabled. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-runtime-capture-opentelemetry-practice.html]
- The architecture can reduce structural and operational blind spots without eliminating semantic failure modes, because AIBOM and telemetry remain weaker than adversarial content and authority misuse at proving behavior is safe. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-aibom-effectiveness-risk-mitigation-limits.html]
- The strongest financial-services regulatory evidence in this cycle comes from APRA and cross-jurisdiction synthesis rather than from a uniform new standard across all reviewed regulators, so some operating-model recommendations remain medium-confidence generalizations. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html]
- Legibility metrics are still composite rather than standardized, so estates may need local scorecards before cross-domain comparisons become reliably meaningful. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-it-system-legibility-measurement-frameworks.html]
- What minimum delegation-receipt schema would let orchestration runtimes preserve subject, actor, scope, target, and approval context across cross-platform tool calls?
- Which minimal composite legibility dashboard can compare catalog coverage, runtime topology coverage, divergence rates, and structural-drift findings without becoming another stale governance artifact?
- When does a policy-conditioned safeguard service justify its infrastructure cost compared with deterministic rules plus sampled human review?
- Which evidence thresholds should trigger automatic rollback versus manual escalation when approved design and observed runtime diverge in high-risk workflows?
- Type: knowledge
- Description: A second-cycle architecture update that widens the prior provenance and governance planes into explicit runtime-evidence, legibility, and operational-assurance capability families for enterprise AI systems. [inference; source: https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html; https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html]
- Links:
- https://davidamitchell.github.io/Research/research/2026-05-06-ai-capability-reference-architecture-security-supply-chain-update.html
- https://davidamitchell.github.io/Research/research/2026-05-07-ai-production-incidents-deep-dive.html
- https://davidamitchell.github.io/Research/research/2026-05-07-ai-regulatory-guidance-update-gap-check.html
Navigation
By Tag
bureaucracy
change-management
coase
constraint-analysis
control-model
decision-rights
delegation
- Q4: Decision rights that should move closer to execution
- Q5: Control model for the best throughput-risk trade-off
delivery-risk
- Operating model synthesis for split-authority delivery systems
- Q6: Leading indicators of instability in split-authority flow systems
demand-segmentation
enterprise
exception-handling
execution
flow
flow-design
flow-metrics
governance
- Operating model synthesis for split-authority delivery systems
- Q1: Dominant flow constraint in split-authority delivery systems
- Q2: Demand segmentation for fast-path vs controlled-path flow
- Q4: Decision rights that should move closer to execution
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
governance-patterns
incentives
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
instability
institutional-economics
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
leading-indicators
operating-model
organisation
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
organisational-design
queue-design
queueing
regulated-enterprise
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
- Barriers to governance reform, leadership failure modes, and reform mechanisms in regulated enterprises
routing
throughput
throughput-risk
transaction-costs
- Conditions under which internal governance controls minimise coordination costs in regulated enterprises
- Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
triage
- Q2: Demand segmentation for fast-path vs controlled-path flow
- Q3: Routing design that isolates exceptions from routine flow
williamson