Releases: Siddardth7/quality-platform
Release list
v1.0.0 — M6 · Cross-platform packaging & release
M6 makes the platform installable and reachable from outside this repository. The eight
distributions are renamed to the quality-* namespace with full PyPI metadata, pinned to one
another exactly, and built by a TestPyPI publish workflow; the MCP server gains registry
listing manifests, an npx skills add publish path, a copy-paste configuration matrix for six
named hosts, and an opt-in OAuth transport mode for web hosts. A MkDocs Material site and an
MCP-first README overhaul document the result. Nothing is published to a package index and no
MCP endpoint is hosted yet, so the affected rows stay PENDING.
Added
- Docs site, README overhaul and worked-example write-up (#296, M6-5). A MkDocs Material
site (mkdocs.yml+ thirteen pages underdocs/) publishes the quickstart, the per-engine
pages, the full 49-tool MCP catalog, the host matrix, the SECOM worked example, the
standards-fidelity story and a standalone limitations page.README.mdis restructured
MCP-first around that material: anchor nav, a "what it is / what it is not" table, a "why
MCP-first" section, and a hosts summary — with the loop's edge relabelled to
"proposed occurrence-rating / CAPA (human reviews)" so the README no longer implies the
SPC → FMEA arrow writes anything. Honesty qualifiers carried through unchanged: analysis is
on-demand rather than continuous,spc/results/*.jsonis a precondition the loop reads and
never produces, nothing is published to a package index (#292) and no MCP endpoint is
hosted (#355), so the web-host rows stayPENDING. A new.github/workflows/docs.yml
deploys the site on push tomain(plusworkflow_dispatch) andpip installs
mkdocs-materialstandalone — it is deliberately absent frompyproject.tomland
uv.lock. Docs-only — no code, noCI / gateimpact. - Opt-in OAuth transport mode for web MCP hosts (#355). The M1-8 HTTP transport gains an
MCP_AUTH_MODE=oauthmode that validates WorkOS AuthKit-issued tokens, reusing FastMCP's
AuthKitProvider(RFC 9728 protected-resource metadata + JWT verification) — no new
dependency. The shared-secret bearer mode stays the zero-config default (unset/bearer,
unchanged); OAuth is configured byMCP_OAUTH_AUTHKIT_DOMAIN+MCP_OAUTH_BASE_URLand
fails closed (a missing var or unknown mode raises, naming the variable). Resource-server
support only: a live Claude.ai/ChatGPT handshake still needs a provisioned public endpoint
and a configured WorkOS account, so theapps/mcp/docs/HOSTS.mdrows stay PENDING. Tests
are hermetic (RSAKeyPairself-signed JWT over in-process ASGI, no live JWKS);
mcp_app.transportstays at 100% line+branch. - Per-host MCP configuration matrix (#295, M6-4).
apps/mcp/docs/HOSTS.mdgives
copy-paste config for the six named MCP hosts — Claude Desktop, Cursor, VS Code, Gemini CLI
(stdio) and Claude.ai, ChatGPT (HTTP) — plus a runbook for turning aPENDINGcell green,
a per-host quirks table and the reference calls (health,version,fmea_score(8, 5, 6)
→rpn 240) shared withskills/COMPATIBILITY.md. This is the native MCP-server
registration layer that M2-6 (#275) explicitly did not attempt; COMPATIBILITY.md's
skill-script results are referenced, not re-derived. Docs-only — no code, noCI / gate
impact. Only Gemini CLI was actually registered (evidence):
the host spawned the server and enumerated its tools, but the call itself was denied by the
host's non-interactive permission gate, so its worked-example cell staysPENDING— as do
the three GUI hosts (not launchable headless) and both web hosts, which are blocked twice
over: no hosted endpoint is provisioned, and the M1-8 bearer-only transport has no path
through connector UIs that expose OAuth only (#355). No cell claimsPASS ✓for a config
that was never invoked. - MCP registry listing manifests (#294, M6-3).
apps/mcp/server.json(official MCP
registry — nameio.github.siddardth7/quality-platform-mcp, onepypipackage entry for
quality-mcp0.15.0 over stdio),apps/mcp/smithery.yaml(stdiostartCommandreusing
the already-documenteduv run python -m mcp_app.serververbatim), and rootglama.json
(maintainers: ["Siddardth7"], for claiming Glama's auto-created listing).
apps/mcp/README.mdgains the<!-- mcp-name: ... -->ownership marker the registry
looks for in the published package's long_description — it must matchserver.json's
nameexactly — plus a "Registry listings" pointer. The newapps/mcp/docs/REGISTRIES.md
tracks all three registries, the canonical tag list, a per-registry runbook and the
release-checklist note thatserver.jsoncarries the workspace version in two places.
All three rows are PENDING and no listing exists yet: the MCP registry verifies
against real PyPI only, andquality-mcpis on TestPyPI alone until the v1.0.0 release
(#292); Smithery and Glama are blocked on an SME account action. Metadata only — no
Python code, no new dependency, no coverage-gate surface touched. npx skills addpublish path verified against the current agentskills.io spec (#293,
M6-2). Re-checkednpx skills add Siddardth7/quality-platformagainst the live
specification and thevercel-labs/skillsCLI behind it: no packaging change is
required — GitHub is the registry, there is no manifest and no registration step, and the
skills/<name>/SKILL.mddirectory is the manifest, which is the shape the repo already
has. Docs-only diff:skills/COMPATIBILITY.mdgains the#test-ref install caveat (a bare
owner/repoinstall pulls the default branch, which does not yet carryquality-research)
andPENDINGmatrix rows forproject-loop(run_project_loop) andquality-research
(qdb_answer_question) —PENDINGmeaning not yet run, with the live multi-host smoke
test tracked as a follow-up — plus a runbook note thatquality-researchneeds
QDB_MCP_URL/QDB_MCP_TOKENin the host sandbox.skills/CONVENTIONS.md§5 records the
re-verification date. No code, CI or dependency change.- TestPyPI publish workflow (#292, M6-1).
.github/workflows/publish.ymlbuilds all eight
distributions and uploads them to TestPyPI over PyPI Trusted Publishing (OIDC, no stored
token),quality-corefirst and the seven that pin it second. It isworkflow_dispatch-only
with TestPyPI as the sole target — no push, merge or tag can publish as a side effect, and
the real-PyPI path is deferred to the v1.0.0 release issue.CI / gate(ci.yml) is
untouched. Nothing has been uploaded anduvx quality-mcpis not yet verified against a
live index: Trusted Publishing is registered index-side per project, so the SME must first
register a "pending" publisher for each of the eight names on TestPyPI and run the workflow
once. The manual step and the verification command are inapps/mcp/README.md. - PyPI metadata on all eight distributions (#292, M6-1).
classifiersand[project.urls]
— the half of #261's metadata deferred to M6 — plus the missingauthorsonquality-mcp.
licenseis still deliberately absent, and noLicense ::classifier was added: the repo
still has no LICENSE file, and an identifier without one would be the same false claim #261
rejected. Both land together before the real-PyPI publish.
Changed
- The eight distributions are renamed to the
quality-*namespace (#292, M6-1):
fmea-app→quality-fmea,spc-app→quality-spc,msa-app→quality-msa,
controlplan-app→quality-controlplan,secom-app→quality-secom,
quality-database-app→quality-database,mcp-app→quality-mcp(quality-corewas
already correct). Only the distribution names changed — no import package name and no
importstatement anywhere in the repo (mcp_app,fmea_app,spc_app, … are
untouched). The rename is what makesuvx quality-mcpresolvable once published:uvx X
resolves the distribution namedX, and the console script inside was already
quality-mcp.uv.lockwas regenerated;uv sync --frozenis unaffected. - Internal dependencies are pinned exactly (#292, M6-1). Every internal entry in a
[project] dependencieslist is now<name>==<workspace version>rather than a bare name,
so a published wheel'sRequires-Diststates the lockstep the workspace already releases in
(one version across all eight, bumped together).[tool.uv.sources] { workspace = true }is
unchanged and still decides where uv resolves them from locally.
packages/quality-core/tests/test_publish_metadata.pyasserts both the new distribution
names (and that the import names did not move) and the pins, read from installed
distribution metadata.
Known limitations at 1.0.0
- Nothing is published to a package index yet (#292). The TestPyPI publish workflow ships in this release but has not run — the pending publishers are not registered — so
uvx quality-mcpdoes not resolve. Install from source. - No MCP endpoint is hosted (#355). The opt-in OAuth transport mode is resource-server support only; a live Claude.ai / ChatGPT handshake still needs a provisioned public endpoint and a configured WorkOS account. The web-host rows in
apps/mcp/docs/HOSTS.mdstayPENDING.
Gate
ruff · mypy (99 source files) · pytest --cov — 2360 passed, 142 skipped, 88% overall · eleven per-surface --cov-fail-under=100 gates with branch coverage · pip-audit clean.
Full changelog: v0.16.0...v1.0.0
v0.16.0 — M5 · Quality Research Skill
[0.16.0] - 2026-08-21 — M5 · Quality Research Skill
The Quality Knowledge Base built in M4 becomes something an engineer can ask questions of. A
query engine turns a question into a grounded, cited answer — or refuses, by a threshold
computed before any generator runs. That engine ships as a private hosted MCP endpoint rather
than a local bundle of licensed text, and a quality-research skill calls it, degrading to a
copyright-safe local fallback when the endpoint is unreachable. Answers can also ground in the
loaded project's own computed artifacts, so "is my %GRR acceptable?" cites both the AIAG band
and the engineer's own number. A citation-accuracy gate guards the whole path in CI.
Added
-
RAG query engine — retrieve, ground, cite, refuse (#287, M5-1).
quality_database_app/query.py
turns a question into a grounded, cited answer: retrieve from the M4-3 store → deterministic
refusal gate → ground → verify citations → enforce the quote cap. The refusal gate is computed
from retrieval scores before any generator call, so refusing is never a model decision.
Citations are parsed from generated text and checked for membership against the retrieved
chunks — a cited locator absent from retrieval, or zero citations, refuses rather than answers.
The ≤50-word quote cap (SME-locked M4 serving policy) is enforced in code and fail-closed,
reusingevalset.QUOTABLE_FLAGS/within_quote_cap. The generator is an injectable
Callable[[str, str], str]— no SDK, no new dependency, no live API call in importable code.
REFUSAL_SCORE_THRESHOLDis measured, not a placeholder:calibrate_refusal_threshold.py,
hand-run againstBAAI/bge-small-en-v1.5over the real 1157-chunk corpus, separated 11
in-corpus gold questions (0.6647–0.8670) from 12 out-of-corpus questions (0.4422–0.6120) with
no overlap, giving the midpoint0.6383290503624621at 0/11 and 0/12 misclassified — recorded
as ASSUMPTIONS_LOG RULE 16. The script is hand-run only and stays outside the coverage gate.
query.pyis gated at 100% line+branch; tests are hermetic and do not depend on the calibrated
value. Three mutations — refusal gate, citation membership, quote cap — were each proven to
fail the suite by the tester and the reviewer independently. -
Private hosted RAG endpoint —
qdb_answer_questionover the M1-8 HTTP transport (#288, M5-2).
Exposes M5-1'sanswer_questionas an MCP tool reading the private corpus, so the research
capability ships as a hosted private endpoint rather than a local bundle of copyrighted data.
The tool returns only theCandidateAnswerfields (item_id,text, the single verified
locator,refused) — never a hit list, raw chunk, prompt context, or vector. The corpus never
leaves the process. Auth is not re-implemented: M1-8's transport-levelSharedSecretVerifier
(bearer, fail-closed) already gates every tool, and unauthenticated HTTP calls get 401. Store
and embedder load lazily once vialru_cache(maxsize=1), so importing the server stays
side-effect-free in CI. The generator is resolved by import path fromQDB_GENERATOR_IMPORT_PATH
(importlib + getattr, inert until called) — wiring a real model is a deploy-time choice. Loader
failures are converted to a structuredToolErrorinside the tool rather than escaping raw, so
a client never depends onmask_error_detailsstaying off. Deploy config is on paper only
(no provisioning, no spend):apps/mcp/README.mdcarries the run command, env-var table, and a
Fly.io hosting recommendation for the SME to action.mcp_app.serverat 100% line+branch;
three negative controls — leak a raw chunk field, refuse→answer, auth accept-any — each proven
load-bearing by tester and reviewer independently. -
quality-researchskill — "ask the standards" cited Q&A (#289, M5-3). An engineer-facing
skill that answers a quality-standards question with a full, standard-supported answer plus
page-accurate citations. HTTP-first against the M5-2qdb_answer_questionendpoint,
degrading cleanly to a copyright-safe local fallback sourced from the apps' own cited
derivations (ASSUMPTIONS_LOG.md) when the endpoint is unreachable. ShipsSKILL.md,
scripts/call_qdb_answer_question.py, and references for the tool contract and fallback
sources. Four citation-verified worked examples (ndc ≥ 5, FMEA no-RPN-threshold, MSA %GRR
bands, SPC stability-before-capability) each preserve the "what the standard publishes vs. what
the platform adds" split. Client env varsQDB_MCP_URL/QDB_MCP_TOKENride the M1-8 bearer
transport. The engine-decides/skill-orchestrates invariant holds; skill-lint clean, and the ndc
denylist was proven load-bearing by negative control. -
Citation-accuracy CI gate for the Quality Knowledge Base (#290, M5-4). Four named
thresholds inquality_database_app/generation_metrics.py—MIN_CITATION_ACCURACY,
MIN_GROUNDEDNESS,MIN_REFUSAL_CORRECTNESS(floors,1.0) andMAX_HALLUCINATION_RATE
(ceiling,0.0) — are enforced byapps/quality_database/tests/test_citation_gate.py, which
scores one hand-authored answer per item of the committed 13-itemdocs/eval/gold_set.json
and failsCI / gateon any regression. Three permanent negative controls prove the gate is
load-bearing: a fabricated citation, anever-shipquestion answered instead of refused, and
a barred or over-cap verbatim excerpt each push a metric across its threshold. The thresholds
are a structural floor over a correct-by-construction fixture, not a measurement of a real
generator — CI has no model and no network — and ASSUMPTIONS_LOG RULE 17 states that ceiling
explicitly; a real-backend-calibrated gate is deferred to M6. No new CI step: the module runs
inside the existing full-suite and Quality Database coverage-gate steps. Server-side serving
policy (quote cap, cite-and-point, fail-closed refusal) was already shipped in M4-4/M5-1 and
is unchanged. -
In-project artifact grounding for
quality-research(#291, M5-5). The skill reads the
loaded project's already-computed M3 artifacts read-only, so an answer grounds in both the
cited corpus standard and the engineer's own value — "is my %GRR acceptable?" reads
msa/gage-rr.json'spgrr_study/ndc/verdictverbatim and cites the AIAG band against
it. Degrades to the M5-3 standards-only flow when no project is loaded.SKILL.mdgains an
additive project-detection and read-one-artifact step (steps 2–5 unchanged, so
fallback-sources.md's cross-reference still resolves), the read-only caveat, and an explicit
"running a study is still not this skill's job" boundary. Newreferences/project-context.md
carries a topic→artifact→field table and a worked %GRR example over the real
examples/secom-quality-loopfixture, with the AIAG bands and RULE 8 caveat quoted verbatim
fromapps/msa's ASSUMPTIONS_LOG.skill_lint.py'sFORMULA_PATTERNgains
pgrr_study|pgrr_tolerance— the new smuggle surface is the artifact's own lowercase field
names — scoped per the per-domain denylist rule, with a mutation-verified negative control.
Skill-only: no new engine tool, module, or coverage gate. Write-back is the post-1.0 co-pilot's job.
v0.15.0 — M2 Agent Skills · M3 closed-loop contract · M4 Quality Knowledge Base
v0.15.0 — M2 · Agent Skills · M3 · Closed-loop contract · M4 · Quality Knowledge Base
Three milestones ship together. M2 puts an Agent Skills layer over the MCP server shipped in
0.14.0 and proves it on four CLI hosts. M3 turns the four engines into one closed loop over a
project directory on disk — FMEA → Control Plan → SPC → back to the FMEA — with MSA gating what
the SPC numbers are worth. M4 stands up the private Quality Knowledge Base the M5 research
skill will query: a licensed corpus ledger, an OCR-aware ingestion pipeline, chunking/embedding
with a file-backed vector store, a citation eval set, and a fail-closed private read path.
Added
-
OCR fallback for image-only corpus PDFs (#335, M4-2c).
pypdfreads an embedded text
layer and nothing else, so the corpus' image-only scans extracted to zero characters.
quality_database_app/ocr.pyadds anOcrProtocol seam — the same shape as M4-3's
Embedder— and a page whose text layer falls belowMIN_WORDSis retried through it. The
defaultNullOcrrecognizes nothing, so CI's behaviour is byte-identical to pre-#335 and
the gate still needs no system binary. The real backend (ocr_tesseract.py, PyMuPDF
rasterization +tesseract) sits behind the optionalocrextra, is excluded from the
coverage gate, and is a hand-run — its output is not bit-stable across tesseract versions or
DPI, which is stated rather than papered over. OCR supplies text only: page numbers still
come from pypdf's page tree andextraction_qualitystill comes solely from the SME-reviewed
ledger (RULE 1), so no code-derived confidence signal is ever invented. Also fixes a
junk-page leak inextract_pdfthat let below-threshold pages through as corpus records. -
PDF text extraction for the corpus pipeline (#329, M4-2b).
extract_pdf.pyis the PDF
counterpart tosegment.py, emitting the sameSegmentshape so everything downstream of
extraction is format-agnostic. Three narrowings, all SME-locked: one Segment per page,
never per heading (raw PDF text carries no Markdown structure, soclauseisNonerather
than guessed from a "first line is a heading" heuristic); page numbers are pypdf's own page
index, 1-indexed, which is strictly more reliable than the Markdown path's footer-marker
regex; and extraction never reclassifies a row — a freshly extracted PDF whose ledger
extraction_qualityis stillnot-extractedproduceslow_confidencerecords until the SME
reviews the ledger by hand. A follow-up SME extraction review then marked 15 text-layer PDFs
clean (PR #337). -
Private corpus storage + access boundary (#286, M4-5). The corpus store is what M4-2/M4-3
already write — the gitignoredapps/quality_database/.corpus_out/(corpus.json+
index/), derived from the on-machine$CORPUS_ROOTtree — so M4-5 creates no new
location and adds no dependency.storage.pyis the one sanctioned read path:
load_index()fails closed, so a missing index is an error rather than a silently empty
store, and M5-2's query endpoint is meant to import it and nothing else. Two limits are
stated rather than implied: this module cannot yet enforce single-reader access (M5-2 does
not exist to be gated), and it adds no auth code because a local filesystem read has no
caller identity — when M5-2 puts it on the network it should reuse M1-8's shared-secret
bearer posture (#267), not invent a second scheme.tests/test_no_corpus_content.pyis the
machine check that keeps corpus text out of version control. -
Citation eval set + RAG metrics (#285, M4-4).
evalset.py/build_gold_set.pybuild a
gold set from the existingCITATIONS.tsvground truth;retrieval_metrics.pyscores
Recall@k / MRR / nDCG andgeneration_metrics.pyscores the answer side, witheval.py
running whichever half it was given the machinery for. What CI can conclude from this is
bounded on purpose:FakeEmbedder's sha256-derived vectors carry no semantic similarity,
so CI asserts plumbing — report shape, item count, filters honoured — and never a numeric
quality threshold; real retrieval numbers are a local hand-run with the real embedder.
never-shipitems score 0 recall by design (the licensing gate makes their chunk
unreachable) and are scored instead byrefusal_correctness, where the right answer is a
refusal.write_reportexists so M5-4 can gate on a report produced that way. -
Chunking, embedding and vector store (#284, M4-3). The retrieval half of the Quality
Knowledge Base, on top of M4-2's ingested corpus.quality_database_app/chunk.pyturns
corpus records into retrievableChunks — one record is one chunk by default, since
M4-2 already split the sources on Markdown heading boundaries; only a record longer than
MAX_CHARS(3000) splits, on paragraph boundaries first, with a hardOVERLAP_CHARS(200)
window as the last resort for a single over-long paragraph. Every chunk carries its record's
standard/clause/page/serving_flag/license_classverbatim, so every hit is
citable and the licensing context never gets lost in the index.embed.pydefines the
Embedderprotocol and a deterministic offlineFakeEmbedder— the only embedder CI
runs; a real local model (fastembed/ONNX) sits behind the optionalembeddependency
group inembed_fastembed.py, excluded from the coverage gate, and real embedding is a
hand-run.store.pyis a file-backedVectorStore(vectors.npy+metadata.json,
brute-force cosine scan) — no server, no ANN index, andnumpy(already aquality-core
dependency) as the only addition.search()filters onstandard/source_id/region
and excludesnever-shipchunks by default: the flag stays in the data so the gap is
auditable, while the query layer is safe by default.index.pywires corpus file -> chunks
-> vectors -> saved index, with no default embedder so a real run can never silently produce
a fake index. Re-indexing unchanged inputs is byte-identical, and the four new modules join
the Quality Database CI gate at 100% line + branch. -
OCR-aware ingestion + cleaning pipeline (#283, M4-2). The ledger becomes text:
pipeline.pyreadsdocs/CORPUS_LEDGER.tsv, skips (and logs) every row it cannot ingest,
reads each remaining source — Markdown split on headings, PDF split on pages — and emits one
CorpusRecordper segment with the ledger's confidence, serving flag and licence class
carried through, so no downstream consumer can lose the licensing context. Cleaning is
format-only (RULE 4): it normalises whitespace and layout artefacts and never rewrites
content. Re-running over unchanged inputs produces a byte-identical output file, the same
write discipline asquality_core.project.io.write_artifact. The output is corpus-derived
text, so it is gitignored and never committed, per M4-1's "the corpus is private; only our
derivations are public". -
Corpus sourcing + licensing ledger (#282, M4-1). The foundation of M4 (Quality Knowledge
Base): every source the RAG may draw on is enumerated once, with its licensing class and an
explicit rule for what a generated answer may do with it.docs/CORPUS_LEDGER.mdis the
policy — the corpus stays private and out of the repo (only metadata and this project's own
derivations are committed), serving flags are tiered by source type (public ISO/SAE/NIST
standardsquote; licensed AIAG/VDA handbooks and textbooksparaphrase-and-point, locator
only; the project's own derivationsserve), and a quoted excerpt in a generated answer is
capped at 50 words / 2 sentences with a locator.docs/CORPUS_LEDGER.tsvis the
manifest, one row per (source, region) so a partially-usable source splits — the AIAG &
VDA FMEA Handbook's clean DFMEA prose isparaphrase-and-pointwhile its OCR-mangled PFMEA
Occurrence/Detection tables (#256) arenever-ship. Known gaps are recorded rather than
hidden: the AIAG SPC edition mismatch (4th Ed. cited, 2nd Ed. held), AIAG FMEA-4 cited but
not located, and the Western Electric / Nelson possible-primaries still logged as
reproductions.tests/test_corpus_ledger.pymakes the completeness claim machine-enforced —
no blank cells, closed vocabularies, unique keys, and the policy's cross-field rules
(quoteimplies a public licence class,serveimplies own derivation,not-heldimplies
nothing to serve). Corpus-presence checks skip on CI, mirroring the MSA/FMEA citation tests. -
Loop orchestration + SECOM worked example (#281, M3-6).
run_project_loop(project_root)
— a new MCP tool plus theskills/project-loop/Agent Skill over it — sequences the four M3
arrows in dependency order (fmea/fmea.json→control-plan/plan.json→spc/config.json→
spc/msa-gate.json→feedback/spc-to-fmea.json+ candidateActions back on the FMEA) and
returns exactly what each arrow's own tool returns, so the one-call path and the four
individual calls can never disagree. Two boundaries are documented rather than fudged: the
loop does not producespc/results/*.json(no arrow does — those come from a prior SPC
charting session and are read as a precondition, and with none on disk the feedback step
legally no-ops), and a characteristic whose MSA gate saysblockstill produces feedback,
because nothing ties the two arrows together today — read the gate alongside the feedback
rather than assuming the loop filtered on it.examples/secom-quality-loop/is the runnable
end-to-end example on the SECOM case study. -
MSA → SPC gate (#280, M3-5).
spc_app/msa_gate_arrow.pyreadsspc/config.jsonand
msa/gage-rr.jsonand writesspc/msa-gate.json— one row per monitored characteristic
saying how far its SPC result may be trusted. No gate policy lives in the arrow; the v...
v0.12.0 — Week 10 · Modern SPC depth
Week 10 · Modern SPC depth. Adds Phase I/II control-limit freezing, EWMA and CUSUM control charts, non-normal (Box-Cox / Yeo-Johnson) capability with Cp/Cpk confidence intervals, and the SPC UI wiring + run-rule gating that exposes them — with a single gated detect_violations chokepoint that blocks WE/Nelson run-rules on autocorrelated EWMA/CUSUM series from every caller.
- Phase I/II control-limit freezing (W10-1, #141)
- EWMA control chart (W10-2, #142)
- CUSUM control chart — tabular two-sided (W10-3, #143)
- Non-normal capability (Box-Cox) + Cp/Cpk CIs (W10-4, #144)
- SPC UI wiring for Week-10 features + run-rule gating + ASSUMPTIONS_LOG (W10-5, #145)
Version note: v0.12.0 (not v0.10.0) — production already shipped v0.11.0 (Week 11 DOE); the version moves forward per SME decision, it does not regress. Full detail in CHANGELOG.md [0.12.0].
v0.11.0 — Week 11 · DOE screening on SECOM
Week 11 · DOE screening on SECOM — an honest screening analysis of which SECOM signals move the pass/fail response, capstoning the real-data story.
Feature (#72 · W11-1)
- DOE screening analysis —
secom_app/doe_screening.py: per-signal univariate screen over theselect_signals()candidate set. Effect = Cohen's d (pooled SD, FAIL−PASS direction); significance = Welch's two-sample t (correct for the 104-fail-vs-1463-pass groups); multiple comparisons = Benjamini–Hochberg FDR,significant = q < 0.05. Reusesscipy.stats— no statistics re-derived.
Honesty over invention (the series line)
SECOM is observational process-monitoring data — factor levels are never set or randomized — so a real DOE screening design is impossible. This is a screening analysis of association, labelled unmistakably as not a designed experiment and not causal. Thresholds (α=0.05) are labelled screening conventions, not quality standards; methods cite Cohen 1988 / Welch 1947 / Benjamini–Hochberg 1995.
Result on the vendored data
463 candidate signals, 23 significant (BH q<0.05); top signal sensor_059 (Cohen's d 0.632).
Quality
All packages → 0.11.0 (v0.10.0 skipped — Week 10 had no issues). 1069 tests; coverage bars 100% — quality_core.io, quality_core.schema (line+branch), SPC, SECOM (incl. doe_screening), MSA, Control Plan. ruff + mypy clean.
Full detail: CHANGELOG.md ## [0.11.0].
v0.9.0 — Week 09 · SECOM semiconductor case study
Week 09 · SECOM semiconductor case study — the platform's honest, end-to-end analysis of the UCI SECOM dataset (1567 wafers × 590 signals), reusing the existing SPC/MSA engines rather than re-deriving them.
Features (#65–#70)
- #65 W09-1 — SECOM dataset ingest (NaN-preserving loader) + signal selection audit
- #66 W09-2 — SPC I-MR control charts (gap-broken moving range, Western Electric / Nelson violations), reusing the SPC engine
- #67 W09-3 — Cp/Cpk against caller-supplied limits, stability-gated (compute + warn on an out-of-control process)
- #68 W09-4 — MSA applicability: an honest refusal — SECOM has no designed measurement study, so Gage R&R does not apply (executable guard + standards doc)
- #69 W09-5 — Yield / DPPM + association (not root-cause) Pareto of failing signals, plus the first SECOM UI page
- #70 W09-6 — SECOM case-study writeup
Engineering line held
No fabricated spec limits (SECOM ships none), MSA correctly refused, the failing-signal Pareto is association not causation, DPPM (defective units) not DPMO, missingness kept qualitative. Every case-study number is test-locked.
Quality
All packages → 0.9.0 (secom 0.7.0 → 0.9.0). 1053 tests; coverage bars 100% — quality_core.io, quality_core.schema (line+branch), SPC, SECOM, MSA, Control Plan. ruff + mypy clean.
Full detail: see CHANGELOG.md ## [0.9.0].
v0.8.0 — Week 08 · MSA / Gage R&R module
Week 08 adds Measurement Systems Analysis (MSA / Gage R&R) as a first-class app on the Quality Platform.
Highlights
- Gage R&R engine — Average-and-Range method (#55).
compute_gage_rrcomputes EV, AV, %GRR (vs study variation and vs tolerance), and ndc, returning an accept / marginal / reject verdict against AIAG thresholds. Formulas anchored to the AIAG MSA 4th-edition reference (derivation in the MSAASSUMPTIONS_LOG). - MSA app UI — study entry, results, verdict + export (#56). Study-entry / results / verdict page with a loop-link note (Control Plan → MSA → SPC), a plain-English verdict sentence, and CSV/Excel/PDF export via
quality_core.io. New standaloneapps/msa/app.py; the platform-shell landing page gains an MSA feature card. - MSA tests + CI coverage gate (#57). AIAG-reference regression test (
compute_gage_rrvs the published "study case 1" EV/AV/%GRR/ndc/verdict, from the newaiag_reference_study.csvfixture) and a new MSA coverage gate enforcing--cov-fail-under=100onmsa_app.gage_rr_engine+msa_app.schema+msa_app.exporter. - MSA scaffold + typed gage-study schema (#54).
Quality gates (green on main)
quality_core.io100% ·quality_core.schema100% line+branch · SPC 100% · MSA engine/schema/exporter 100%
Release mechanics
- All workspace packages bumped
0.7.0 → 0.8.0(#126). - Milestone Week 08 · MSA / Gage R&R module complete — issues #54, #55, #56, #57 closed.
Full changelog: v0.7.0...v0.8.0
v0.7.0 — Week 07 · Close the loop
Completes the AIAG improvement loop end to end: FMEA → Control Plan → SPC → FMEA.
Highlights
- Control Plan → SPC (#88) — a characteristic auto-configures the SPC view (spec/tolerance, sample size/frequency, recommended chart type), no manual re-entry.
- SPC → FMEA (#89) ⭐ — an out-of-control SPC signal emits a candidate occurrence-rating / CAPA payload back to the source FMEA cause: human-in-the-loop, never auto-committed, anchored to the AIAG FMEA-4 (2008) / SAE J1739 occurrence table.
- Loop integration tests + gate ratchet (#90) — end-to-end cross-app test on real sample data (join-key round-trip + never-auto-commit invariant); SPC coverage floor ratcheted 95 → 100%.
Quality
877 tests pass · ruff + mypy clean · coverage floors: quality_core.io 100%, quality_core.schema 100%, Control Plan 100%, SPC 100%.
Full details in CHANGELOG.md.
v0.6.0 — Week 06 Control Plan
FMEA → Control Plan connector engine (#84), authoring UI with injection-safe CSV/Excel/PDF export (#85), 100% line+branch coverage gate (#86), and the Control Plan app added to the mypy gate (#95). Closes the FMEA→Control Plan half of the AIAG loop. All W06 issues closed; every coverage bar green line+branch; full suite 815 passing.
v0.5.0 — Relational domain model + cross-tool schema contracts
Week 05: relational FMEA domain model + cross-tool schema contracts.
- Schema promoted to
quality_core.schema; AIAG/VDA relational model (Function → Failure Mode → Effect/Cause/Control) with loss-less flat adapters. - Scalar risk scoring promoted to
quality_core.scoring(RPN + the AIAG-VDA Action Priority table). - Action tracking + effectiveness (before→after S·O·D, RPN/AP delta); end-to-end relational validate→score→export with action columns in Excel/CSV and a PDF Action Tracking page.
- Relational + action-tracking Streamlit UI in the FMEA app; flat uploads auto-convert.
- Engineering system adopted (Definition of Done, playbook, PR-per-issue workflow) and branch coverage turned on across every gate.
Full details in CHANGELOG.md.