Skip to content

Releases: Siddardth7/quality-platform

v1.0.0 — M6 · Cross-platform packaging & release

Choose a tag to compare

@Siddardth7 Siddardth7 released this 11 Sep 18:28
0ee19a7

M6 makes the platform installable and reachable from outside this repository. The eight
distributions are renamed to the quality-* namespace with full PyPI metadata, pinned to one
another exactly, and built by a TestPyPI publish workflow; the MCP server gains registry
listing manifests, an npx skills add publish path, a copy-paste configuration matrix for six
named hosts, and an opt-in OAuth transport mode for web hosts. A MkDocs Material site and an
MCP-first README overhaul document the result. Nothing is published to a package index and no
MCP endpoint is hosted yet, so the affected rows stay PENDING.

Added

  • Docs site, README overhaul and worked-example write-up (#296, M6-5). A MkDocs Material
    site (mkdocs.yml + thirteen pages under docs/) publishes the quickstart, the per-engine
    pages, the full 49-tool MCP catalog, the host matrix, the SECOM worked example, the
    standards-fidelity story and a standalone limitations page. README.md is restructured
    MCP-first around that material: anchor nav, a "what it is / what it is not" table, a "why
    MCP-first" section, and a hosts summary — with the loop's edge relabelled to
    "proposed occurrence-rating / CAPA (human reviews)" so the README no longer implies the
    SPC → FMEA arrow writes anything. Honesty qualifiers carried through unchanged: analysis is
    on-demand rather than continuous, spc/results/*.json is a precondition the loop reads and
    never produces, nothing is published to a package index (#292) and no MCP endpoint is
    hosted (#355), so the web-host rows stay PENDING. A new .github/workflows/docs.yml
    deploys the site on push to main (plus workflow_dispatch) and pip installs
    mkdocs-material standalone — it is deliberately absent from pyproject.toml and
    uv.lock. Docs-only — no code, no CI / gate impact.
  • Opt-in OAuth transport mode for web MCP hosts (#355). The M1-8 HTTP transport gains an
    MCP_AUTH_MODE=oauth mode that validates WorkOS AuthKit-issued tokens, reusing FastMCP's
    AuthKitProvider (RFC 9728 protected-resource metadata + JWT verification) — no new
    dependency. The shared-secret bearer mode stays the zero-config default (unset/bearer,
    unchanged); OAuth is configured by MCP_OAUTH_AUTHKIT_DOMAIN + MCP_OAUTH_BASE_URL and
    fails closed (a missing var or unknown mode raises, naming the variable). Resource-server
    support only: a live Claude.ai/ChatGPT handshake still needs a provisioned public endpoint
    and a configured WorkOS account, so the apps/mcp/docs/HOSTS.md rows stay PENDING. Tests
    are hermetic (RSAKeyPair self-signed JWT over in-process ASGI, no live JWKS);
    mcp_app.transport stays at 100% line+branch.
  • Per-host MCP configuration matrix (#295, M6-4). apps/mcp/docs/HOSTS.md gives
    copy-paste config for the six named MCP hosts — Claude Desktop, Cursor, VS Code, Gemini CLI
    (stdio) and Claude.ai, ChatGPT (HTTP) — plus a runbook for turning a PENDING cell green,
    a per-host quirks table and the reference calls (health, version, fmea_score(8, 5, 6)
    → rpn 240) shared with skills/COMPATIBILITY.md. This is the native MCP-server
    registration
    layer that M2-6 (#275) explicitly did not attempt; COMPATIBILITY.md's
    skill-script results are referenced, not re-derived. Docs-only — no code, no CI / gate
    impact. Only Gemini CLI was actually registered (evidence):
    the host spawned the server and enumerated its tools, but the call itself was denied by the
    host's non-interactive permission gate, so its worked-example cell stays PENDING — as do
    the three GUI hosts (not launchable headless) and both web hosts, which are blocked twice
    over: no hosted endpoint is provisioned, and the M1-8 bearer-only transport has no path
    through connector UIs that expose OAuth only (#355). No cell claims PASS ✓ for a config
    that was never invoked.
  • MCP registry listing manifests (#294, M6-3). apps/mcp/server.json (official MCP
    registry — name io.github.siddardth7/quality-platform-mcp, one pypi package entry for
    quality-mcp 0.15.0 over stdio), apps/mcp/smithery.yaml (stdio startCommand reusing
    the already-documented uv run python -m mcp_app.server verbatim), and root glama.json
    (maintainers: ["Siddardth7"], for claiming Glama's auto-created listing).
    apps/mcp/README.md gains the <!-- mcp-name: ... --> ownership marker the registry
    looks for in the published package's long_description — it must match server.json's
    name exactly — plus a "Registry listings" pointer. The new apps/mcp/docs/REGISTRIES.md
    tracks all three registries, the canonical tag list, a per-registry runbook and the
    release-checklist note that server.json carries the workspace version in two places.
    All three rows are PENDING and no listing exists yet: the MCP registry verifies
    against real PyPI only, and quality-mcp is on TestPyPI alone until the v1.0.0 release
    (#292); Smithery and Glama are blocked on an SME account action. Metadata only — no
    Python code, no new dependency, no coverage-gate surface touched.
  • npx skills add publish path verified against the current agentskills.io spec (#293,
    M6-2).
    Re-checked npx skills add Siddardth7/quality-platform against the live
    specification and the vercel-labs/skills CLI behind it: no packaging change is
    required
    — GitHub is the registry, there is no manifest and no registration step, and the
    skills/<name>/SKILL.md directory is the manifest, which is the shape the repo already
    has. Docs-only diff: skills/COMPATIBILITY.md gains the #test-ref install caveat (a bare
    owner/repo install pulls the default branch, which does not yet carry quality-research)
    and PENDING matrix rows for project-loop (run_project_loop) and quality-research
    (qdb_answer_question) — PENDING meaning not yet run, with the live multi-host smoke
    test tracked as a follow-up — plus a runbook note that quality-research needs
    QDB_MCP_URL / QDB_MCP_TOKEN in the host sandbox. skills/CONVENTIONS.md §5 records the
    re-verification date. No code, CI or dependency change.
  • TestPyPI publish workflow (#292, M6-1). .github/workflows/publish.yml builds all eight
    distributions and uploads them to TestPyPI over PyPI Trusted Publishing (OIDC, no stored
    token), quality-core first and the seven that pin it second. It is workflow_dispatch-only
    with TestPyPI as the sole target — no push, merge or tag can publish as a side effect, and
    the real-PyPI path is deferred to the v1.0.0 release issue. CI / gate (ci.yml) is
    untouched. Nothing has been uploaded and uvx quality-mcp is not yet verified against a
    live index
    : Trusted Publishing is registered index-side per project, so the SME must first
    register a "pending" publisher for each of the eight names on TestPyPI and run the workflow
    once. The manual step and the verification command are in apps/mcp/README.md.
  • PyPI metadata on all eight distributions (#292, M6-1). classifiers and [project.urls]
    — the half of #261's metadata deferred to M6 — plus the missing authors on quality-mcp.
    license is still deliberately absent, and no License :: classifier was added: the repo
    still has no LICENSE file, and an identifier without one would be the same false claim #261
    rejected. Both land together before the real-PyPI publish.

Changed

  • The eight distributions are renamed to the quality-* namespace (#292, M6-1):
    fmea-app → quality-fmea, spc-app → quality-spc, msa-app → quality-msa,
    controlplan-app → quality-controlplan, secom-app → quality-secom,
    quality-database-app → quality-database, mcp-app → quality-mcp (quality-core was
    already correct). Only the distribution names changed — no import package name and no
    import statement anywhere in the repo
    (mcp_app, fmea_app, spc_app, … are
    untouched). The rename is what makes uvx quality-mcp resolvable once published: uvx X
    resolves the distribution named X, and the console script inside was already
    quality-mcp. uv.lock was regenerated; uv sync --frozen is unaffected.
  • Internal dependencies are pinned exactly (#292, M6-1). Every internal entry in a
    [project] dependencies list is now <name>==<workspace version> rather than a bare name,
    so a published wheel's Requires-Dist states the lockstep the workspace already releases in
    (one version across all eight, bumped together). [tool.uv.sources] { workspace = true } is
    unchanged and still decides where uv resolves them from locally.
    packages/quality-core/tests/test_publish_metadata.py asserts both the new distribution
    names (and that the import names did not move) and the pins, read from installed
    distribution metadata.

Known limitations at 1.0.0

  • Nothing is published to a package index yet (#292). The TestPyPI publish workflow ships in this release but has not run — the pending publishers are not registered — so uvx quality-mcp does not resolve. Install from source.
  • No MCP endpoint is hosted (#355). The opt-in OAuth transport mode is resource-server support only; a live Claude.ai / ChatGPT handshake still needs a provisioned public endpoint and a configured WorkOS account. The web-host rows in apps/mcp/docs/HOSTS.md stay PENDING.

Gate

ruff · mypy (99 source files) · pytest --cov — 2360 passed, 142 skipped, 88% overall · eleven per-surface --cov-fail-under=100 gates with branch coverage · pip-audit clean.

Full changelog: v0.16.0...v1.0.0

v0.16.0 — M5 · Quality Research Skill

Choose a tag to compare

@Siddardth7 Siddardth7 released this 22 Aug 19:03
061b876

[0.16.0] - 2026-08-21 — M5 · Quality Research Skill

The Quality Knowledge Base built in M4 becomes something an engineer can ask questions of. A
query engine turns a question into a grounded, cited answer — or refuses, by a threshold
computed before any generator runs. That engine ships as a private hosted MCP endpoint rather
than a local bundle of licensed text, and a quality-research skill calls it, degrading to a
copyright-safe local fallback when the endpoint is unreachable. Answers can also ground in the
loaded project's own computed artifacts, so "is my %GRR acceptable?" cites both the AIAG band
and the engineer's own number. A citation-accuracy gate guards the whole path in CI.

Added

  • RAG query engine — retrieve, ground, cite, refuse (#287, M5-1). quality_database_app/query.py
    turns a question into a grounded, cited answer: retrieve from the M4-3 store → deterministic
    refusal gate → ground → verify citations → enforce the quote cap. The refusal gate is computed
    from retrieval scores before any generator call, so refusing is never a model decision.

    Citations are parsed from generated text and checked for membership against the retrieved
    chunks — a cited locator absent from retrieval, or zero citations, refuses rather than answers.
    The ≤50-word quote cap (SME-locked M4 serving policy) is enforced in code and fail-closed,
    reusing evalset.QUOTABLE_FLAGS / within_quote_cap. The generator is an injectable
    Callable[[str, str], str] — no SDK, no new dependency, no live API call in importable code.
    REFUSAL_SCORE_THRESHOLD is measured, not a placeholder: calibrate_refusal_threshold.py,
    hand-run against BAAI/bge-small-en-v1.5 over the real 1157-chunk corpus, separated 11
    in-corpus gold questions (0.6647–0.8670) from 12 out-of-corpus questions (0.4422–0.6120) with
    no overlap, giving the midpoint 0.6383290503624621 at 0/11 and 0/12 misclassified — recorded
    as ASSUMPTIONS_LOG RULE 16. The script is hand-run only and stays outside the coverage gate.
    query.py is gated at 100% line+branch; tests are hermetic and do not depend on the calibrated
    value. Three mutations — refusal gate, citation membership, quote cap — were each proven to
    fail the suite by the tester and the reviewer independently.

  • Private hosted RAG endpoint — qdb_answer_question over the M1-8 HTTP transport (#288, M5-2).
    Exposes M5-1's answer_question as an MCP tool reading the private corpus, so the research
    capability ships as a hosted private endpoint rather than a local bundle of copyrighted data.
    The tool returns only the CandidateAnswer fields (item_id, text, the single verified
    locator, refused) — never a hit list, raw chunk, prompt context, or vector. The corpus never
    leaves the process.
    Auth is not re-implemented: M1-8's transport-level SharedSecretVerifier
    (bearer, fail-closed) already gates every tool, and unauthenticated HTTP calls get 401. Store
    and embedder load lazily once via lru_cache(maxsize=1), so importing the server stays
    side-effect-free in CI. The generator is resolved by import path from QDB_GENERATOR_IMPORT_PATH
    (importlib + getattr, inert until called) — wiring a real model is a deploy-time choice. Loader
    failures are converted to a structured ToolError inside the tool rather than escaping raw, so
    a client never depends on mask_error_details staying off. Deploy config is on paper only
    (no provisioning, no spend): apps/mcp/README.md carries the run command, env-var table, and a
    Fly.io hosting recommendation for the SME to action. mcp_app.server at 100% line+branch;
    three negative controls — leak a raw chunk field, refuse→answer, auth accept-any — each proven
    load-bearing by tester and reviewer independently.

  • quality-research skill — "ask the standards" cited Q&A (#289, M5-3). An engineer-facing
    skill that answers a quality-standards question with a full, standard-supported answer plus
    page-accurate citations. HTTP-first against the M5-2 qdb_answer_question endpoint,
    degrading cleanly to a copyright-safe local fallback sourced from the apps' own cited
    derivations (ASSUMPTIONS_LOG.md) when the endpoint is unreachable. Ships SKILL.md,
    scripts/call_qdb_answer_question.py, and references for the tool contract and fallback
    sources. Four citation-verified worked examples (ndc ≥ 5, FMEA no-RPN-threshold, MSA %GRR
    bands, SPC stability-before-capability) each preserve the "what the standard publishes vs. what
    the platform adds" split. Client env vars QDB_MCP_URL / QDB_MCP_TOKEN ride the M1-8 bearer
    transport. The engine-decides/skill-orchestrates invariant holds; skill-lint clean, and the ndc
    denylist was proven load-bearing by negative control.

  • Citation-accuracy CI gate for the Quality Knowledge Base (#290, M5-4). Four named
    thresholds in quality_database_app/generation_metrics.py — MIN_CITATION_ACCURACY,
    MIN_GROUNDEDNESS, MIN_REFUSAL_CORRECTNESS (floors, 1.0) and MAX_HALLUCINATION_RATE
    (ceiling, 0.0) — are enforced by apps/quality_database/tests/test_citation_gate.py, which
    scores one hand-authored answer per item of the committed 13-item docs/eval/gold_set.json
    and fails CI / gate on any regression. Three permanent negative controls prove the gate is
    load-bearing: a fabricated citation, a never-ship question answered instead of refused, and
    a barred or over-cap verbatim excerpt each push a metric across its threshold. The thresholds
    are a structural floor over a correct-by-construction fixture, not a measurement of a real
    generator
    — CI has no model and no network — and ASSUMPTIONS_LOG RULE 17 states that ceiling
    explicitly; a real-backend-calibrated gate is deferred to M6. No new CI step: the module runs
    inside the existing full-suite and Quality Database coverage-gate steps. Server-side serving
    policy (quote cap, cite-and-point, fail-closed refusal) was already shipped in M4-4/M5-1 and
    is unchanged.

  • In-project artifact grounding for quality-research (#291, M5-5). The skill reads the
    loaded project's already-computed M3 artifacts read-only, so an answer grounds in both the
    cited corpus standard and the engineer's own value — "is my %GRR acceptable?" reads
    msa/gage-rr.json's pgrr_study / ndc / verdict verbatim and cites the AIAG band against
    it. Degrades to the M5-3 standards-only flow when no project is loaded. SKILL.md gains an
    additive project-detection and read-one-artifact step (steps 2–5 unchanged, so
    fallback-sources.md's cross-reference still resolves), the read-only caveat, and an explicit
    "running a study is still not this skill's job" boundary. New references/project-context.md
    carries a topic→artifact→field table and a worked %GRR example over the real
    examples/secom-quality-loop fixture, with the AIAG bands and RULE 8 caveat quoted verbatim
    from apps/msa's ASSUMPTIONS_LOG. skill_lint.py's FORMULA_PATTERN gains
    pgrr_study|pgrr_tolerance — the new smuggle surface is the artifact's own lowercase field
    names — scoped per the per-domain denylist rule, with a mutation-verified negative control.
    Skill-only: no new engine tool, module, or coverage gate. Write-back is the post-1.0 co-pilot's job.

v0.15.0 — M2 Agent Skills · M3 closed-loop contract · M4 Quality Knowledge Base

Choose a tag to compare

@Siddardth7 Siddardth7 released this 19 Aug 22:05
952ad5c

v0.15.0 — M2 · Agent Skills · M3 · Closed-loop contract · M4 · Quality Knowledge Base

Three milestones ship together. M2 puts an Agent Skills layer over the MCP server shipped in
0.14.0 and proves it on four CLI hosts. M3 turns the four engines into one closed loop over a
project directory on disk — FMEA → Control Plan → SPC → back to the FMEA — with MSA gating what
the SPC numbers are worth. M4 stands up the private Quality Knowledge Base the M5 research
skill will query: a licensed corpus ledger, an OCR-aware ingestion pipeline, chunking/embedding
with a file-backed vector store, a citation eval set, and a fail-closed private read path.

Added

  • OCR fallback for image-only corpus PDFs (#335, M4-2c). pypdf reads an embedded text
    layer and nothing else, so the corpus' image-only scans extracted to zero characters.
    quality_database_app/ocr.py adds an Ocr Protocol seam — the same shape as M4-3's
    Embedder — and a page whose text layer falls below MIN_WORDS is retried through it. The
    default NullOcr recognizes nothing, so CI's behaviour is byte-identical to pre-#335 and
    the gate still needs no system binary. The real backend (ocr_tesseract.py, PyMuPDF
    rasterization + tesseract) sits behind the optional ocr extra, is excluded from the
    coverage gate, and is a hand-run — its output is not bit-stable across tesseract versions or
    DPI, which is stated rather than papered over. OCR supplies text only: page numbers still
    come from pypdf's page tree and extraction_quality still comes solely from the SME-reviewed
    ledger (RULE 1), so no code-derived confidence signal is ever invented. Also fixes a
    junk-page leak in extract_pdf that let below-threshold pages through as corpus records.

  • PDF text extraction for the corpus pipeline (#329, M4-2b). extract_pdf.py is the PDF
    counterpart to segment.py, emitting the same Segment shape so everything downstream of
    extraction is format-agnostic. Three narrowings, all SME-locked: one Segment per page,
    never per heading
    (raw PDF text carries no Markdown structure, so clause is None rather
    than guessed from a "first line is a heading" heuristic); page numbers are pypdf's own page
    index, 1-indexed
    , which is strictly more reliable than the Markdown path's footer-marker
    regex; and extraction never reclassifies a row — a freshly extracted PDF whose ledger
    extraction_quality is still not-extracted produces low_confidence records until the SME
    reviews the ledger by hand. A follow-up SME extraction review then marked 15 text-layer PDFs
    clean (PR #337).

  • Private corpus storage + access boundary (#286, M4-5). The corpus store is what M4-2/M4-3
    already write — the gitignored apps/quality_database/.corpus_out/ (corpus.json +
    index/), derived from the on-machine $CORPUS_ROOT tree — so M4-5 creates no new
    location and adds no dependency
    . storage.py is the one sanctioned read path:
    load_index() fails closed, so a missing index is an error rather than a silently empty
    store, and M5-2's query endpoint is meant to import it and nothing else. Two limits are
    stated rather than implied: this module cannot yet enforce single-reader access (M5-2 does
    not exist to be gated), and it adds no auth code because a local filesystem read has no
    caller identity — when M5-2 puts it on the network it should reuse M1-8's shared-secret
    bearer posture (#267), not invent a second scheme. tests/test_no_corpus_content.py is the
    machine check that keeps corpus text out of version control.

  • Citation eval set + RAG metrics (#285, M4-4). evalset.py / build_gold_set.py build a
    gold set from the existing CITATIONS.tsv ground truth; retrieval_metrics.py scores
    Recall@k / MRR / nDCG and generation_metrics.py scores the answer side, with eval.py
    running whichever half it was given the machinery for. What CI can conclude from this is
    bounded on purpose
    : FakeEmbedder's sha256-derived vectors carry no semantic similarity,
    so CI asserts plumbing — report shape, item count, filters honoured — and never a numeric
    quality threshold; real retrieval numbers are a local hand-run with the real embedder.
    never-ship items score 0 recall by design (the licensing gate makes their chunk
    unreachable) and are scored instead by refusal_correctness, where the right answer is a
    refusal. write_report exists so M5-4 can gate on a report produced that way.

  • Chunking, embedding and vector store (#284, M4-3). The retrieval half of the Quality
    Knowledge Base, on top of M4-2's ingested corpus. quality_database_app/chunk.py turns
    corpus records into retrievable Chunks — one record is one chunk by default, since
    M4-2 already split the sources on Markdown heading boundaries; only a record longer than
    MAX_CHARS (3000) splits, on paragraph boundaries first, with a hard OVERLAP_CHARS (200)
    window as the last resort for a single over-long paragraph. Every chunk carries its record's
    standard / clause / page / serving_flag / license_class verbatim, so every hit is
    citable and the licensing context never gets lost in the index. embed.py defines the
    Embedder protocol and a deterministic offline FakeEmbedder — the only embedder CI
    runs
    ; a real local model (fastembed/ONNX) sits behind the optional embed dependency
    group in embed_fastembed.py, excluded from the coverage gate, and real embedding is a
    hand-run. store.py is a file-backed VectorStore (vectors.npy + metadata.json,
    brute-force cosine scan) — no server, no ANN index, and numpy (already a quality-core
    dependency) as the only addition. search() filters on standard / source_id / region
    and excludes never-ship chunks by default: the flag stays in the data so the gap is
    auditable, while the query layer is safe by default. index.py wires corpus file -> chunks
    -> vectors -> saved index, with no default embedder so a real run can never silently produce
    a fake index. Re-indexing unchanged inputs is byte-identical, and the four new modules join
    the Quality Database CI gate at 100% line + branch.

  • OCR-aware ingestion + cleaning pipeline (#283, M4-2). The ledger becomes text:
    pipeline.py reads docs/CORPUS_LEDGER.tsv, skips (and logs) every row it cannot ingest,
    reads each remaining source — Markdown split on headings, PDF split on pages — and emits one
    CorpusRecord per segment with the ledger's confidence, serving flag and licence class
    carried through, so no downstream consumer can lose the licensing context. Cleaning is
    format-only (RULE 4): it normalises whitespace and layout artefacts and never rewrites
    content. Re-running over unchanged inputs produces a byte-identical output file, the same
    write discipline as quality_core.project.io.write_artifact. The output is corpus-derived
    text, so it is gitignored and never committed, per M4-1's "the corpus is private; only our
    derivations are public".

  • Corpus sourcing + licensing ledger (#282, M4-1). The foundation of M4 (Quality Knowledge
    Base): every source the RAG may draw on is enumerated once, with its licensing class and an
    explicit rule for what a generated answer may do with it. docs/CORPUS_LEDGER.md is the
    policy — the corpus stays private and out of the repo (only metadata and this project's own
    derivations are committed), serving flags are tiered by source type (public ISO/SAE/NIST
    standards quote; licensed AIAG/VDA handbooks and textbooks paraphrase-and-point, locator
    only; the project's own derivations serve), and a quoted excerpt in a generated answer is
    capped at 50 words / 2 sentences with a locator. docs/CORPUS_LEDGER.tsv is the
    manifest, one row per (source, region) so a partially-usable source splits — the AIAG &
    VDA FMEA Handbook's clean DFMEA prose is paraphrase-and-point while its OCR-mangled PFMEA
    Occurrence/Detection tables (#256) are never-ship. Known gaps are recorded rather than
    hidden: the AIAG SPC edition mismatch (4th Ed. cited, 2nd Ed. held), AIAG FMEA-4 cited but
    not located, and the Western Electric / Nelson possible-primaries still logged as
    reproductions. tests/test_corpus_ledger.py makes the completeness claim machine-enforced —
    no blank cells, closed vocabularies, unique keys, and the policy's cross-field rules
    (quote implies a public licence class, serve implies own derivation, not-held implies
    nothing to serve). Corpus-presence checks skip on CI, mirroring the MSA/FMEA citation tests.

  • Loop orchestration + SECOM worked example (#281, M3-6). run_project_loop(project_root)
    — a new MCP tool plus the skills/project-loop/ Agent Skill over it — sequences the four M3
    arrows in dependency order (fmea/fmea.json → control-plan/plan.json → spc/config.json →
    spc/msa-gate.json → feedback/spc-to-fmea.json + candidate Actions back on the FMEA) and
    returns exactly what each arrow's own tool returns, so the one-call path and the four
    individual calls can never disagree. Two boundaries are documented rather than fudged: the
    loop does not produce spc/results/*.json (no arrow does — those come from a prior SPC
    charting session and are read as a precondition, and with none on disk the feedback step
    legally no-ops), and a characteristic whose MSA gate says block still produces feedback,
    because nothing ties the two arrows together today — read the gate alongside the feedback
    rather than assuming the loop filtered on it. examples/secom-quality-loop/ is the runnable
    end-to-end example on the SECOM case study.

  • MSA → SPC gate (#280, M3-5). spc_app/msa_gate_arrow.py reads spc/config.json and
    msa/gage-rr.json and writes spc/msa-gate.json — one row per monitored characteristic
    saying how far its SPC result may be trusted. No gate policy lives in the arrow; the v...

Read more

v0.12.0 — Week 10 · Modern SPC depth

Choose a tag to compare

@Siddardth7 Siddardth7 released this 26 Jul 23:53
a315f1b

Week 10 · Modern SPC depth. Adds Phase I/II control-limit freezing, EWMA and CUSUM control charts, non-normal (Box-Cox / Yeo-Johnson) capability with Cp/Cpk confidence intervals, and the SPC UI wiring + run-rule gating that exposes them — with a single gated detect_violations chokepoint that blocks WE/Nelson run-rules on autocorrelated EWMA/CUSUM series from every caller.

  • Phase I/II control-limit freezing (W10-1, #141)
  • EWMA control chart (W10-2, #142)
  • CUSUM control chart — tabular two-sided (W10-3, #143)
  • Non-normal capability (Box-Cox) + Cp/Cpk CIs (W10-4, #144)
  • SPC UI wiring for Week-10 features + run-rule gating + ASSUMPTIONS_LOG (W10-5, #145)

Version note: v0.12.0 (not v0.10.0) — production already shipped v0.11.0 (Week 11 DOE); the version moves forward per SME decision, it does not regress. Full detail in CHANGELOG.md [0.12.0].

v0.11.0 — Week 11 · DOE screening on SECOM

Choose a tag to compare

@Siddardth7 Siddardth7 released this 24 Jul 10:49
215993b

Week 11 · DOE screening on SECOM — an honest screening analysis of which SECOM signals move the pass/fail response, capstoning the real-data story.

Feature (#72 · W11-1)

  • DOE screening analysis — secom_app/doe_screening.py: per-signal univariate screen over the select_signals() candidate set. Effect = Cohen's d (pooled SD, FAIL−PASS direction); significance = Welch's two-sample t (correct for the 104-fail-vs-1463-pass groups); multiple comparisons = Benjamini–Hochberg FDR, significant = q < 0.05. Reuses scipy.stats — no statistics re-derived.

Honesty over invention (the series line)

SECOM is observational process-monitoring data — factor levels are never set or randomized — so a real DOE screening design is impossible. This is a screening analysis of association, labelled unmistakably as not a designed experiment and not causal. Thresholds (α=0.05) are labelled screening conventions, not quality standards; methods cite Cohen 1988 / Welch 1947 / Benjamini–Hochberg 1995.

Result on the vendored data

463 candidate signals, 23 significant (BH q<0.05); top signal sensor_059 (Cohen's d 0.632).

Quality

All packages → 0.11.0 (v0.10.0 skipped — Week 10 had no issues). 1069 tests; coverage bars 100% — quality_core.io, quality_core.schema (line+branch), SPC, SECOM (incl. doe_screening), MSA, Control Plan. ruff + mypy clean.

Full detail: CHANGELOG.md ## [0.11.0].

v0.9.0 — Week 09 · SECOM semiconductor case study

Choose a tag to compare

@Siddardth7 Siddardth7 released this 24 Jul 06:16
e7dbe79

Week 09 · SECOM semiconductor case study — the platform's honest, end-to-end analysis of the UCI SECOM dataset (1567 wafers × 590 signals), reusing the existing SPC/MSA engines rather than re-deriving them.

Features (#65–#70)

  • #65 W09-1 — SECOM dataset ingest (NaN-preserving loader) + signal selection audit
  • #66 W09-2 — SPC I-MR control charts (gap-broken moving range, Western Electric / Nelson violations), reusing the SPC engine
  • #67 W09-3 — Cp/Cpk against caller-supplied limits, stability-gated (compute + warn on an out-of-control process)
  • #68 W09-4 — MSA applicability: an honest refusal — SECOM has no designed measurement study, so Gage R&R does not apply (executable guard + standards doc)
  • #69 W09-5 — Yield / DPPM + association (not root-cause) Pareto of failing signals, plus the first SECOM UI page
  • #70 W09-6 — SECOM case-study writeup

Engineering line held

No fabricated spec limits (SECOM ships none), MSA correctly refused, the failing-signal Pareto is association not causation, DPPM (defective units) not DPMO, missingness kept qualitative. Every case-study number is test-locked.

Quality

All packages → 0.9.0 (secom 0.7.0 → 0.9.0). 1053 tests; coverage bars 100% — quality_core.io, quality_core.schema (line+branch), SPC, SECOM, MSA, Control Plan. ruff + mypy clean.

Full detail: see CHANGELOG.md ## [0.9.0].

v0.8.0 — Week 08 · MSA / Gage R&R module

Choose a tag to compare

@Siddardth7 Siddardth7 released this 21 Jul 01:59
853214f

Week 08 adds Measurement Systems Analysis (MSA / Gage R&R) as a first-class app on the Quality Platform.

Highlights

  • Gage R&R engine — Average-and-Range method (#55). compute_gage_rr computes EV, AV, %GRR (vs study variation and vs tolerance), and ndc, returning an accept / marginal / reject verdict against AIAG thresholds. Formulas anchored to the AIAG MSA 4th-edition reference (derivation in the MSA ASSUMPTIONS_LOG).
  • MSA app UI — study entry, results, verdict + export (#56). Study-entry / results / verdict page with a loop-link note (Control Plan → MSA → SPC), a plain-English verdict sentence, and CSV/Excel/PDF export via quality_core.io. New standalone apps/msa/app.py; the platform-shell landing page gains an MSA feature card.
  • MSA tests + CI coverage gate (#57). AIAG-reference regression test (compute_gage_rr vs the published "study case 1" EV/AV/%GRR/ndc/verdict, from the new aiag_reference_study.csv fixture) and a new MSA coverage gate enforcing --cov-fail-under=100 on msa_app.gage_rr_engine + msa_app.schema + msa_app.exporter.
  • MSA scaffold + typed gage-study schema (#54).

Quality gates (green on main)

  • quality_core.io 100% · quality_core.schema 100% line+branch · SPC 100% · MSA engine/schema/exporter 100%

Release mechanics

  • All workspace packages bumped 0.7.0 → 0.8.0 (#126).
  • Milestone Week 08 · MSA / Gage R&R module complete — issues #54, #55, #56, #57 closed.

Full changelog: v0.7.0...v0.8.0

v0.7.0 — Week 07 · Close the loop

Choose a tag to compare

@Siddardth7 Siddardth7 released this 19 Jul 03:09
ce48f06

Completes the AIAG improvement loop end to end: FMEA → Control Plan → SPC → FMEA.

Highlights

  • Control Plan → SPC (#88) — a characteristic auto-configures the SPC view (spec/tolerance, sample size/frequency, recommended chart type), no manual re-entry.
  • SPC → FMEA (#89) ⭐ — an out-of-control SPC signal emits a candidate occurrence-rating / CAPA payload back to the source FMEA cause: human-in-the-loop, never auto-committed, anchored to the AIAG FMEA-4 (2008) / SAE J1739 occurrence table.
  • Loop integration tests + gate ratchet (#90) — end-to-end cross-app test on real sample data (join-key round-trip + never-auto-commit invariant); SPC coverage floor ratcheted 95 → 100%.

Quality

877 tests pass · ruff + mypy clean · coverage floors: quality_core.io 100%, quality_core.schema 100%, Control Plan 100%, SPC 100%.

Full details in CHANGELOG.md.

v0.6.0 — Week 06 Control Plan

Choose a tag to compare

@Siddardth7 Siddardth7 released this 18 Jul 08:03
2611267

FMEA → Control Plan connector engine (#84), authoring UI with injection-safe CSV/Excel/PDF export (#85), 100% line+branch coverage gate (#86), and the Control Plan app added to the mypy gate (#95). Closes the FMEA→Control Plan half of the AIAG loop. All W06 issues closed; every coverage bar green line+branch; full suite 815 passing.

v0.5.0 — Relational domain model + cross-tool schema contracts

Choose a tag to compare

@Siddardth7 Siddardth7 released this 10 Jul 06:09
0063a0a

Week 05: relational FMEA domain model + cross-tool schema contracts.

  • Schema promoted to quality_core.schema; AIAG/VDA relational model (Function → Failure Mode → Effect/Cause/Control) with loss-less flat adapters.
  • Scalar risk scoring promoted to quality_core.scoring (RPN + the AIAG-VDA Action Priority table).
  • Action tracking + effectiveness (before→after S·O·D, RPN/AP delta); end-to-end relational validate→score→export with action columns in Excel/CSV and a PDF Action Tracking page.
  • Relational + action-tracking Streamlit UI in the FMEA app; flat uploads auto-convert.
  • Engineering system adopted (Definition of Done, playbook, PR-per-issue workflow) and branch coverage turned on across every gate.

Full details in CHANGELOG.md.