Repository navigation
The release that fills the report. v0.2.0 carried six safety displays over a single study; this one carries twenty-six, and the twenty new ones are the parts of a clinical study report that were previously declared and empty — both primary efficacy endpoints, the study's one figure, the disposition tables every CSR opens with, and vital signs, weight and concomitant medications.
The claim that matters is not the count. Every one of the twenty is measured against the clinical study report the CDISC pilot published in 2006, from SAS programs sharing no code with this repository, and the comparison is a build failure rather than a note. Where a figure differs, the display says so on its own face and names which of the two is ours. Where the study contradicts itself, the display records which side it took.
-
The report's disposition section is now two tables the CDISC pilot itself published, and they agree with it cell for cell.
t-populations(DST02) andt-end-of-study(DST03) rebuild Tables 14-1.01 and 14-1.02 of the clinical study report the pilot released in 2006: the analysis populations by treatment group, and completion status with every reason for early termination. Both read the study's own ADaM packaging, which is the only one that carries the population flags and collected discontinuation reasons they report. -
The 2006 report is used as a second measurement, not as decoration.
qc/reference-report-agreement.Rrequires three routes to land on the same string for all ninety published cells: a from-scratch recomputation that reads the vendoredadsl.xpt.gzand never loads{opencsr}, the cell text parsed back out of the committed rendered HTML, and the printed report transcribed cell by cell. Any two disagreeing is a non-zero exit, and--self-testperturbs each route in turn to prove the comparison can still fail. Both run in CI. -
The transcription is the one route nothing in the repository could contradict, so it now checks itself — and it was wrong twice.
--verify-transcriptionre-reads the source document at a pinned SHA-256 and re-derives the transcribed cells from it. It found a p-value of 0.5000 recorded against a row the report leaves blank, and the two tables' population labels recorded the wrong way round. Neither was visible to the three-route comparison, because the display specifications had been written from the same misreading — all three routes agreed on the same mistake. The p-value is gone, the population lines now read as the report prints them, and the check that would have caught it earlier is committed alongside the record. -
Where the reference defines nothing, the display says so on its own face. The report prints a Complete Study row and states no definition for it; open.csr derives it as the complement of the study's own
DISCONFLand footnotes that the derivation is the project's rather than the report's, with the printed figures it reproduces as the evidence. The reference's note calling the column N the number of subjects "entered in study" is likewise not repeated, because 306 were screened and it is the 254 randomised that N counts — and the footnote says that too. -
A p-value appears on exactly the three rows the study's analysis plan names — protocol completion, adverse event and lack of efficacy — and on no other row, which is now a requirement rather than a coincidence of the spec. Nine new requirements (
DSP-POP-*,DSP-EOS-*,DSP-REF-001) carry the two displays in the matrix, verified against ADSL read straight from the vendored file rather than throughprepare_data(). -
Nine efficacy displays, and every figure in them measured three ways. The library now carries the CDISCPILOT01 ADAS-Cog programme — the Week 24 primary and its Week 8 and Week 16 companions, the Week-24 completers, the male and female subgroups, the mean-and-mean-change-over-time summary and the repeated-measures supportive analysis — together with the NPI-X secondary endpoint. Section 14.2 of the assembled report is populated where it used to be declared and empty.
-
They reproduce the reference report. 764 of the 769 published figures match the CDISC pilot's own Section 14 tables exactly, cell for cell, including the ANCOVA p-values, least-squares mean differences, standard errors and confidence intervals. All five that differ are in the repeated-measures display, all at the last digit shown, and all from one declared cause: Kenward-Roger degrees of freedom and Prasad-Rao-Jeske-Kackar-Harville standard errors are not implemented. The display says so in its own footnotes and names the five.
-
The repeated-measures model is identified, not asserted. The reference report ships its own
PROC MIXEDoutput, and the fit here reproduces it: the same 539 observations from 234 subjects, all six unstructured covariance parameters to six significant figures, and a REML criterion of 3087.843035 against SAS's 3087.84303515. Those numbers are carried into the analysis results dataset as data, because two programs can agree on a rounded least-squares mean by accident and cannot agree on that by accident. -
Qualification runs three routes, and
qc/efficacy-agreement.Rfails if any of them moves. Route A is the committed ARD; route B recomputes every statistic from the vendored transport files in base R without loading{opencsr}, solving the ANCOVA by an explicit design matrix and QR rather thanlm(), and agrees with route A on all 750 values at full double precision; route C is the reference report, extracted mechanically from the source document rather than transcribed. A new difference from the reference is a build failure even though five known ones are not. -
Where the study contradicts itself, the display says which side it took. CDISCPILOT01's analysis results metadata selects
NPTOTMNfor the NPI-X endpoint; its own reference report published a different set.NPTOTMNcovers 210 of the 222 efficacy subjects with an assessment in the Week 4 to Week 24 span — it omits twelve and adds none. The display follows the analysis plan's wording, which is also what the reference report did, reproduces that report to every cell, and records the divergence in its footnotes and inquality/data/efficacy-agreement.json. -
Two displays state a definition that came from us rather than from the study, in their own footnotes: the over-time table is presented with treatment in the columns rather than the reference's row-group layout, no number or record selection differing; and the repeated-measures least-squares means are the visit-averaged treatment main effect rather than the Week-24-conditioned estimate — which is what the reference computed, despite its title, and is a materially different quantity (1.6 with a standard error of 0.49 against 2.3 with 0.69 for placebo).
-
Analysis entries can now declare a
derive:block, for an endpoint that is a statistic of a subject rather than of a record. Acarrycolumn that varies inside a subject is an error, not a silently chosen first value. -
A trap worth naming for anyone reading
adqsnpix: it shipsAVISITright-aligned in a sixteen-character field, soAVISIT == 'Weeks 4-24'selects nothing while looking correct. The efficacy specifications select visits byAVISITNthroughout. -
The demonstration study now has all of its data, not just its safety spine. open.csr demonstrates on CDISCPILOT01, the CDISC pilot submission's xanomeline Alzheimer's study, and until now it read that study from
{pharmaverseadam}— a re-derivation that ships five domains and no efficacy data at all. The study's own ADaM package is public, and it is the one the study'sdefine.xmland data guide describe. It is now vendored in the repository: both primary efficacy endpoints (ADAS-Cog (11) and CIBIC+), the secondary NPI-X, the three laboratory datasets including Hy's-law parameters, time to first dermatologic event, and concomitant medications. Thirteen datasets are preparable where there were five. (#39) -
Which packaging each dataset comes from is a registry,
data_sources(), not a guess. Every domain{pharmaverseadam}already served stays there, so the six committed displays read byte-identical data — the prepared datasets hash the same before and after this change. Only the domains it had no answer for moved. -
The two packagings of the same study do not agree, and the repository now says exactly where. They describe the same 254 subjects and agree exactly on age, sex, race, ethnicity, site, planned treatment, treatment start and every clinical field of all 1,191 adverse events. They disagree on the actual treatment arm of twelve subjects, on how age is grouped, on the treatment-emergent flag of four adverse event records, and on the derived vital-sign records ADVS carries. Each divergence is measured and committed at
quality/data/source-agreement.json, andqc/source-agreement.Rre-measures them by a route that never loads the package and exits non-zero if the answer changes. -
Design decision D12 was wrong and has been corrected in place, dated. It read "no CDISCPILOT01 efficacy ADaM exists". That was true of
{pharmaverseadam}and false of the study. The reasoning it was built on — derive what is missing in a tested layer rather than substitute other data — is unchanged and now has less to derive. -
The efficacy analysis set is real rather than notional:
analysis_set: efficacyresolves toEFFFLand covers 234 of 254 subjects on the lane that states it. On the lane that does not, asking for it fails and names the missing flag rather than quietly returning every subject. -
Concomitant medications needed cleaning before they could be trusted. The only ADCM PHUSE publishes for this study had half its subjects relabelled into a synthetic second study; the relabelling is reversed, and every restored subject must match the subject-level table on age, sex and actual treatment or the preparation fails. It is also the one dataset here that is not part of the CDISC pilot package, and the vendored README says so.
-
The record turned out to hold across more than it was measured on. It was written under
{pharmaverseadam}1.1.0 on macOS and re-measured under 1.3.0 on Linux in CI: ninety-nine of its hundred facts reproduced exactly, and the hundredth was the version label itself — which is now recorded as environment metadata and reported rather than compared, because a version bump that moves no measured fact is news and not a failure. -
Provenance is recorded per file: upstream repository, pinned commit, upstream path, git blob SHA-1, SHA-256 and byte count, alongside the upstream MIT licence text. The vendoring script verifies each download's blob SHA against the upstream commit's tree before writing it, and CI re-checks every vendored file on every run.
Written when the data landed and no efficacy display existed yet; the nine ADAS-Cog and NPI-X displays above are that separate work. Section 11.1's prose still says no efficacy analysis dataset exists — that sentence is stale and needs a reviewed text change, which no display specification can make on its behalf.
No efficacy display is specified. The data is available; specifying a display against it is separate work, and the study's own analysis results metadata is the reference for it. Section 11.1's prose still says no efficacy analysis dataset exists — that sentence is now stale and needs a reviewed text change. (Superseded later in this release: the CIBIC+ and time-to-event displays below are specified against that data. Section 11.1's prose is still stale.)
-
The Report Template Library is plural. A template object is a document model plus a per-report assembly, and the library now holds two: the full ICH E3 clinical study report, and a new ICH E3 Annex I study synopsis. Both are assembled against the same study, CDISCPILOT01, from the same analysis results datasets and the same store of named values — so where the two documents quote the same quantity they quote the same number, and the build fails if that stops being true. (#28)
-
The assembler takes
--template <id>and--all, and CI assembles every template object rather than only the default.npm run assembleand every published link are unchanged. -
The synopsis makes the numbering rule visible: the six displays are
Table 14.1.1toListing 14.3.2.1in the report andTable 13.1toListing 13.6in the synopsis, from the same specifications and the same prose, because display identity is the slug and the number is assigned at build time. -
Efficacy fields in both documents are declared and left unpopulated rather than dropped — no efficacy ADaM exists for this study in
pharmaverseadam, and the reference report for the same dataset devotes thirteen tables to efficacy. -
The demo site publishes both documents. It reads the template library rather than one configured directory, so every assembled template object gets a reader page, a document-model page, a card on the front page and an entry in the demo's navigation tree. The report keeps
/reader/and/templates/; the synopsis is at/reader/e3-synopsis.htmland/templates/e3-synopsis.html. A third template object needs no change to the site build and none tosite/config.json. (#32) -
Every document says whether its prose has been reviewed. The approval gate holds generated-tier blocks only, so an unapproved boilerplate block still assembles — "assembled" is not "reviewed", and the site now states which one it is. The synopsis carries a notice above its first section saying all eighteen of its prose blocks are unapproved drafts that nobody has read; the report says its ten are approved. Both are read off the assembled document rather than set by hand.
-
Two more documents, and neither cost a sentence of prose. The library now holds four template objects. The post-text display package is the report's tables and listings delivered on their own — what a statistician reviews before any narrative exists — and carries no prose at all. The abbreviated clinical study report is the reduced report ICH E3 contemplates for a study not intended to support a claim: the same fifteen text blocks and the same six displays as the full report, in a document model with the efficacy-analysis apparatus and the uncited appendices removed. (#34)
-
Both are expressed as restrictions of the full ICH E3 model — every section they declare keeps the number, title, slug and content declaration it has in the report, unchanged — and a test walks both against the full model section by section, so "subset" is a property of the files rather than a claim in a comment.
-
The numbering rule now shows both of its faces. The synopsis renumbers the six displays because it declares a different structure; the display package and the abbreviated report keep
Table 14.1.1because they declare the same one. Neither file writes a number down. -
Adding the two changed nothing outside
library/templates/,quality/anddocs/: no file underscripts/names either template id, and a test enforces that. The two new documents appear on the demo site with no site change at all — declaring their titles insite/config.jsonis the one thing still outstanding. -
Both documents open in the demo, in the same viewer. The demo's Documents view holds every assembled document, one visible at a time — the arrangement the Displays view already used for the six tables. Selecting the synopsis in the explorer switches the document in place rather than leaving the demo, the trace panel behind any bound number works from either document, and each document carries its own prose notice, so the synopsis reads as unreviewed draft inside the app exactly as it does on its own page. A third and fourth template object become panels with no change to the site build. (#36)
-
The two documents are guarded against disagreeing. Every display both documents place resolves to the same analysis results dataset with a byte-identical payload, and both carry the same store of named values with no value differing — checked on the committed artifacts on every run, so a template that recomputed a number rather than citing a display fails the build before anyone reads two different figures on one screen.
-
The explorer's sidebar has one state where it had two. Each document used to carry its own disclosure arrow, remembered independently of which document was open — so an arrow could be "expanded" on a document nobody was reading, revealing nothing, and could be collapsed on the document that was open, hiding the contents the tree exists to list. The arrow is gone rather than fixed: the open document shows its table of contents, every other document shows its title, and selecting a different document moves the table of contents with it. Sections a template declares but does not populate stay dimmed, navigable and labelled, exactly as before. (#40)
-
A display says which documents use it, and a document links back to the display store. The store lists every document that places a display, with a link to each and the number that document assigned it — so the AE overview reads as
Table 14.3.1.2in the clinical study report andTable 13.4in the synopsis, side by side, which is the numbering rule made visible rather than a contradiction. Underneath the table in any document is the display's slug, linking to its store entry: inside the demo that opens the Displays view in place, and on a standalone reader page it navigates, from the same link. It is the gesture a text block already offers to reach the Text Library. (#42) -
The explorer no longer states a display number. The number belongs to the assembly, not to the display, and the sidebar serves four documents at once — so a number there was one report's fact printed where four assemblies disagree. Numbers stay everywhere they are unambiguous: inside a document, and on the display's own page beside the document that assigned each one. The list of documents using a display is now read from the whole template library rather than from the report alone, which is the same single-document assumption #36 removed from the shell, one layer down.
-
Four safety displays that reproduce a submitted report, cell for cell. The TFL Library gains the vital signs summary, the vital signs change from baseline, the weight summary and the concomitant-medication table — the four displays the reference clinical study report for this study numbers Table 14-7.01 to Table 14-7.04. Every one of the 1,341 statistics they publish is measured twice by routes sharing no code, and the 1,179 of them the reference report also prints agree with what a SAS implementation printed for the same study in 2006 -- including all 600 rendered cells of the two vital-signs tables and the weight table, which reproduce the report's printed strings character for character. Section 14 of the demonstration report now carries ten displays rather than six.
-
They are qualified on three routes, not two.
Rscript qc/vitals-conmeds-agreement.Rrecomputes every publishable statistic in the four committed analysis results datasets from the rawpharmaverseadamdata with code that shares nothing with the pipeline — its own population filter, its own baseline join, its own end-of-treatment selection, its own rounding — and compares it both with the pipeline and with the transcribed 2006 figures. It exits non-zero on any disagreement, refuses to run if the pipeline is loaded, and fails if any statistic a cell could print went unchecked. It is a required CI step, and the record it writes is committed and checked for staleness. -
The third route earned its place immediately: with a rounding tolerance a thousand times looser than the pipeline's, the independent implementation printed the median weight change in one group as
-0.5where the pipeline and the 2006 report both print-0.4. The value is-0.4499999999999890— genuinely below the half rather than a half stored imprecisely — and the pipeline was right. -
Three definitions these displays own rather than inherit, each stated in the footnotes of the table that uses it: they group by planned treatment, as the reference report does and unlike the rest of this library (twelve subjects received a treatment other than the one planned); baseline and end of treatment are derived rather than taken as shipped, because
pharmaverseadam'sBASEfalls back to a screening measurement and its end-of-treatment record exists for only 189 of 254 subjects; and the concomitant-medication table orders classes by frequency and prints percentages to one decimal where the reference did neither. -
ICH E3 gains a subsection it does not have. E3 enumerates four subsections under Section 14.3 and provides no slot for the vital signs, weight and concomitant-medication data its own Sections 12.5 and 12.6 require the report to present. The document model now declares
14.3.5, marked in the model itself as an addition rather than something taken from the guideline. The abbreviated report declares neither Section 12.4 nor 12.5 and carries neither these displays nor the laboratory ones, which is what a restriction is for.
The synopsis prose is drafted and not yet reviewed: its eighteen TXT-SYN-* blocks are draft, the site says so on the page, and the build says so on every block.
- Section 14.2 is no longer an empty heading. Five efficacy displays are specified, generated and published: the study's primary endpoint analysis of CIBIC+ at Week 24, its supporting analyses at Weeks 8 and 16, the categorical view of the same scale at all three visits, and the study's one figure — time to a first dermatologic event. Each is specified against the study's own analysis results metadata, which carries the SAS that produced the published table, and each reproduces what the sponsor's report printed: every summary statistic, every analysis-of-covariance p-value, difference and confidence limit, all 147 categorical cells, and the Kaplan-Meier medians and their intervals. (#44)
- A figure display now draws a figure. The display contract has always named
figureas a type and the engine could always compute one, but nothing drew it — the library's only figure rendered its statistics as a table under a heading that promised a curve. The Kaplan-Meier curve, its censoring tick marks, its treatment-group legend, the numbers-at-risk strip beneath it and the log-rank result annotated on its face are now drawn as inline SVG from the committed analysis results dataset and from nothing else, so the picture cannot drift from the numbers printed beneath it. No plotting package is involved, no image is embedded, and the result stays a self-contained file that diffs as text. - A figure has to survive being published, not just being rendered. The demo site embeds a rendered display by lifting its
<body>and discarding its<head>, so a plot whose colours live in a document stylesheet arrives with no strokes and nofill: none— every survival curve publishes as a solid black wedge, and nothing errors on the way there. Each drawn element now carries its appearance twice: as an SVG presentation attribute, which no sanitiser can remove, and as a rule in a stylesheet that travels inside the<svg>and adds the dark-scheme palette. The floor is a correct figure in the light palette; there is no path that reaches a reader as a black wedge, and a test holds it. - The figure's colours are the engine's, not the specification's.
display.yamlnames and orders the series; the palette that dresses them is Okabe-Ito, which stays legible under every common form of colour vision deficiency, and each series carries a dash pattern too so the figure survives being printed in greyscale. A hex value in a display specification would be a presentation choice masquerading as part of the display's definition, and it would have no dark-scheme counterpart. - A statistic that belongs to the study is no longer parked in a treatment column. The log-rank test compares the three arms; an earlier draft of this figure addressed its chi-square, degrees of freedom and p-value to the last arm's column, where a reader would take them for that arm's result. They now carry no treatment group at all, the qualification asserts the absence rather than the value, and the test is annotated on the figure where it describes all three curves.
- A median that was never reached now says so. The placebo curve never falls to one half, so its median is not estimable — and the cell was blank, which in a clinical study report reads as an omission. Rows may now declare
na_text:, and this one saysNE. - Two workers built the same figure twice, and only one of them shipped. Figure 14-1 existed in two independent implementations. They were measured against each other and against R's
{survival}package: identical to the last bit on every statistic — the same medians, the same confidence limits, the same log-rank chi-square, the same curve at every one of 200 days. The choice was decided by everything except the arithmetic — which one carried censoring into the figure, which one survived the site's embed path, which one addressed a study-level statistic correctly, and which one validated its own specification. The rejected one contributed its numbers-at-risk strip, its colour-vision-safe palette and its p-value boundary rule before it was dropped. - Qualification is three measurements, only two of them ours. A safety display is right when the pipeline and a direct recomputation agree. A display that carries a model is not: two implementations of the same misunderstanding agree perfectly. So
qc/efficacy-reference.Rrecomputes every published statistic from the vendored.xpt.gzfiles without loading{opencsr}— least squares from the normal equations rather thanstats::lm(), the Cochran-Mantel-Haenszel statistic from per-stratum score sums rather than in the form the display uses, and the Kaplan-Meier estimator, Greenwood's variance, the Brookmeyer-Crowley median limits and the log-rank test from risk sets without{survival}— and compares it both with the committed datasets and with the 2006 report. 563 comparisons, run by CI, non-zero on any disagreement. There is no--writemode: a script that could rewrite the transcription would turn the one external anchor into an echo. - Three things these displays report are not settled by the reference, and each says so in its own footnotes: the transformation behind the median's confidence limits (the report gives the intervals but not the method; only the linear form reproduces both, and all three candidates are recorded), the log-rank chi-square and degrees of freedom (open.csr's addition, annotated beside the p-value so it can be checked rather than trusted), and the choice among three statistics all answering to the phrase "CMH test" (only the row-mean-scores form reproduces the published p-values; the other two are recorded so the choice is inspectable).
- A p-value is no longer allowed to print as zero. The log-rank probability here is about 8 in a hundred trillion; rounded to the display's four decimals it printed
0.0000, which asserts that the probability is zero. Any p-value past the precision its display declares is now reported at the boundary —<0.0001, or>0.9999at the other end — while the unrounded value stays in the analysis results dataset, addressable and unrounded. - A digit plan written the obvious way was being silently discarded. YAML 1.1 reads a bare
Nornas a boolean, including as a map key, sodigits: {N: 0, n: 0}parses cleanly into a mapping with two entries both namedFALSE— and a display declaring a precision the renderer never saw would have rendered at the engine default with nothing complaining. Two specs would not parse at all and three were quietly lying. Keys that resolve to booleans are now a build failure naming the file and the path, alongside the guard that already covered row-plan values. qc/regenerate-library.Rtakes its display list from the library rather than from a list inside itself, and prepares data once per declared packaging rather than once for everything — a display specified against the study's own ADaM package must not be regenerated from the pharmaverse re-derivation. The whole library was regenerated from its specifications in a throwaway tree and compared with what is committed: 3,980 statistics across eleven displays, none of them different.
The five new displays are not registered in site/config.json, so the demo publishes them with their slug as the title and without a requirement extract or an evidence page. Registering them is outstanding.
Earlier releases
- v0.2.0 — the editing release — 2026-07-27.
- v0.1.0 — early prototype — 2026-07-26.
Requirements delivered: #44, #49, #51, #55
Drafted by Claude Code using Opus 5 and reviewed by @jwildfire