Skip to content

Releases: CrossRead/scholion

v0.5.0

Choose a tag to compare

@l1nkberry l1nkberry released this 12 Sep 23:54

What you can do now

Every body system now has its genetic half, composed from a base with a
version, without waiting for anyone's signature.
Eleven systems of the radar
carry a gene list taken from the Gene Curation Coalition's export: which genes
have an asserted link to the diseases of that system, who asserted it, how
strongly, under which mode of inheritance, and on what date — 945 genes in
all, from 31 for the growth axis to 218 for the kidneys, every submitter's row
kept side by side. The disease groups behind each list are named per system
and tuned title by title against the export; a term that matched the wrong
things — a bare «gout» that reached a neurological syndrome, «short stature»
that reached sixty syndromic genes — was vetoed by name, and the tool now
carries such vetoes so a refresh keeps them. The primary analysis runs on
every system at once: what was read, what was not, what ClinVar holds, and the
questions that follow. A clinician's signature is still what the curated
layer waits for — the positions, the phrases per genotype state, the signed
exclusions — but a list composed from the base no longer waits for it.

COMT rs4680 stands on the radar, under the adrenal axis, and says what it
does not decide.
The catalogue held the locus since the previous release;
the adrenal system now carries it as an authored position: one or two copies
of Met are printed with the sentence that CPIC names the gene in its opioid
guideline and issues no recommendation by it, that the behavioural and
cognitive associations of the 2000s do not reproduce, and that nothing is
decided by the genotype — it is shown because it is asked about. It is printed, and it is not a finding: a
position never read prints as unread, and the named allele not found prints as
read and absent, with the depth.

Polygenic scores stand beside every body system, as scores. Of the seventy
pinned polygenic models, thirty-eight are placed with the system they measure —
LDL and coronary disease with the lipids, type 2 diabetes and glycated
haemoglobin with carbohydrate metabolism, Crohn's and lupus with inflammation,
gout and kidney disease with the kidneys — and each system's card prints its
scores under their own heading, never interleaved with the gene rows and never
inside the 0–100 index or the verdict: a genotype cannot be refuted by the next
blood draw. A reliable score at or above the eightieth percentile becomes one
question for the clinician — whether a screening is worth discussing — with the
caveats in the sentence: a percentile is not a probability, and the panels are
mostly European. The thirty-two models that fit no system of the radar are
named, with the reason, in the file that maps them. A profile with no scores
says so, and says that scores are computed from a full genome.

The card says what a full genome would add. For a person whose genome file
is an array, an exome or a panel — or who has none yet — the «to read in the
genome» basket names, per system, the genes of the list a full genome would
read and the polygenic scores it would make possible. A full genome prints
nothing there, because nothing is missing.

The «Second opinion» tab is now «Radar», and it carries the whole picture.
The figure and the radar with two rings per system, and one block per system
below them: what is outside the corridor, the genetic half composed from a base
with a version and how much of it was read, the polygenic scores, the
prescriptions acting on it, what to test and the questions for the clinician —
with the system's full card one click away. An explanation of what the rings
and the genetic half mean stands on the page itself, and the page still prints
as the sheet to take to an appointment.

A body system answers as one card on every face. Click a segment of the
radar or an organ on the figure, run scholion system thyroid, ask the local
API or the assistant's tool, and the same card comes back: the laboratory now
and its movement since the previous draw, the genetic half and how much of it
was actually read, the prescriptions acting on the system, the target a
clinician set, what to test, and the questions to bring to the appointment. The
next step comes in three baskets — laboratory, genome, ask — and a basket that
is empty says why it is empty. The card has two densities, one for the patient
and one for the clinician; they differ in how much is said, never in the
verdict, and no line in either is an instruction.

The second look is laid out by body system. What is outside the corridor,
what is prescribed for that system, what to test and what to ask now stand
together under one heading per system, where before the thyroid appeared three
times in three layers without being called one subject. Each block prints
whole, and a system with nothing to say is named as silent rather than left
out.

Every system carries two rings, never merged. How much of its laboratory
panel has been measured, and how much of its genetic half has been read, drawn
as two rings on the figure and on the radar. A system whose genetic list is not
composed shows that ring dotted — an absence, not a zero.

Markers, genes and prescriptions point to their system. A marker card says
which system it belongs to and opens that system's card; a gene says in which
systems it is named; a prescription says which system it acts on. A
prescription's genes now come from the system it acts on, and a clinician's own
rows about that drug stand as exceptions on top — a sentence, a class of link,
or a signed exclusion; an unsigned exclusion is refused and counted. Thirteen
drug classes are mapped to a system; the classes whose target is a use rather
than an organ — blood pressure, anticoagulation, pain — are left unmapped, and
the file says why.

Three entries refuse in one voice. A prescription, a class of disease and a
body system all judge their gene list through one form: the same gate (printed
only with a sentence and a source, the rest counted), the same pending row, the
same four-state verdict with the number of unread genes inside it, and the same
sentence. A list that was not read end to end is never called clear, whichever
door was used.

The genetic half of a body system is composed from a base with a version, and
a clinician signs exceptions.
For the thyroid the monogenic composition comes
from GenCC — 38 genes, every submitter's assertion side by side with its
classification, mode of inheritance and date, and the export the list came
from printed with it. A clinician's own positions, phrases and exclusions are
read from a curated file that ships empty and says why. Ten systems say plainly
that their genetic half is not composed; the twelfth, built from wearables,
says it has none by design — three different answers where an empty block used
to be the only one.

What the engine says about genetics is bounded. Every row carries its
evidence mode and the three modes never print alike; an assertion classified
Limited, Disputed or Refuted is never a finding; one copy of an allele in a
recessive gene is printed as carriership and raised as a question, never as a
risk line; a single common variant prints only with an effect size from a named
study; and genetics does not enter the 0–100 score — a genotype cannot be
refuted by the next blood draw, so it stands beside the score, not inside it.

A target your treating clinician set has a place of its own. «TSH between
1 and 2», «free T3 5.0» — beside the reference interval printed on the form and
separate from your personal goals. You enter it with who set it and when; a
target without either is refused, and the product never proposes a figure of
its own. The figures convert to the marker's standard unit the way a laboratory
value does. The labs report, the marker cards and the marker charts show the
target beside the corridor, drawn in a different stroke so the laboratory's
range and the treatment's aim are never confused. When a value sits inside the
corridor but outside the target, the product asks whether it is worth
discussing at the next visit — a question, not an instruction. Standing outside
a target is never counted as an abnormality: the flag and the out-of-range
count are exactly what they were before. Targets live in their own file, so
re-importing a folder of forms rewrites the series and leaves them in place;
scholion target and a matching block on the labs page list, record and
withdraw them, and name in one place the values that are in the corridor yet
outside a target.

The monogenic half of a system's gene list can come from a curated base
instead of from one person's memory.
A single command pulls the Gene
Curation Coalition's weekly export, keeps only the genes whose diseases belong
to a system, and writes them down with everything a reader needs to weigh each
line: who asserted the link, how strongly, on what date, under which mode of
inheritance, and which version of the base it came from. Every submitter's row
is kept, so two groups disagreeing about one gene are shown side by side rather
than averaged, and a weak or refuted assertion travels marked as such rather
than being silently dropped or silently promoted. The thyroid system ships
composed this way — thirty-eight genes with the recessive ones named recessive,
which is the difference between «carrier» and «at risk» for a heterozygote. A
freshness check says how far behind the base the shipped copy is, and the tool
refuses to touch the network when told to work offline. The other ten systems
remain unfilled on purpose: their terms are entered from a named panel or
report, and none has been named yet, so each says so instead of looking like a
system nobody asked about.

A gene answers with what has been decided about it, not only with genotypes.
Asked about a gene, the build now prints the curated verdict written about that
gene before the findings and again under them — includi...

Read more

v0.4.11

Choose a tag to compare

@l1nkberry l1nkberry released this 11 Sep 23:09

What you can do now

A variant file with no index is read. Every reader of a genome file used to
need an index beside it, and the tools that build one are not on a physician's
machine. Such a file is now read once from beginning to end — the catalogue's
positions in both builds, the counts that tell what kind of file it is, and the
header — and no index is written, because one written subtly wrong would be
trusted by every other tool. What that reading cannot answer is refused by
name: a question about a whole region says it needs an index and names the two
commands that make one, rather than answering with an empty list that reads as
«this gene carries no variants».

The screen for what is worth acting on runs from an installed package. The
secondary-findings screen over the 84 ACMG genes does not compute when asked: it
reads a table, and the pass that writes that table lived in the data-preparation
directory, which does not travel with pip install. scholion acmg-scan runs it
now, with no external tools and no index, from the person's variant file and the
published ClinVar file; when that file is missing the command prints the one
download it needs, for the build the person's file is in. Each file's build is
established from the file itself, and a crossed pair — a GRCh37 genome against a
GRCh38 ClinVar — is refused before a single position is compared, because a
crossed pair does not fail: it finds nothing, or matches a position that belongs
to a different base, and both look like an answer.

An exome is told apart from a panel. Breadth used to be probed in windows
that are gene-poor by construction, where an exome is empty, so an exome was
classed as a sparse panel — and that class shuts the ClinVar and ACMG paths, on
the one input where a screen for known pathogenic variants is most obviously
worth running. A second set of probes in gene-dense windows tells the two apart
by contrast. The two paths open on an exome and say what an exome can and cannot
answer; polygenic scores stay shut, because a score summed over the coding two
per cent of the genome has no distribution behind it.

The status says which questions this input can carry, before any finding.
Six paths — catalogue loci, pharmacogenetics, a whole region, ClinVar, the ACMG
panel, polygenic scores — each marked open or shut for this file, with the
reason; «this input cannot carry the answer» is told apart from «the annotation
this path reads has not been produced», which used to be invisible because the
narrower gate fired first and blamed the file for a missing table.

A gene is asked of every shelf that holds anything about one. What this build
knows about a gene sits in four places, each keyed differently — the curated
catalogue by rsID, the ClinVar scan by coordinate, the shipped ACMG secondary
findings panel by symbol, coverage by gene — and nothing joined them. Asking about
a gene read the catalogue alone and answered «not in the coordinate reference»,
which is a statement about one shelf and reads as a statement about the genome. A
clinician asked about the bile-acid transporters and was told there was nothing to
say «even in general terms from your genome», while the ClinVar table on that
machine held 386 findings.

scholion genome --gene X now prints a frame before the findings: what each layer holds,
that it holds nothing, or that it could not be asked and what would let it be. A
gene in the ACMG panel is recognised with no network and no annotation file,
because those 84 symbols travel inside the build — so a question about BRCA1 on a
machine that cannot reach Ensembl gets an answer rather than a shrug.

ClinVar can be asked by gene. scholion clinvar --gene X matches findings to a gene by
COORDINATE — the scan table carries no gene column, and this needs no re-scan of
the genome. A gene whose coordinates cannot be obtained is reported as unresolved,
with what would resolve it, rather than as a gene with no findings.

What is fixed

A row at the position was taken for the locus. The reader took the first
variant row at a catalogue position and built the genotype from that row's own
alleles. A coordinate is not an identity: an insertion, a neighbouring
substitution or a multi-allelic row can stand on the same base, and the genotype
printed then belonged to somebody else's variant while carrying the locus's name
and the label «called». On three clinical files the alleles at the Factor V
Leiden position disagreed with the catalogue in every one. The row is now
chosen, not taken — its alleles must be the catalogue's — and the two ways of
failing to find one are named: another variant on this base, and a reference
base that is not ours at all, which is a statement about the file's build, never
about the person.

A genome file with no index that had been cut short was read up to the cut
and answered «reference» at every position past it.
A copy or a download that
did not finish leaves such a file; a heterozygous APOE ε4 carrier could be
printed as a non-carrier under a status line saying the genome was connected.
Such a file is now refused by name — the status says the file ends before its
end — and a reading that dies part-way is answered as a failure rather than as
an empty result.

Two shapes of a gVCF row printed the strongest label the layer has over
stretches that were never read.
A reference block whose genotype is a no-call
(./., depth 0 — what a caller writes over an unread region) answered
«confirmed reference»; and a row whose only alternative is a spanning deletion
(*), left behind when a multi-allelic row is split and one half filtered,
answered «confirmed reference» at a base that lies inside a deletion on one
chromosome. The first is now a no-call; the second a refusal naming the allele
found.

A deletion or multi-base change standing on a catalogue base was reported as
«the file is in another build»
— a claim about the file made from a row that
merely carried a different variant; it is now «another variant at this
position». A multi-allelic row written with a shared trailing base rendered a
heterozygote as a four-letter string labelled «called»; it now prints the
locus's own allele pair, and a genotype naming an allele of another length
beside ours is refused by name instead of concatenated.

An owner of a genotyping chip saw a status that said the array was connected
and, six lines later, «no genome is connected» on every question
— closing
the locus catalogue and the pharmacogenetics a chip answers as designed. The
frame now belongs to the input that answered: the catalogue and
pharmacogenetics are open, a gene region is closed because a chip reads chosen
positions and a gene is a stretch nobody chose, and the screens are closed as
too narrow; a gene query on a chip refuses with the same reason instead of
proceeding without a file.

The secondary-findings screen could report a variant a person does not carry,
and miss one they do.
A row whose genotype was not called at all — the
ordinary shape of a family file or a gVCF — was written as a finding, because
«not a reference» was the only test. A row with two alternate alleles was
matched on the pathogenic one and judged from the other, so a person homozygous
for a benign second allele was reported homozygous for the pathogenic first. And
in a file holding several samples the first column was read as the person's: a
mother's BRCA1 heterozygote, screened as her child's. Each is now a named step —
the column is chosen, the allele index is required in the genotype, a no-call is
counted and said with the result — and a call the caller itself flagged as not
PASS is written as filtered rather than decided.

The refusal against matching two builds could be switched off by the one
variable meant to arm it.
Declaring the build of a headerless personal file
also declared it for the ClinVar file, so a declared GRCh37 against the
downloaded GRCh38 ClinVar read as «GRCh37 against GRCh37» and ran. The ClinVar
build is now read from the ClinVar file alone. And a personal file whose build
could not be told at all went into the scan with a clean «ok»; it now refuses
and names what closes it, the same way the ClinVar side already did.

The table was written beside the variant file while the screen read it from
the genome folder
, so with the file anywhere else the command said «written»
and the screen said «not run». It now lands where the screen looks, and carries
a sidecar saying which builds it was matched in, which ClinVar release, when,
and how many positions were unread — so a stale or crossed table can be told
from a current one.

The advice about external tools described a product that needed them. The
first command a new person types offered four binaries and explained that
without them «the genome layer does not work at all». True when written, false
since a file without an index became readable and the ACMG screen became
runnable. It now says what the four buy — seeking a region or a whole gene, and
the wide ClinVar annotation — and what answers without them.

The gene frame said the wrong thing about a scan that had not run. Asking
about a gene prints, above its findings, which shelves of the build hold
anything about it. When the ClinVar scan had simply never been run, that frame
said the gene's coordinates could not be obtained — sending a person to fetch an
annotation file they did not need — and clinvar --gene blamed a missing or
broken index. Worse, the ACMG line printed «in it, 0 findings» for a panel
nobody had scanned: a false «clean» on a hereditary-cancer gene. Each layer now
says which of three things is true — what it holds, that the scan has not been
run, or that it could not be asked and why — and a warning that only part of the
findings table was read now reaches the frame instead of dying on the way.

The web page still showed a genotype for a position with no row. A position
a...

Read more

v0.4.10

Choose a tag to compare

@l1nkberry l1nkberry released this 09 Sep 12:37

What you can do now

The Genome tab is one page at a time, and it opens with findings. Seven
headings on one scroll — all open, the maintenance block above every conclusion,
and the same gene appearing in four of them without any of them saying so — are
now panes of one tab. The first is neither ClinVar nor polygenic nor longevity: it
is one list, ordered by how much a line can change and marked with the source each
came from. Clicking a gene anywhere on the tab gathers what every source holds
about it into a single card: the loci and their genotypes with coverage, what
ClinVar says at those positions, what the longevity layer and the lipid card say
about that gene. Nothing was removed — the sections are still there, one at a
time, and the service text each of them opened with is folded rather than first.

The documents inside the package open as pages, and the second opinion offers
the one written for a clinician.
The product's own output names files — «see
PREPARING-THE-GENOME.md» — and after pip install the only way to read one was
scholion doc <name> at a terminal, which the person reading the interface may
not have open. The local server now serves the same nine documents at
/doc/<name>, each as a self-contained page that prints on paper and pulls
nothing from the network. «Second opinion» carries a button to the one-page
description for clinicians and researchers: what the program is, what it
computes, and where it refuses to answer — meant to be handed over at the
appointment.

The first screen is the person. It opened with the one experiment being run
this month and reached the reader before anything about the reader. The order is
now the order the questions come in: who the data belong to — sex, age, height,
body-mass index, the reference panel — then their own indicators, then the
targets they are aiming at, then the body systems and the figure, then what is
out of range right now, then what is worth measuring, and last the focus of
attention. Every block still names the tab that owns it and goes there on a
click.

The Profile tab is retired into that first block. It held four cards and two
forms, and the four cards were the ones a watch measures every day; as a tab of
its own it had become the emptiest page in the product. Nothing it could do was
lost — the profile form, the manual measurement form and the list of what the
profile is still missing all open at the top of the first screen.

The state of the genome is said in one place, including when it is not being
read.
The tab used to show a grey badge and leave the reason unprinted; a
person whose folder held more than one candidate file met «not read» in six
places and the cause in none. The state now carries the file, the assembly, the
sample, what was set aside and why — and, where a choice is open, the names with
a button beside each. The choice is kept, so it is asked once
(scholion choose-genome at the command line).

«What to test» is part of the second opinion. Every row of it is a line for
the same conversation, and it was a tab that handed half of its own list back to
the tab beside it. The routine controls travel with it, folded.

The lifestyle brief's «needs review» can be answered. The flag is raised when
a marker a block watches is measured after the block's wording was last read, and
until now nothing in the product could lower it: it went up once and stayed up, at
the top of the tab, above the content. There is now one button — the wording still
holds — which records the date that has always been what the flag compares against
(scholion brief-reviewed <block>). It also moves down the page to sit beside the
wording it is about.

What is fixed

Accepting the reach baseline moved the whole file to whichever machine ran
it.
A number in test_reach_baseline.json is not a property of the code
alone: it is what the suite reached on one interpreter, with one backend — the
two do not count a line identically — and, for at least one module, only where a
file the repository does not carry happens to sit. --strict had printed a
warning whenever the run and the recorded stamp disagreed; --accept rewrote
every number regardless, restamped the file, and reported it in one line.

There are three writing modes now, and only one of them moves the file.
--accept records the whole measurement and refuses when the baseline was taken
elsewhere — before measuring, since the answer never depended on the ninety
seconds. --accept-new records only modules that have no accepted number yet,
which is what the suite's own guard asks for when a module is added; it changes
no other number, not the overall, and not the stamp, and it does not run the
suite at all when there is nothing to add. --rebaseline is the deliberate
whole-file move and prints every number it lowers, old → new, before writing.

A document name was tolerant about spelling, which was free until a route
handed it one.
DATA_LAYOUT, data-layout.md and data-layout are one
request, and refusing two of them teaches nothing but the exact spelling. While
the only caller was a person typing at their own prompt, that was the whole
story; a URL is composed by whoever holds it, and ../../../etc/passwd builds a
path outside the package as readily as a name builds one inside it. The
tolerance stays and the shape does not: a document name is one file name, or it
is nothing. Anything else opens the list of documents instead.

Seven of the thirteen rows of the goal table said «—» while the numbers sat in
the file.
Weight, body-mass index, body fat, muscle mass, VO₂max, resting heart
rate and steps come from a wearable device, and five of the goal charts drew
nothing at all. The lifestyle layer stores a measurement together with the device
that made it — two watches do not measure resting heart rate the same way, and
one series built out of both shows a step on the month the second export was
loaded — and the goal reader had been written before that was true. It looked for
the metrics where they used to sit, found nothing, and returned an empty series.
Nothing failed and nothing was logged: «—» is what that table prints when there
is no data, and there was a decade of it.

Two things follow, and both are new behaviour rather than a repair. The reader
goes through the accessor that knows the file's shape, so a file written by an
older version answers exactly as a current one does. And a row with no number now
says which of three things is the matter — nothing carries this series, the
series is empty, or more than one device measures it and the goal has to say
whose. The last of those is refused rather than averaged, and a goal may name the
device (wear:garmin:RestingHeartRate) to answer it.

A number the watch already had was reported as missing, or as three weeks
old.
Some indicators are kept twice: what a person types in, and what a device
records every day. Only the first was read. Steps stood at a single figure
entered in July and were called «below target», while the device series had the
month just gone above it; sleep showed nothing at all beside seventy-five months
of nightly data; a resting heart rate from December stood as the current one in
September. The two are joined now: the newest measurement is the one shown, each
card says which store it came from and on what date, and the other store is
printed beside it rather than instead of it. A tie goes to the hand-entered
reading — a monthly mean and a measurement taken on a day are not the same
statement. Neither file is written to.

The pairing is declared once, in the shipped wearable reference, and only where
the two are the same quantity. Where they are merely similar it is left out and
the row goes on saying it has nothing behind it: intensity minutes are not
«minutes of activity», and a pairing that is nearly true prints a number nobody
can act on.

The goal board was dated by one of the files behind it. The heading read
«data as of» the timestamp of the wearables file, while half the rows come from
the laboratory — so a table carrying a draw from the 3rd was headed with the 23rd
of the month before. Every row carries its own date now, and the heading carries
the newest of them.

A stored result was deciding what language the product speaks. Polygenic
risks and longevity markers printed in Russian while the interface was English.
Neither catalogue is missing a translation — both carry every name in both
languages. The names were coming from prs_results.json and
longevity_findings.json, which are stored RESULTS: each label is a copy of the
catalogue made on the day of the run, in whatever language that run was speaking.
The catalogue now decides what a thing is called and the file decides what the
number is; a trait or a marker the catalogue does not carry keeps the name it was
stored with, because a percentile with no name is worse than a name in one
language.

The longevity layer showed genotypes and explained none of them. Each row
printed a gene, an rsID, a genotype — and an explanation line that was always
empty, because the page asked for a field these rows do not carry. Everything
that says what a marker MEANS was in the catalogue, in both languages, unread: what
the allele is, what a second copy does, what it argues for, what population the
direction holds in, and the papers behind it. Rows now carry all of it, sorted so
that what was found comes first and what was checked-and-quiet folds away. The
verdict and the strength of the sources are written out as sentences rather than
as the internal words they are — and a word the catalogue does not recognise is
never turned into a message key, which is how «⟦longevity.verdict.…⟧» used to
reach a reader.

APOE says what APOE is. The card led with «APOE — status» over two rsID
numbers, which tells a reader nothing about the gene or about their own
combination. It now opens with what the gene i...

Read more

v0.4.9

Choose a tag to compare

@l1nkberry l1nkberry released this 08 Sep 19:28

What you can do now

The Overview draws the body beside the radar. The same systems, twice:
the radar as before, and a figure on which every system with a place on a body is
lit by its score. Around each mark a ring shows how much of that system was
actually measured, which is the number the radar has never been able to show. A liver scoring 100 out of 100 on one marker of three is a full
point at the edge of the radar and a ring closed a third of the way on the body.
Picking a system on either picture marks it on the other.

The endocrine system is four systems, and each is scored on the panel a
laboratory actually issues.
One domain called «Hormones», averaged from
testosterone, IGF-1 and TSH, is gone. In its place the radar carries the thyroid
axis (TSH, free T4, free T3, anti-TPO), the adrenals (cortisol, DHEA-S), the
gonads (testosterone, DHT, estradiol) and the growth axis (IGF-1, growth
hormone). The ring beside each now says how much of that panel exists, instead of
how much of a mixture of three markers nobody orders together; a system with
nothing measured is not drawn at all rather than averaged into the picture.

A hormone is marked at the gland that makes it, not at the organ it is read
for: TSH and growth hormone at the pituitary that secretes them, free T4, free T3
and anti-TPO at the thyroid, cortisol and DHEA-S at the adrenals, IGF-1 at the
liver that writes it, testosterone at the gonads. Each mark carries the score of
what is made there — the average of a whole axis would print the same number in
two places and mean it in neither. Two hormones answer instead that there is no
one place to mark: DHT is converted in the tissues that respond to it, and
estradiol comes from the ovary in one person and from adipose tissue in another.
An antibody is the one named exception to the rule — anti-TPO is made by
lymphocytes, which are everywhere and mean nothing as a place, so it is marked at
the gland it is raised against. A reason is recorded beside every placement, and
beside every refusal to place.

The pancreas is on the map. It is a system of its own now — amylase, its
pancreatic fraction, lipase and C-peptide, the acinar cell and the beta cell
together — and insulin, which stays in the carbohydrate panel because HOMA-IR is
computed from it, is marked at the gland that secretes it. Nothing else about
carbohydrate metabolism moves: glucose, HbA1c and the index itself say plainly
that they are made nowhere in particular.

A reference interval that belongs to one sex is no longer lent to the other.
Six markers carry both corridors and were already handled. Every other marker
carries one, transcribed from the forms of one person — and it was borrowed by
anybody whose own laboratory printed no range. HDL's floor of 1.0 mmol/L is a
man's; so are AST's ceiling of 40 U/L, GGT's of 60, and the intervals for DHEA-S,
DHT and estradiol. Borrowed for a woman, none of them failed: each produced a
verdict, and the wrong one. Every marker a body system is scored on now declares
whether its corridor depends on sex — silence was what made the class invisible,
so silence now fails the build — and a corridor is withheld, with the reason
printed, rather than lent across that boundary. A person's own form is unaffected:
it is always preferred, and this only ever governed what fills a hole.

A number with no reference interval beside it no longer costs its system
points.
Such a marker scored 55 out of 100 — the score for «a deviation whose
size cannot be assessed» — when there is no deviation, because there is nothing to
deviate from; and it was counted among the system's deviations, while the marker
list on the same screen said, correctly, that it was not one. It is now left out
of the score and out of that count, and still counted as measured: a system whose
only measured markers carry no corridor reports no score rather than a mediocre
one.

The five systems with no place on a body say so rather than being drawn on an
organ they do not belong to; which is which is recorded with a reason beside each
entry, and a system added to the radar without one now fails the build. The figure
is drawn male or female from the profile; where the profile does not say, it is
not drawn at all and the radar answers alone.

Windows is a supported platform. The package, the command line, the local
web application and the assistant skill run there, and a cell of the test matrix
now says so on every commit rather than leaving it to hope. The promise and the
check are tied together: a platform named in the package metadata with no runner
behind it, or a runner with no promise in front of it, now fails the suite.

Two things stay Unix-only and both are optional. The crossread wrapper is a
shell script — on Windows the installed scholion command is the same core, so
nothing is lost but the second name. Building a genome from raw reads drives
external alignment and coverage tools through shell scripts; every such tool is
looked for before it is used, so where one is missing the answer names it
instead of failing. The data directory on Windows is %USERPROFILE%\.scholion,
and SCHOLION_REPO_DIR overrides it as everywhere else.

The suite can be started without a shell. Until now the only way to run it
was a bash script, including the run the release procedure requires INSIDE the
unpacked package. There is now a second entry point written in Python that does
the same run, so a person who received the package can check it on a machine
that has no shell at all. The build carries it, and a check makes sure of that:
every way of starting the suite that the project's own automation uses must be
present in the package a recipient receives, or the build fails.

Any gene can be asked about, not only one in the curated catalogue. scholion genome --gene CASR used to answer "Gene CASR is not in the coordinate reference"
— true about loci.json, which is a book of pharmacogenetic loci, and easily read
as a statement about the genome, whose reads were there the whole time. The gene
name is now resolved to coordinates from a local Ensembl annotation (or from
Ensembl live, cached afterwards), the interval is cut out of the personal VCF,
coding variants are separated from the rest by the coding exons of the canonical
transcript, and what each one does to the protein is computed against the same
reference the genome was called against — no web service is asked, and no
coordinate leaves the machine.

A gene report prints coverage before it prints findings. Every reassuring
thing such a report can say has the form "no such variant", and that sentence is
empty until the region is known to have been read. Read depth is computed
directly from the alignment, without samtools or pysam, and it reproduces the
project's own native callability run exactly — mean and every threshold, to the
last base, on four genes across three chromosomes. Where a piece is missing the
report names it and what would supply it: coverage without an alignment prints as
"not measured — because", and "changing the protein: 0" is replaced by "not
computed" wherever the reference was absent, because a zero and an uncomputed
number look identical on the page and mean opposite things. The two blind spots
of short reads — large exon-level deletions, and deep intronic or regulatory
variants — are printed on every answer, not only on the reassuring ones.

Three analytes the dictionary did not know are now recognised. A urine
albumin-to-creatinine ratio, a urine microalbumin and a total T3 went through
ingest-labs without a word: an unrecognised line is skipped, and a skipped line is
indistinguishable from a line that was never on the form. All three sit on real forms
of an ordinary laboratory, and the ratio is one of the two axes of KDIGO staging — a
value whose absence changes what may be said about the kidneys. Total T3 gets its own
table rather than joining free T3: the two are reported in different units (ng/dL
against pg/mL) and the same number under the wrong one is out by an order of
magnitude. The unit table learned mg/g, so the ratio prints with a unit rather than
a code.

A polygenic percentile is printed with two quantities beside it: how stable
the number is, and how informative the model is.
A bare percentile looked
equally convincing for coronary artery disease and for intelligence, and it is
not: for the second, the choice of reference population moves the number by
some fifty points while the model's whole range separates the outcome by a
few. The two are different properties and are shown as such, under every
percentile in the prs report and in its JSON. Stability belongs to the
measurement — the spread of the percentile across the five reference
populations, its spread across the models scored for the trait, and the share
of the model your file actually covers. Informativeness belongs to the model —
its discrimination (AUROC), and what the score's range from the 10th to the
90th percentile does to the outcome, read per standard deviation as the
catalogue reports it. A figure that is not on the machine is named as «not
recorded» rather than left out, because a line with one number and a line
with three look alike only when the missing two are silent. The two are not
folded into one «signal to method» ratio: the effect size cancels out of such
a ratio, and what remains is the population spread written in other letters.

What is fixed

A corridor transcribed at one age was lent at every age, and most corridors
did not say whose they were.
The rule that keeps one sex's interval from the
other covered only the markers the body systems score; the other three hundred
and forty held a corridor and were silent about it, so a man's prolactin or
LH interval could still be borrowed for a woman whose form printed no range.
And age was not modelled at all: IGF-1 and DHEA-S depend on it more than on
sex, every laborat...

Read more

v0.4.8

Choose a tag to compare

@l1nkberry l1nkberry released this 31 Aug 09:29

What you can do now

ingest-studies names every file it took nothing from. The report used to be
four counts, and the number of files seen was never reconciled against them: a
file that was read and then dropped moved no counter at all. Each such file is
now listed with the reason it produced nothing — the PDF gave up no text, it is a
laboratory form that the other loader owns, it reads like a study but no
conclusion could be lifted out of it, or this loader cannot place it at all. The
counts and the named files add up to the files that were looked at.

A PDF holding several studies is reported with what is inside it. A discharge
summary — several examinations, and often a laboratory panel, in one file — is
listed section by section with each section's own date, together with the plain
statement that this loader cannot split such a file yet and that nothing from it
reached the profile. Until now such a file was declared a laboratory form and
handed to the laboratory loader, which dropped it for a reason of its own, and
neither report held a line about it.

A decision you made about a wearable series survives the next import. A
rebuild is written from the export every time, so an edit made inside
wearable_trends.json was lost at the next run — a weight point removed by hand
as physically impossible came back as soon as a fresh export was read. Put the
decision in profile/wearable_corrections.local.json instead: remove or
replace a month of one metric of one device, with a reason. The import applies
them over the rebuilt series and says which it applied, which it refused and why,
and which no longer match anything — a reason is required, and a correction
without one is refused rather than applied. LOADING-DATA.md describes the file.

A monthly point from a wearable now says what it stands on, and a movement too
small to see is no longer given a direction.
Each month carries the number of
days that held a reading, the days the month had, the median beside the mean and
the spread. From those the lifestyle answer works out the smallest difference the
data can tell from its own sampling, and when a shift is smaller than that it says
so instead of naming a direction — for a monthly sleep series measured on a
typical month that threshold is on the order of several minutes, and movements
under it were previously reported as improvement or decline. A month measured on
part of itself — nineteen nights of thirty-one — says that too. A series written
by an earlier version carries no sample description and behaves exactly as before:
nothing is invented for it.

A laboratory arrow says whether the change can be told from the marker's own
scatter.
«↑ 44 % since the previous measurement» is arithmetic on two numbers
and it stays; what follows it now is whether this marker's own history has ever
shown itself able to tell a change that size from its own wobble. The threshold
is measured from that history — no outside table is required and no coefficient
is assumed — and it is reported with it, so it can be argued with. A rule that
suggests a test because a marker is trending no longer fires on a movement the
series cannot distinguish: a suggested test costs a real draw of blood. Where
the history is too short to support the statement, none is made.

A genotype read from a VCF says how well the reads actually supported it.
Depth said how many reads there were; nothing said how they were divided. A
heterozygote whose reads split 15/85 — a mosaic, a duplicated region, an
artefact — was presented in the same words as one that split evenly. The allele
fraction and the caller's own quality score are now read, and everything worth
confirming about a call arrives as one list with the measurement that raised each
entry, instead of three separate flags a reader had to know to look for. These
are quality-control heuristics and they say «have this confirmed by another
method», never «this is wrong».

A gene a variant file cannot answer now says so, instead of showing a tag
SNP.
CYP2D6 decides codeine, tramadol, tamoxifen and the tricyclics, and it is
settled by the full diplotype — copy number and phase — not by any single
position. Asked about codeine, the answer used to print the gene, the word
«important», a bare genotype and the promise that a full genome would bring the
rest, which for this gene is not true. It now states that a diplotype is what
decides it, shows the tag SNPs under a name that says what they are, and names
the two things that actually close the gap: a star-allele call over your reads,
or a laboratory report that states the diplotype — the only route open to
somebody who has no alignment file. A diplotype already on file is used as
before, and the sentence beside it now follows where it came from instead of
asserting the caller's name over both cases.

The genes a «nothing found» cannot rest on can be handed to a laboratory.
The coverage table knew which genes were under-read and could give the list to
nobody: a percentage names a gene, and re-reading something needs coordinates.
The intervals the measurement was made over are now recorded and the weak list
exports as a BED, worst gene first, each interval carrying its own percentage.
The track line states that these are gene loci with a margin rather than coding
sequence — a small dropout inside an exon barely moves a locus-wide percentage,
so the file is a worklist of genes to look at again, not a map of the bases that
were missed. A coverage table written before the coordinates existed is refused
with the run that would fill them in: a gene name is not an interval and none is
invented from it.

What is fixed

A conclusion written in mixed case produced an empty record, and an empty
record was dropped without a word.
The recogniser accepts the heading whatever
its case; the extractor beneath it demanded the word in capital letters and
required the text to begin on the following line. A document that writes the
heading in ordinary case, with the text after a colon on the same line, was
therefore accepted as a study and then yielded nothing at all. Both shapes are
read now, and the printed disclaimer that contains the same word is not mistaken
for the finding.

A laboratory point is dated in one of three shapes, and anything else is
refused.
A month, a day, or a day with the clock time of the draw — the last
one because two draws can fall on one day and a key no finer than the day loses
the second. Nothing checked the string before: a value that was not a date at all
would have become a key of the series and been sorted and charted beside real
dates. And when the same period is already present at another resolution — a
month point and a dated point of that month, one measurement standing twice — the
import now says so rather than quietly holding both.

A doubled measurement was found and then not mentioned, for anyone whose
results arrive as a table.
The importer notices when one measurement ends up in
a series twice at two resolutions — a month point and a dated point of that
month. It said so for results read out of a PDF and stayed silent for results
read out of a CSV or TSV export, because the two kinds of file reach the profile
by different routes and only one of them carried the report. The doubling was
detected in both cases and shown in one. Both now say it.

What needs recomputing

Run scholion ingest-studies --force over your folder of documents once.
Studies that were dropped in silence may now be read, and whatever is still not
taken will be named in the report.

Measure coverage once more if you want the BED of under-read genes: the
measurement now records the intervals its percentages were taken over, and a
table produced before it did does not carry them.

Run scholion ingest-garmin once. Existing months gain the description of
their sample; the values themselves do not change, so charts and comparisons are
unaffected. Until an import has run, a series simply has no sample description
and no statement about what is distinguishable is made for it.

v0.4.7

Choose a tag to compare

@l1nkberry l1nkberry released this 27 Aug 07:22

What you can do now

The Ouroboros Hub skill now says where your files go, on a page of its own.
Installed there, Scholion arrives by a click into a machine whose paths you have
never seen and cannot list. Until now it registered thirty tools and nothing
else: no data directory was created, and the first thing the skill said was a
tidy report — markers: 0, genome: not connected — about a person it had
never been given anything about. A statement of absence, phrased as a finding.

Enabling the skill now adds a Scholion tab to the Widgets page, and that tab
is the first thing to read:

  • it names the data directory this installation actually uses, and the exact
    folders for laboratory forms and for the genome;
  • a button lays that directory out — empty templates, plus a README in every
    folder saying what belongs in it. Pressing it a second time writes nothing,
    and it never touches a file that already exists;
  • if your files already live in a folder of their own, a field points at that
    folder instead of copying them.

The directory the skill creates is one the host keeps, so it survives restarts
and is not scattered inside a container. To put it somewhere else, set
SCHOLION_REPO_DIR (the whole layout) or SCHOLION_PROFILE_DIR (the profile
alone) in the host environment before enabling — either is respected, and
nothing is moved.

Nothing changed for the command line or the local web application: the same
layout is what scholion init has always written, and the tab creates it
through the same command.

What is fixed

A freshly created, empty directory was described as a connected genome. The
introduction the skill prints was reading the genome section of the report as
«present» rather than reading its ready flag, so a person who had just pressed
the button — with no VCF anywhere — was told the genome was connected and then
handed a list of loci «still unread». Both halves were wrong and the second made
the first look substantiated. It now says not connected until a VCF is
actually there, which is what scholion genome-status said all along.

A profile that cannot be read is no longer reported as an empty one. If a
file in the profile is malformed, the tab now says so, with the reason, instead
of showing zeroes that look like a person with no history.

What is retracted

Nothing. No stored value changes, and no conclusion drawn from a previous
version needs revisiting: the corrections above are to what the skill said about
its own state, never to a reading of anybody's data.

What needs recomputing

Nothing.

v0.4.6

Choose a tag to compare

@l1nkberry l1nkberry released this 24 Aug 21:16

What you can do now

A tool that installs skills from a repository can now find this one.
npx skills add and its kind read a repository rather than a package: the root
if it holds an entry, then skills/, then the agent folders — the documented
shape being skills/<name>/SKILL.md. None of those existed here, so the entry
was reachable only by such a tool's fallback recursive search, at a path whose
folder is called skill while the file inside calls itself scholion — and the
format requires those two to agree.

The published repository now carries skills/scholion/, and it holds the entry
and nothing else. That is the licensing decision rather than economy: what is
published as a skill is the entry; the long instruction and the canon of rules
stay with the package they are printed from, under their own licence, and a
folder carrying both would put one licence over two bodies of text.

Unlike the standalone skill folder in the downloadable archive, this one is not
excluded from the repository — a tool fetches it from there, and a folder kept
out of version control does not exist for the purpose it was made for.

An assistant that reads skills from a folder can now be given this one, in one
line.
Several hosts read the same path, with no registry, no account and
nobody's moderation in between — and until now nothing in this project mentioned
it:

mkdir -p ~/.agents/skills/scholion
cp "$(scholion skill --path)" ~/.agents/skills/scholion/SKILL.md

That single file is the whole installation, which is what makes it worth doing
and also what it had to be repaired for. The entry did not name the tool
server.
A host with a plugin mechanism finds scholion mcp through its own
means; a host that reads only this file had no way to learn the door existed —
so the door built for exactly those runtimes was invisible to them. It names the
tool server and the in-process module now, and says that neither writes to a
profile.

And it pointed at reference texts that a one-file install does not have. The
full instruction and the canon of safety rules sit beside the entry in the
downloadable bundle and nowhere else; the entry named them by path regardless, so
a model following the pointer would report a missing file rather than ask for
what it needed. Each is now named twice — the path, for the bundle, and the
command that prints it out of the installed package, for every other way in.

scholion capabilities --json lists this door beside the others, derived from
the build: the size of the entry and whether it names the tool server are read
off the file rather than asserted about it. scholion doc connecting-an-agent
describes it in words, and the heading there has stopped counting the doors —
it said «four» through the arrival of a fifth and a sixth.

The checks now run on the oldest Python the package promises, before a release
rather than after one.
pyproject.toml says this runs on Python 3.10 and
later. That sentence is a promise to everybody who installs it, and until now
nothing checked it before a release: the matrix that runs every promised version
lives with the published repository, so it answers about a version that is
already out.

It answered twice in two days, both times too late — once with a release build
that failed and never reached the registry, once with a red matrix on a version
that had. Both were the same shape: a repair verified on one interpreter and
promised about four.

./run_tests.sh now runs the whole suite a second time under the oldest promised
version, taking that version from the promise itself rather than from a number
written into the script — scholion doc contributing describes how to run the
checks that come with the package. Where uv is available it fetches the interpreter and
caches it; where it is not, the run says the check could not be made instead of
passing over it in silence. SCHOLION_SKIP_OLDEST=1 skips it while editing.

Three places name that floor — the promise, the versions the matrix runs, and
this step — and a test now compares them, so a version can no longer be promised
and never run, or run and never promised.

v0.4.5

Choose a tag to compare

@l1nkberry l1nkberry released this 24 Aug 07:30

What was wrong

A check that travels with the package disagreed with itself between Python
versions.
The tool that measures how much of the code the test suite actually
runs asked the compiler whether a file has anything to execute. For a file with
no statements in it at all — the vendored package carries an empty __init__.py
— what compiles is an implicit return, and the line it is numbered at is 1 on
Python 3.10 and 0 on 3.11 and later. The tool ignores line 0.

So one empty file was measured on one interpreter and skipped on another, and
the recorded baseline, taken on a newer Python, failed the suite on 3.10 over a
module with no code in it. Anybody running the checks on the oldest Python this
package supports would have met it — the checks travel beside the package and
scholion doc contributing describes how to run them. The question is answered
from the source now: no statements, nothing to reach, the same answer
everywhere.

v0.4.4

Choose a tag to compare

@l1nkberry l1nkberry released this 23 Aug 21:45

This release is mostly about a medical record read from a file: what such a record
holds that is not a laboratory number, and how the product names the codes it could
not place. The rest is about statements being true — a number and the word after it,
the version the local server gives out about itself, what a connectivity check is
allowed to fetch, whether a download checked who answered it, and whether an index
still describes the genome it was built from.

What you can do now

*The facts the application cannot derive can be given from the page at last.
Sex, year of birth, height and which wearable answers are preconditions: without
them a dozen reference intervals are withheld rather than guessed, the age-banded
rows of a laboratory form cannot be read, and there is no body-mass index. There
is now a Profile tab, and it offers all four. The main wearable had no field
anywhere but the command line.

The reference panel for percentiles is no longer a question. It used to be
one: --ancestry EUR|AFR|EAS|SAS|AMR, a superpopulation code nobody knows about
themselves in those terms. What such a box collects is a guess, and a guess
stored where a measurement goes is indistinguishable from one afterwards — while
every polygenic percentile depends on it.

Your own genome answers it. Comparing a few hundred of your genotypes against
the five 1000 Genomes panels is a step of preparing a genome, and the result is
used from where that step writes it; the Profile tab shows which panel applies
and says whether it was measured from your DNA or set by hand. Until it has been
determined, percentiles go on naming the default panel they used. The
command-line flag remains as a deliberate override.

Saying there is no wearable is an answer of its own. Until now «I do not own one»
and «nobody asked» were one blank field, so anything listing what the profile
still needs asked a person with no watch about their watch for ever. Choose «No
wearable device» once, or scholion profile --wearable none, and it stops.

What the profile is still missing is now part of what the data cannot answer.
scholion limits already answered one question — what cannot be said here, and
what would close it — and a missing precondition is exactly that. Each one now
appears in that list with the sentence it prevents and the command that records
it, and disappears from it once answered. An assistant reading the list at the
start of a conversation therefore asks for what is actually absent, rather than
from a list somebody typed into an instruction and then had to keep in step.

A body measurement in a bundle is now kept, not dropped.* This product has
always taken a weight — from the command line and from the page — and an import
that met one inside a medical record threw it away because it is not a
laboratory analyte. It goes to the metrics layer where it belongs, with the same
mark of whose measurement it is as everything else written today. Only where one
of this product's own metrics holds the same quantity IN THE SAME UNIT: nothing
here converts, so a weight in pounds is refused rather than joined to a series in
kilograms. scholion import-fhir <file> reports the two layers separately,
because one number for both would say nothing about either.

Height can be set from the command line at last. The page has always had the
field; the command had not, and the body-mass index needs it —
scholion profile --height-cm 178. A bundle that states a height still does not
apply it: a file may hold a relative or two people, and a height is one field of
the profile rather than a series, so it is reported for you to set yourself.

A marker may now carry more than one LOINC code, because a LOINC code is not
one per analyte: the same substance has a different code by material and by
method, and four of that bundle's unplaced observations were analytes this
dictionary knows under a code it does not — glucose in whole blood beside
glucose in plasma, LDL measured directly beside LDL calculated. Every additional
code has to state why the two are the same measurement, and none ships: whole
blood and plasma differ by about a tenth, so pairing them because the names
match would put a systematic error into a series. The mechanism is here; the
pairings wait for a source, and the refusal now prints the code, so each one can
be looked up.

The suite now says how much of the code it runs, and will not quietly run
less.
A thousand green tests is not a measurement, and this project had been
reading it as one. The number, once taken, was 69.9% — with the module that
implements provenance for every answer at 12.8%, and the VCF reader used whenever
pysam is absent, which is most installations, at 35.4%. Nothing could have said
so, because nothing was counting.

check_test_reach.py, which travels with the package beside the other checking
tools, counts — on the standard library alone and
including the tests that run the command line in a real subprocess. It reports
per module, --strict fails when a module falls below the reach recorded for it,
and --accept records a new number deliberately. It is the last step of
./run_tests.sh; SCHOLION_SKIP_REACH=1 skips it while editing.

No module is now under half. Seven that were — polygenic scoring, the provenance
audit, the study loader, the watch import, every route of the local web
interface, the online drug lookup and the VCF reader — carry tests for what they
actually decide: which polygenic model is allowed to speak for a trait, that a
score is not computed for an organ the person does not have, that a stored number
which no report holds is told apart from one that a second method explains, that
a rebuilt watch export cannot erase months it no longer mentions, that a
judgement written about a study survives the file being read again, and that an
empty answer from a database that was never reached is never printed as «nothing
was found».

Nothing about anyone's data changes, and no command behaves differently.

What was wrong

*A form the interface never showed. The page holding sex, year of birth and
height existed, worked, and was in no tab: it had never been reachable, in the
whole history of the file, so in practice those facts could only be set from the
command line — while the command's own help said the opposite. Every view the
page defines is now checked for a way in, because a view nobody can open fails no
test that calls its functions directly.

A recorded age and a recorded sex could both show as «—». The age was
computed from a year of birth alone, so a profile carrying a full birth date —
which is what the demonstration writes, and what an imported medical record
writes — reported no age at all while the file held one. And the page compared
the stored sex against male, while the file is allowed to say m; a sex
already given then showed as blank and was asked for again. Both are read
through the one function that knows the spellings.

A value that cannot mean anything was stored rather than refused. --ancestry EURO went in, and every polygenic percentile afterwards was computed against a
reference population that does not exist — printed as an ordinary number, and
without the caveat about a default one, because a value was set. Nothing
downstream could tell. Populations, devices and spellings of a sex are now
checked on the way in, by the same lists in every face; an unrecognised value is
refused with what is accepted. A profile field a person did not touch is also no
longer rewritten by a save of a different one.

«1 markers are printed without a reference range».* A number and the word
after it did not agree — on the screen that says what the data cannot support,
which is the screen this project argues from. The machinery for it existed and
was in use; twenty-three messages simply did not go through it, and in Russian,
where a noun after a number takes three different forms, the same lines read as
carelessness about everything else on them. All twenty-three are repaired, and
the rule that replaces them is mechanical: a message may not put a number
immediately in front of a word unless it carries all the forms of that word. The
page and the package now also agree on which form to choose, checked against
each other rather than each against its author.

A body weight was reported as a laboratory code nobody knows. Importing a
FHIR bundle, ten of the twenty-three observations it could not place were not
analytes at all — height, weight, body mass index, temperature, heart rate,
respiratory rate, oxygen saturation. Calling them «a code this build does not
know» sends the reader looking for a dictionary entry that should never exist.
They are named for what they are now, each with the code and, where one of this
product's own metrics holds the same quantity, with that named too — and where
it does not, the reason is written down rather than left as silence: a heart
rate at a visit is not the resting home pulse, and a body-mass-index percentile
against an age-and-sex reference is not the index this product keeps. Nothing is
written from them yet; what changed is that the list of what a bundle held is no
longer misleading. On the same bundle the count of «not in the dictionary» falls
from 23 to 13.

An observation the dictionary could not place named only its label. Importing
a FHIR bundle lists what it did not take and why; for a code the dictionary does
not know it printed «Calcium» — while the code, which is the thing an entry is
keyed by and the only way to look the analyte up, sat in the record one line
short of the screen. It is printed now: «Calcium (49765-1)».
scholion import-fhir <file> --dry-run shows the list without writing anything.

A profile with no birth year stopped a whole batch of laboratory forms.
Reference ranges on a form are often given by age band — one row for 40 to 49,
another for 50 and over — and choosing the ro...

Read more

v0.4.3

Choose a tag to compare

@l1nkberry l1nkberry released this 22 Aug 10:58

This release is about files that arrive wrapped, laboratory forms printed for an
American reader, and — the larger half — a report that stops claiming more than
the file it was given can support.

What you can bring it now

A consumer DNA test still inside whatever the provider wrapped it in. A VCF
compressed with bzip2, a VCF inside the provider's zip, a VCF whose file name
carries URL-encoded brackets — all of them ordinary VCFs, and all of them now
read. A file is identified by its bytes before its name, so an export that
arrived zipped is no longer «an archive, not opened blind here»: looking inside
is the opposite of blind. None of these can be seeked into, but the catalogue is
fifty-four loci, so one cached pass over the file answers all of them — and the
same pass measures the call set exactly rather than by probe.
scholion genome-status names what it found.

A genotype table from a chip, with the ceiling of the chip attached. It is
read the same way, and a position nobody typed is reported as not read — never
as the reference.

Three formats stay named and unread, each with its reason, because a gap
left unexplained is indistinguishable from an oversight: a VCF that went through
a spreadsheet (its structure is gone rather than hidden — the honest answer is
the original export), a Complete Genomics var table (the vendor ships a
converter that is correct by construction, and a second implementation here
would be a worse one), and an archive of the FTDNA era keyed by internal SNP
numbers with no rsID anywhere in it.

Laboratory forms printed for an American reader. A LabCorp report prints its
dates as a table — the column headings on one line, the values on the next — and
a reader expecting the label and the date side by side finds neither. The
heading and the line under it are now read, and only while «collected» is the
leftmost date column; where that cannot be established, nothing is read rather
than the wrong column filed. A two-digit year is read. And a page carrying
12/15/2008 has already said which order it prints in, because there is no
fifteenth month, so 12/10/2008 beside it is the tenth of December and not a
coin toss — the evidence is on the same page, in the same table, from the same
instrument. A page whose dates contradict each other is still refused, and so is
a lone ambiguous date with nothing on the page to settle it. Two weaker
witnesses are used only after the page and are named as what they are:
«Ordered Date», which some lipid panels print instead of a draw date, and the
file name, where a month spelled out is read and a numeric one is not.
scholion ingest-labs reports what each file gave.

An exome or a panel whose header does not say which build it is in. The
build decides whether the genomic layer may answer at all, and it was
established from the lengths of the chromosomes in the header, or — failing
that — from a variant lying past the end of chromosome 1, which only a GRCh37
file can have. A capture panel has neither: no contig lengths, and no rows out
at the telomere to probe with. Three providers stamp their own pipeline into the
header and build against one reference only, so that signature is now read as
well. Where a pipeline serves both builds, the provider's name alone settles
nothing and is not used on its own: DRAGEN counts only together with the
reference path in the same header. scholion genome-status prints which of the
four witnesses answered — a build inferred from who made the file is a weaker
claim than one measured off the file, and it no longer looks the same.

A WHOOP export. The zip that arrives by email — or the folder it was
unpacked into — is now read: recovery, day strain, resting heart rate, heart
rate variability, respiration, blood oxygen, skin temperature, sleep with its
stages, sleep debt, consistency and efficiency, and workouts.
scholion ingest-wearable <folder-or-zip> takes either device and says which
one it recognised; the Garmin command still does exactly what it did.

WHOOP does not publish the layout of that export, and a layout nobody published
is one that can change without telling anybody, so the column names are read
rather than assumed. Each header is looked up in a table that ships as data, and
a column that is not in the table is listed by name with nothing read from
it — an export carrying a column this has never seen says so instead of quietly
dropping a measurement.

The reader can be named. Three readers can put a genome through: two
external programs and the one that ships here. Until now whichever was installed
won, which means the same file on two machines was read two different ways with
nothing saying so. SCHOLION_GENOME_ENGINE=tabixlite scholion genome-status
pins it — useful for anybody comparing two runs, and the only way to see what
somebody with no external tools installed actually gets. A name that is not one
of the three, or a reader that is not installed, stops the genomic layer and
says which: falling back quietly would answer a different question from the one
asked.

Two devices no longer become one line. A Garmin and a WHOOP both report
resting heart rate, heart rate variability, respiration and sleep, and they do
not measure them the same way: different window, different algorithm, different
place on the body. A measurement is now stored together with the device that
made it. Where both measured the same thing both series are shown, neither is
averaged into the other, and no conclusion is drawn from either until the
question is answered — scholion profile --wearable whoop names the one that
speaks. For anybody with a single device nothing changes at all. A lifestyle
file written by an earlier version is brought to the new layout when it is read,
and the device it is filed under is taken from what that file says about itself:
a file naming no device is filed as unspecified rather than assigned one.

Every number says which device measured it. The strip above the lifestyle
section names the devices in the profile rather than the file they are kept in,
and each metric card carries measured by … under the value. Where two devices
report the same thing and none has been named to answer, the page says so at the
top of the section and on each affected card, and offers to settle it in one
click. A column the reader of a WHOOP export does not recognise is listed by
name after every import, and can be named once in
profile/wearable_metrics.local.json so the next import reads it — additions
merge per column and cannot delete what already works.

What was wrong

Your own measurement joined a fictional person's history without a word.
The demonstration profile marks itself as invented, and the mark was on the
FILE. Adding a real value to it therefore worked: the point joined a series of
generated numbers, the file went on declaring itself synthetic — by then untrue
— and the overview counted the abnormalities of somebody half imaginary. Every
datum now records whose it is, and one profile holds one person. The first real
value written — by scholion add-lab, scholion add-metric, scholion add-med,
an import, or the same actions on the page — erases the demonstration and says
so, naming every file it removed. Nothing is lost by that: the demonstration is
generated from a fixed seed, so scholion init --demo --dir <folder> builds it
again exactly as it was, while a series mixing invented values with measured ones
could not be separated afterwards by anybody. Only data carrying the mark can be
erased, so an ordinary profile — which carries no mark at all — is never
touched.

A published reference genome could be read under somebody else's laboratory
history.
A genome anybody may look at can be fetched to see the genomic layer
work, and it belongs to a real, consented, published person — not to the reader,
and not to the fictional one of the demonstration. Read beside either it produced
one case out of two people: this genotype, that history, a single report. The
fetched folder now says whose genome it is, and the genomic layer stays silent
while the two are in one profile, naming both sides and the folder to give it
instead. A genome folder that says nothing is still read as before: silence is
not a claim, and every genome anybody already has is unmarked.

A chip stopped being a chip by arriving as a VCF. ClinVar findings, the
secondary-findings list and polygenic scores can only be answered from a broad
call set, and the gate that closes them keyed on the CARRIER — a consumer array
file — rather than on what the input actually holds. A genotyping panel
distributed as a VCF (553 197 variants) and a table of chosen positions (48 838)
both went the other way and were told «this has not been annotated yet — run the
preparation», which is an invitation to do the exact thing the gate exists to
prevent. Breadth now decides: a chip, a genotype table, a panel, a low-pass
screen, a mostly-imputed file and half a call set all close those three paths,
whatever they arrived in. An input whose breadth could not be measured closes
them too — refusing a genome that could not be probed costs one command, and
opening a screen that was never measured costs a finding somebody may act on.

An approximate date stopped looking approximate the moment it was stored. A
form printing no draw date is read from two weaker witnesses — an «Ordered
Date», which some lipid panels print instead of one, and the file name. That
caveat was printed once, while importing, and afterwards the point sat in the
series indistinguishable from one dated by the draw itself. Every laboratory
point now records which of the three answered for it. Points stored before this
version say «not recorded» rather than claiming a form; scholion limits counts
them in one line instead of marking each, and re-running scholion ingest-labs
over the original forms fills them in.

**«A whole gen...

Read more