Running the 784-worm chimeric plate: how it was driven, and what the driving taught us #450
Replies: 3 comments
FINAL — the run completed 2026-08-21. Every open item below is resolved.784/784 cells, The eight open items
Estimates against outcome
The fan-in estimate is the instructive one. It would have been correct at the declared 8 threads — Corrections to this log
Peak memory figures quoted mid-run were node-total. The samplers in these scripts sum RSS across What this run produced beyond the dataFour issues, one PR, one measurement, all from driving it:
The result5.886% pooled bacterial share, 0.753% by the typical worm, working range 4.93-5.89%. The rise is real; the monotonic reading is not — day3 falls below day1, day7 returns to baseline, For the next session
|
|
Another engineering finding off this run, filed as #457. Opening Cause: none of the three map modules set any splice parameter, so STAR's derived The Numbers, method and what is not claimed are in #457. Not fixing it here — the right ceiling is |
The multimapper question this run opened is answered — #427 is closedThis log recorded that 32.593% of kept records are multiply placed, and used that to call #425's They are ribosomal RNA. The chrI 45S rDNA array is collapsed in ce11 to ~2.1 copies of a 7,197 bp And the Two things from this log get sharper as a result:
The full report, with every table and the ruled-out list, is Three follow-ups filed: liuhlab/liulab-genome#111, #471, #472. The recommendation was no bar and no |
Uh oh!
There was an error while loading. Please reload this page.
What is being run
The in-house SMART-seq3 aging plate — 784 worms, 9 ages x 3 strains, one animal per library — mapped
against the chimeric
ce11_ecHT115reference so the bacterial share of each library becomes a number.Chimeric arm only; the plain arm was deliberately not re-run at 784 (see "Decisions taken" below).
Workspace
aging_SS3/seqforge, pipeliness3-784-chimera-c54b1418fa98, six shards acrosscpu03-cpu06 + fat01 + fat02,
WORKFLOW_VERSION 2026.8.18.How it was actually driven
Enough detail to repeat it. Every seqforge/snakemake invocation goes through
script/sfrun, thecontainer shim.
1. Refresh the code — on the LOGIN node, because compute nodes have no internet and
gitis noteven installed on them (CentOS 7 base).
2. Rebuild the metadata chain —
make_metadata.py->harvest verify->mk_assertions.py.Deterministic and offline: the drafts are parsed from directory names, so no model runs. Regenerated
rather than reused; the output was byte-identical to the August files, which confirms
harvest verifyhas not moved.
3. Fill the manifest —
manifest fill <1568 files> --records --assertions --organism 6239 --offline.Not
seqforge run --no-llm: that path skips harvest and leavesassertions_path = None, so themanifest would come out with no age and no strain and nothing would say so.
4. Fresh recipe, fresh
--id—processing new --assembly ce11_ecHT115 --annotation wormbase_ws298+refseq_rs_2025_06_26 --threads 8 --mem-gb 48 --id ss3-784-chimera. Fresh becauserun_idfolds the recipe's fill-time workflow stamp, so re-composing an older recipe re-stages intoa directory that run already owns (#447).
5. Compose, then dry-run — 5493 jobs over 784 cells, module
map/star-umi-chimeraselected offthe assembly name with no second flag.
6. Materialize the genome index ONCE —
snakemake --cores 1 results/index/ce11_ecHT115— beforeany shard starts. See hazard 2.
7. Partition —
partition.pydoes longest-processing-time-first bin packing weighted by fragmentcount against heterogeneous nodes, emitting one list per node. fat nodes get 234 cells (20 slots),
cpu nodes 79 (7 slots), so all six are projected to finish together.
8. Six shards, each
sbatch --partition=... --nodelist=<node> --cpus-per-task=<n> --export=ALL,BATCH=<node> run-chimera-shard.sbatch, each asking for its own cells' four leaves andnever for
all.9. The fan-in, one job, gated
--dependency=afteranyon all six —afteranynotafterok,because
--keep-goingmakes a non-zero exit the normal outcome of a shard that refused one cell.Logistic observations
Things that would have silently produced a wrong or incomplete result
The obvious file glob is wrong.
fastq/*/*.fq.gzyields 1597 files; the dataset has 1568.The extras are 28 per-lane source files under
day13_CF_26..32— already concatenated into themerged pairs, so passing both would have doubled every read for seven worms — plus one orphan mate.
It surfaced only because Singularity caps its argv JSON at 128 KiB and refused, initially with a
misleading error about the runtime engine; bisecting argument counts produced the real message. The
shell never complained. Use the file list
records.yamldeclares, not a glob.Six shards race
rule genome_index. It is shared upstream work and is not part of the disjointper-cell split. Its
run:block doessymlink_to(), which raises on an existing link, and snakemakeremoves an output before re-running a job. Up to five of six shards can lose their entire run.
Materialize it once first. By contrast
rule load_genome'stouch()flag is safe to race, and thedifference is not visible from either rule.
--export=BATCH=<node>strips the environment. The bare form setsSLURM_EXPORT_ENVto that onename, so the job starts with no
PATHorHOME. Must be--export=ALL,BATCH=<node>.--notempis mandatory for a shard, and the reason is module-specific.split_chimera'sper-Component BAMs are
temp()with two consumers:unique_to_cramper cell, and the plate-wideumi_count. A shard's DAG sees only the first and reclaims them when the CRAMs land; the fan-in thenfinds every input missing, re-runs
split_chimera, whose own input is the STAR BAM — alsotemp(),also reclaimed — and silently re-maps the whole plate.
--nolockis required and throws away the guard that catches a bad split. Six instances cannotshare one lock. The only thing between that and a corrupt output is a human verifying the lists are
disjoint. Do it explicitly: union must be 784 lines with zero duplicates.
<cell>.qc.json.gzis NOT a completion signal.rule qc_bundledoes not depend on the CRAMs, soit lands before them. A shard naming only the QC bundle — the natural shorthand, since it is what the
report reads a cell's row from — would finish every cell with its archives never written, and nothing
downstream consumes an archive, so nothing would say so. Name all four leaves.
Measurement traps
sacctis unavailable on this cluster (accounting DB refuses connections), soMaxRSSper job doesnot exist. Any resource figure has to come from your own sampler.
The peak-RSS sampler in the run scripts measures NODE-TOTAL, not per-process. It sums RSS across the
process session, which double-counts the shared genome segment once per attached STAR job. Honest
per-rule memory needs
/proc/<pid>/smaps_rollup(Pss, private RSS), because the index is one ~1.3 GBSysV segment with
nattchequal to the number of attached STAR processes — not one copy per process.pspercentages are instantaneous. For setting athreads:declaration what you want is(utime+stime)/elapsedintegrated over a process's whole life.Scheduling and utilisation
Nodes run well below saturation, and the cause is a thread declaration. Observed mid-run: cpu03 at
44% (load 24.6/56), fat01 at 72% (115.8/160).
split_chimeraholds an 8-core slot and delivers~1.7 cores — 24 OS threads, most blocked behind a serial per-record Python routing loop — stranding ~6
cores per running job. STAR by contrast converts its 8 threads at ~94% efficiency. Filed as #449; the
structural point is that the module already solves this on the memory axis (
memory.pyturns one recipefigure into three per-rule-class requests) and has no equivalent for threads.
Work is back-loaded. Each shard runs most of its mapping before any of its conversion, so a
completed-cell count stays near zero for hours and then finishes in a rush. Partly the thread
reservation above; possibly also
--notempremoving snakemake's temp-reclamation scheduling incentive.Do not read early low completion as trouble.
A canary shard was worth less than expected. One shard was submitted alone to validate before the
other five. Because conversions queue behind all of that shard's mapping, the end-to-end check it
existed to provide was hours away, while the target list was already evidenced by a finished
pilot16r2cell under the same workflow file. Cost: the other five nodes idle. If the stage you wantto validate is at the back of the DAG, a canary does not validate it — find the evidence elsewhere.
Things that were fine, contrary to expectation
Log volume is a non-issue: 0.24-0.61 MB per shard after four hours. Worth recording because the
August run left a 7.2 GB
.errfor a single mapping job. The difference is #429 declaring STAR's runfiles as rule outputs instead of letting them stream to stderr — a fixed regression, noted so nobody
re-opens it.
Snakemake's coordinator overhead is modest at 5493 jobs: 199 MB RSS,
.snakemakemetadata 6.6 MBover 5404 files. (It does carry ~384 threads at ~47% CPU sustained, which is unexplained but harmless
here.)
Memory is nowhere near the request. The recipe asks 48 GB per mapping job and derives a 36 GiB
--limitBAMsortRAMcap; measured peak on a whole 20-slot node is a small fraction of that. The cappermits rather than reserves. A
mem_gbticket is still owed.Decisions taken, and what they cost
Chimeric arm only, no plain 784 arm. #425's stated methodology was both arms under the same code,
because two of its five items are arm-vs-arm comparisons. Those now rest on
pilot16r2's 16-cellcomparison — defensible, since both arms there ran under #429-era code on the same cells, but the
research artifact must say so and must not present a 784-cell bacterial share beside a 16-cell
multimapper bound as one result.
--mem-gb 48kept unchanged despite being ~20x the measured peak. In the shard model(
--cores N, no--resources mem_mb) it throttles nothing, so cutting it buys no throughput whilerisking a FATAL on the deep tail; and keeping it identical to the pilot's recipe is what keeps the
16-cell result comparable.
Two fat nodes and four cpu nodes; cpu07/cpu08 left free for other lab members.
Corrections to earlier claims in this run
Recorded because they were acted on before being checked:
multimapper". It is not —
split.pyroutes it by its representative record's Component and keepsit, marked.
multiplacedis a subset ofkept;droppedholds only unmapped/secondary/supplementary.The point value is contaminated in both directions, and the honest bracket is wide.
share of kept records (1.445%), per-cell median of the same (0.978%), and per-cell median of
unique/kept_total(the 0.144% -> 2.081% day1/day17 figures). They differ by up to 1.5x. Label whichone you mean.
partition.pyconstant (~6xconservative), then upward again once node utilisation turned out to be 44-72% rather than saturated.
Open items, to be resolved before this is finalized
combined.ce11.h5adandcombined.ecHT115.h5ad--notempmultiplier (provisional: ~3.1 GB/cell, ~3 TB peak)split_chimeracost normalised to records/sec, and the profiling research task on split_chimera spends 81% of its wall time rebuilding records in Python, over half of that round-tripping aux tags #449seqforge reportbehaviour at 784 samples (that code was exercised on 16)All reactions