First real in-house dataset: 784 C. elegans Smart-seq3 worms, compiled end to end #350
lhqing
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
aging_SS3is the first dataset seqforge has compiled that nobody deposited anywhere: 784 single C. elegans worms, Smart-seq3, sequenced in-house, sitting on a lab filesystem with no accession, no paper and no archive record. That is the case the compiler was built for and the one it had never actually been run against, so this is a report on what it decided on its own, what it refused, and what had to be added.The dataset
day1–day17,CF/DA/N2)ce11+ WormBaseWS298{age}_{strain}_{sample_id}What seqforge worked out without being told
Only three things were stated: where the FASTQs are, the directory-name convention, and the genome. Everything below came out of the bytes or the code.
The chemistry.
smartseq3was resolved from the reads, not from the name of a folder. R1 begins withATTGCGCAATGat offset 0 — the Smart-seq3 TSO tag the KB entry gates on — and the UMI structure fell out as:Read roles. R1 →
umi_cdna, R2 →cdna. Worth noting because the KB entry'sfile_hintis_R1_/_R2_and these files are named_1.fq.gz/_2.fq.gz. The hint never matched; the roles were decided from the reads.Sample grouping. 1568 files → 784 samples, by filename, with no record set required to say so.
What it refused, and why that is the point
One directory,
day7_N2_19, held only_2.fq.gz— its R1 was missing. R1 is the read carrying the SS3 tag, so on its own bytes that sample is indistinguishable from bulk RNA-seq. seqforge stopped the entire compile at exit 3:A majority vote would have swallowed it: 783 cells say Smart-seq3. The refusal names the cell, the evidence hash, and two remedies. This is the single most valuable thing that happened in the whole run — a corrupt sample was caught by the compiler rather than by someone puzzling over a sparse column in the final matrix months later.
A second data problem surfaced the same way. Seven
day13_CF_*directories held both per-lane files and the merged file they were concatenated from — verified byte-exactly (L01 + L02 == merged, all 7 samples, both mates). Passing both would have doubled every read and invented 14 spurious samples.Per-sample metadata, and the change it needed
The directory name is the metadata:
day13_CF_26says age, strain and replicate. Getting that into the manifest turned out to be the one thing seqforge could not do.Two paths existed and neither worked. A user-written record set declares structure only — ADR-0034 forbids typed facts on it, because a typed slot reaches
assertedwith no span to check. And a README works but is dataset-scoped:_basis_forgrades those claimsinferredand fans one value across every sample. For a homogeneous batch that is merely weaker. For 9 ages and 3 strains it is wrong — one age stamped on 784 worms.ADR-0047 resolves it: a user record set may carry
free_text, neverattributes.The distinction is a slot versus a sentence. An attribute is believed for where it sits. Prose is believed for nothing — it becomes that sample's own document, and a claim leaves it only carrying a quote that greps back and a value that quote entails. The objection to allowing it was that such a quote greps back against a file its own author wrote to be grepped; true, and not a distinction, since ADR-0034's own remedy ("write it in a README and harvest that") is equally a file the author wrote that morning. What prose on a record actually buys is not credence but subject — which sample the words are about — plus a mandatory
labelrecording where they came from, sostrain=CFread out ofdirectory_name: day13_CF_26is visibly a filename convention and not a bench measurement.Generalisation was the design constraint, and the compiler learns nothing about this lab: it accepts prose attached to a record and a label saying where the prose came from. A batch that keeps its metadata in a filename, a sample sheet or a bench note writes a different generator script and changes nothing upstream.
No model was called. Harvest's LLM only ever proposes
{field, value, quote}; the verification is deterministic code. 1568 drafts (784 samples x age + strain) were generated mechanically from the directory names and checked byseqforge harvest verify --recordson a compute node with no network:Every sample now carries
ageandstrainatbasis: asserted,confidence: 1.0, and they render in the HTML report's metadata table with their provenance attached.The deliverable
2356 jobs:
umi_extractx784,star_umi_mapx784,umi_to_cramx784, plusgenome_index,load_genome, the plate-wideumi_countfan-in, andall. 0 blockers, 0 warnings, 0 of 784 cells dropped by the admission gate. Deliberately not submitted by seqforge — the last artifact is a Snakefile the user submits.The whole thing, from a cold start
This is the complete prompt. Nothing else was needed.
Note what is absent: no chemistry, no read roles, no adapter, no UMI offsets, no sample sheet, no accession. Those are the compiler's job. What a human has to supply is the genome (the one decision with no safe default), where the bytes are, and what the folder names mean.
Prerequisites
liulab-runtimealign-rnaimage carrying seqforge, STAR and liulab-genome.ce11assembly withWS298registered, and its STAR index built ahead of the run — seqforge only ever resolves an index (get_star_index), it never builds one. Building isliulab-genome's job:Genome("ce11").build_star_index(gtf="WS298", sjdb_overhang=149), wheresjdb_overhangis read length − 1.Two things that catch people on ircbc
Do not pass
--software-deployment-method apptainer. The Snakefile header recommends it and it is right on a modern host. ircbc's bare OS is CentOS 7 / glibc 2.17, so snakemake cannot run outside a container at all and singularity does not nest. Run snakemake inside the samealign-rnaimage the config names; the pinning then holds by construction rather than by the flag.Size the job by cores, not by the declared
mem_mb. The per-cell memory figure is a BAM-sort budget that scales with read depth, not with the genome — see #349. On a small genome likece11(1.3 GB index) enforcing it as a scheduler resource collapses concurrency for no reason.What the run cost
On one 160-core fat node,
umi_extractis single-threaded so ~160 run at once;star_umi_maptakes 8 threads so ~20 do. Whole dataset end-to-end is a single overnight job.Takeaways
assertedmetadata with no model call. Agents propose, code decides — the LLM proposes drafts and every one is span-verified deterministically, so an offline compute node with no credential produces the same manifest.All reactions