-
Notifications
You must be signed in to change notification settings - Fork 3
Home
Kenji Fukushima edited this page Jun 8, 2026
·
9 revisions
AMALGKIT builds RNA-seq expression datasets from public SRA records and private FASTQ files. The current CLI is Python-only: the main pipeline no longer requires R, Rscript, or R packages.
Good first pages:
The usual public-SRA workflow is:
dataset -> metadata -> select -> getfastq -> quant -> merge -> finalize
Cross-species projects often add ortholog-aware normalization and filtering after merge:
merge + BUSCO/orthogroups -> cstmm -> wsfilter -> csfilter -> finalize
Private FASTQ projects start by integrating local files into metadata:
integrate -> getfastq -> quant -> merge -> finalize
Validation and recovery are separate maintenance steps:
sanity -> rerun
amalgkit dataset --rule_set base --out_dir ./ --overwrite yes
amalgkit metadata \
--out_dir ./ \
--entrez_email example@email.com \
--search_string '("platform illumina"[Properties]) AND ("type rnaseq"[Filter])'
amalgkit select --out_dir ./
amalgkit getfastq --out_dir ./
amalgkit quant --out_dir ./ --build_index yes
amalgkit merge --out_dir ./
amalgkit finalize --out_dir ./ --batch_effect_alg noAdd BUSCO-based cross-species steps when you have compatible transcriptome FASTA files or BUSCO tables:
amalgkit busco --out_dir ./ --lineage eukaryota_odb12
amalgkit cstmm --out_dir ./ --dir_busco ./busco
amalgkit wsfilter --out_dir ./
amalgkit csfilter --out_dir ./ --metadata ./wsfilter/metadata.tsv --dir_busco ./busco
amalgkit finalize --out_dir ./ --metadata ./csfilter/metadata.tsv --batch_effect_alg latent_glm| Command | Role |
|---|---|
amalgkit dataset |
extract bundled datasets, initialize workspaces, and export select_rules.tsv
|
amalgkit metadata |
retrieve SRA metadata by Entrez query or species-wise batch files |
amalgkit integrate |
add local FASTQ files to AMALGKIT metadata |
amalgkit select |
apply select_rules.tsv and mark rows for downstream processing |
amalgkit getfastq |
download public SRA data or stage private FASTQ files |
amalgkit quant |
quantify transcript abundance with kallisto or oarfish |
amalgkit merge |
merge per-run quantification into per-species tables |
amalgkit busco |
create BUSCO tables for ortholog-aware downstream steps |
amalgkit cstmm |
run cross-species TMM normalization using single-copy genes |
amalgkit wsfilter |
perform within-species outlier filtering |
amalgkit csfilter |
perform cross-species outlier filtering |
amalgkit finalize |
export final expression tables and optional batch-corrected outputs |
amalgkit sanity |
check expected outputs in a workspace |
amalgkit rerun |
rerun failed targets recorded by sanity
|
These commands existed in older AMALGKIT releases but are not part of the current CLI:
| Old command | Current workflow |
|---|---|
amalgkit config |
use amalgkit dataset --rule_set ... to write select_rules.tsv, then run amalgkit select
|
amalgkit curate |
run amalgkit wsfilter, optionally amalgkit csfilter, then amalgkit finalize
|
amalgkit csca |
use amalgkit csfilter for cross-species filtering and amalgkit finalize for final tables |
-
Start
-
Active commands
-
Migration and release