-
Notifications
You must be signed in to change notification settings - Fork 0
Home
Map protein-domain amino-acid coordinates to their underlying genomic CDS/UTR/intron structure, using any GENCODE, Ensembl, or NCBI RefSeq GTF.
For each input query - a protein_id or a transcript_id, optionally with an aa range - fastCDS answers two related questions:
- Mapping - which exact genomic bases code this domain?
- Structure - how is the whole transcript organised into 5'UTR / CDS / 3'UTR / intron, and where does the domain fall on it?
Three steps: get an index (build it from a GTF with index, or fetch a pre-built one), map your queries onto it, then plot.
flowchart LR
GTF[GTF annotation] --> INDEX([fastCDS index])
ZEN[Zenodo] --> FETCH([fastCDS fetch])
INDEX --> IDX[index<br/>human.idx]
FETCH --> IDX
IDX --> MAP([fastCDS map])
BED[query BED<br/>protein + aa range] --> MAP
MAP --> TSV[isoform_structure.tsv]
MAP --> B12[domain_blocks.bed<br/>BED12 for IGV/UCSC]
TSV --> PLOT([fastCDS plot])
PLOT --> STATIC[static figure<br/>.pdf / .png / .svg]
PLOT --> INTER[interactive viewer<br/>.html: js or plotly]
classDef cmd fill:#2f6db0,color:#ffffff,stroke:#1c4a7d,stroke-width:1px;
classDef file fill:#eef2ff,color:#111111,stroke:#9aa7d0,stroke-width:1px;
class INDEX,FETCH,MAP,PLOT cmd;
class GTF,ZEN,IDX,BED,TSV,B12,STATIC,INTER file;
The same four commands, in order:
fastCDS index gencode.v49.primary_assembly.annotation.gtf # 1a build an index
fastCDS fetch human --out human.idx # 1b get a pre-built index from Zenodo
fastCDS map --index human.idx \ # 2. map queries -> see Mapping
--bed queries.bed --out-dir results --output all
fastCDS plot --isoform results/isoform_structure.tsv \ # 3. plot -> see Plotting
--input-id TP53_DBD --out tp53.pdf| Command | Does | Page |
|---|---|---|
index |
Build a binary index from a GTF. | Building an index |
fetch |
Download a pre-built index from Zenodo. | Building an index |
map |
Map protein/domain queries to genomic structure. | Mapping |
plot |
Render a static (PDF/PNG) or interactive (HTML) figure. | Plotting |
The same workflow from Python:
import fastCDS as fc
idx = fc.fetch_index("human")
mapper = fc.Mapper(index=idx)
result = mapper.map_batch([
{"protein_id": "ENSP00000269305", "aa_start": 102, "aa_end": 292, "domain_id": "TP53_DBD"},
])
result.summary # DataFrame, one row per query
fc.plot(result, input_id="TP53_DBD", out="tp53.pdf")New here? Start with Installation, then walk the sidebar top to bottom. The Tutorials and Notebooks page has a copy-paste run from zero to a figure.
1 - How to install
2 - Building an index
(fastCDS index, fastCDS fetch)
3 - Mapping
(fastCDS map)
4 - Plotting
(fastCDS plot)
6 - Performance and benchmarking
7 - Reference