-
Notifications
You must be signed in to change notification settings - Fork 0
Home
George Muñoz edited this page Jun 3, 2026
·
12 revisions
Map protein-domain amino-acid coordinates to their underlying genomic CDS/UTR/intron structure, using any GENCODE, Ensembl, or NCBI RefSeq GTF.
For each input query — a protein_id or a transcript_id, optionally with an aa range — prot2exon answers two related questions:
- Mapping — which exact genomic bases code this domain?
- Structure — how is the whole transcript organised into 5′UTR / CDS / 3′UTR / intron, and where does the domain fall on it?
The whole workflow is four commands, used in order:
prot2exon index gencode.v49.primary_assembly.annotation.gtf # 1a build an index
prot2exon fetch human --out human.idx # 1b get a pre-built index from Zenodo
prot2exon map --index human.idx \ # 2. map queries -> see Mapping
--bed queries.bed --out-dir results --output all
prot2exon plot --isoform results/isoform_structure.tsv \ # 3. plot -> see Plotting
--input-id TP53_DBD --out tp53.pdf| Command | Does | Page |
|---|---|---|
index |
Build a binary index from a GTF. | Building an index |
fetch |
Download a pre-built index, or download a GTF and build one. | Building an index |
map |
Map protein/domain queries to genomic structure. | Mapping |
plot |
Render a static (PDF/PNG) or interactive (HTML) figure. | Plotting |
The same workflow from Python:
import prot2exon as p2e
idx = p2e.fetch_index("human")
mapper = p2e.Mapper(index=idx)
result = mapper.map_batch([
{"protein_id": "ENSP00000269305", "aa_start": 102, "aa_end": 292, "domain_id": "TP53_DBD"},
])
result.summary # DataFrame, one row per query
p2e.plot(result, input_id="TP53_DBD", out="tp53.pdf")New here? Start with Installation, then walk the sidebar top to bottom. The Tutorials and Notebooks page has a copy-paste run from zero to a figure.
1 - How to install
2 - Building an index
(fastCDS index, fastCDS fetch)
3 - Mapping
(fastCDS map)
4 - Plotting
(fastCDS plot)
6 - Performance and benchmarking
7 - Reference