Skip to content
Specifications of SAM/BAM and related high-throughput sequencing file formats
Branch: master
Clone or download
Latest commit 82f7867 Jul 8, 2019
Permalink
Type Name Latest commit message Commit time
Failed to load latest commit information.
.circleci Adding circle-CI integration (PR #328) Mar 25, 2019
_includes Redo the index web page using Jekyll Nov 11, 2016
_layouts Add ga4gh-retrieval.md to README and index Dec 13, 2016
img Make images non-executable Mar 14, 2014
pub Fixed refget-openapi.yaml (#410) May 21, 2019
scripts Run BibTeX if needed Apr 8, 2019
.gitattributes Fix typo Mar 7, 2014
.gitignore Run BibTeX if needed Apr 8, 2019
BCFv1_qref.pdf Added compiled PDF docs to github page Nov 27, 2013
BCFv1_qref.tex Note that BCF1 is obsolete Jul 3, 2013
BCFv2_qref.pdf Added compiled PDF docs to github page Nov 27, 2013
BCFv2_qref.tex Fix minor typo in BCFv2 Jun 24, 2019
CRAMv2.1.pdf Update PDFs (RNAME allowed chars; CRAMv3 tag encodings etc; OA tag) Feb 1, 2019
CRAMv2.1.tex Avoid \usepackage{ulem} as it conflicts with latexdiff [minor] Jan 22, 2019
CRAMv3.pdf Update PDFs (VCF formatting; CRAM rANS/slice/container/index) Apr 9, 2019
CRAMv3.tex Merge pull request #412 from cmnbroad/cn_substitution_matrix Jun 26, 2019
CSIv1.pdf Update CSIv1.pdf, SAMv1.pdf, tabix.pdf to f35504b Nov 21, 2014
CSIv1.tex Rerun LaTeX by checking .aux/etc files rather than scraping .log Oct 1, 2018
CSIv2.pdf Added PDF version of CSIv2 and VCFv4.3 drafts Jan 14, 2015
MAINTAINERS.md add @jeromekelleher and @jmarshall as additional htsget maintainers (#… Jun 4, 2019
Makefile Also clean BibTeX output and log Apr 9, 2019
README.md Version 0.2 of the refget API as discussed in PR #327 Aug 15, 2018
SAMtags.pdf Update PDFs (RNAME allowed chars; CRAMv3 tag encodings etc; OA tag) Feb 1, 2019
SAMtags.tex SAM*.tex: Avoid end-of-sentence spacing after i.e. and e.g. [minor] May 21, 2019
SAMv1.pdf Update SAMv1.pdf (`@SQ-TP` added) May 14, 2019
SAMv1.tex Wording improvements to footnote [minor] Jun 20, 2019
VCFv4.1.pdf Update PDFs (VCF formatting; CRAM rANS/slice/container/index) Apr 9, 2019
VCFv4.1.tex Fix minor typo in the predecessors of BCF2 (#427) Jul 8, 2019
VCFv4.2.pdf Update PDFs (VCF formatting; CRAM rANS/slice/container/index) Apr 9, 2019
VCFv4.2.tex Fix minor typo in the predecessors of BCF2 (#427) Jul 8, 2019
VCFv4.3.pdf Update PDFs (VCF formatting; CRAM rANS/slice/container/index) Apr 9, 2019
VCFv4.3.tex Fix minor typo: alelle -> allele (#422) Jun 18, 2019
_config.yml Add latexdiff-vc rule to make diff/*.pdf Mar 25, 2019
htsget.md set htsget v1.2.0 May 14, 2019
htsget_interop.html Update htsget_interop.html Oct 4, 2018
index.md Version 0.2 of the refget API as discussed in PR #327 Aug 15, 2018
refget.md Product review committee changes v1 (#336) Oct 4, 2018
tabix.pdf Update PDFs (VCF formatting; CRAM rANS/slice/container/index) Apr 9, 2019
tabix.tex Cosmetic improvement in [beg,end) formatting Jun 6, 2018

README.md

SAM/BAM and related specifications

Links in bold point to the corresponding PDFs on this repository's GitHub Pages website.

Please request improvements or report errors using this repository, but see also the list of maintainers if you need to contact them directly.

Alignment data files

SAMv1.tex is the canonical specification for the SAM (Sequence Alignment/Map) format, BAM (its binary equivalent), and the BAI format for indexing BAM files. SAMtags.tex is a companion specification describing the predefined standard optional fields and tags found in SAM, BAM, and CRAM files. These formats are discussed on the samtools-devel mailing list.

CRAMv3.tex is the canonical specification for the CRAM format, while CRAMv2.1.tex describes its now-obsolete predecessor. Further details can be found at ENA's CRAM toolkit page. CRAM discussions can also be found on the samtools-devel mailing list.

The tabix.tex and CSIv1.tex quick references summarize more recent index formats: the tabix tool indexes generic textual genome position-sorted files, while CSI is htslib's successor to the BAI index format.

Unaligned sequence data files

We do not define or endorse any dedicated unaligned sequence data format. Instead we recommend storing such data in one of the alignment formats (SAM, BAM, or CRAM) with the unmapped flag set. However for completeness, we list the commonest formats below with external links.

FASTA is an early sequence-only format originally defined by William Pearson's tool of the same name.

FASTQ was designed as a replacement for FASTA, combining the sequence and quality information in the same file. It has no formal definition and several incompatible variants, but is described in a paper by Cock et al.

Variant calling data files

VCFv4.3.tex is the canonical specification for the Variant Call Format and its textual (VCF) and binary (BCF) encodings, while VCFv4.1.tex and VCFv4.2.tex describe their predecessors. These formats are discussed on the vcftools-spec mailing list.

BCFv1_qref.tex summarizes the obsolete BCF1 format historically produced by samtools. This format is no longer recommended for use, as it has been superseded by the more widely-implemented BCF2.

BCFv2_qref.tex is a quick reference describing just the layout of data within BCF2 files.

Transfer protocols

Htsget.md describes the hts-get retrieval protocol, which enables parallel streaming access to data sharded across multiple URLs or files.

Refget.md enables access to reference sequences using an identifier derived from the sequence itself.

You can’t perform that action at this time.