Skip to content

Specifications of SAM/BAM and related high-throughput sequencing file formats

Notifications You must be signed in to change notification settings

carsonhh/hts-specs

 
 

Repository files navigation

SAM/BAM and related specifications

Quick links

HTS-spec GitHub page SAMv1.pdf CRAMv2.1.pdf BCFv1.pdf BCFv2.1.pdf CSIv1.pdf tabix.pdf VCFv4.1.pdf VCFv4.2.pdf

Alignment data files

SAMv1.tex is the canonical specification for the SAM (Sequence Alignment/Map) format, BAM (its binary equivalent), and the BAI format for indexing BAM files.

CRAMv2.1.tex is the canonical specification for the CRAM format. Further details can be found at ENA's CRAM toolkit page.

The tabix.tex and CSIv1.tex quick references summarize more recent index formats: the tabix tool indexes generic textual genome position-sorted files, while CSI is htslib's successor to the BAI index format.

Variant calling data files

VCFv4.1.tex and VCFv4.2.tex are the canonical specifications for the Variant Call Format and its textual (VCF) and binary encodings (BCF 2.x).

BCFv1_qref.tex summarizes the obsolete BCF1 format historically produced by samtools. This format is no longer recommended for use, as it has been superseded by the more widely-implemented BCF2.

BCFv2_qref.tex is a quick reference describing just the layout of data within BCF2 files.

About

Specifications of SAM/BAM and related high-throughput sequencing file formats

Resources

Stars

Watchers

Forks

Releases

No releases published

Packages

No packages published