Skip to content

Download

github-actions[bot] edited this page Aug 24, 2026 · 23 revisions

Download

The Download page is the entry point for reusable files underlying the SAR11 Genome Atlas. Compact web-facing files are served directly from the atlas repository, while larger sequence, annotation, HMM, and analysis archives are distributed through the versioned Zenodo dataset. Direct files are marked Available, mixed cards Partially available, and Zenodo-derived files Embargoed.

Directly Available Atlas Datasets

The following compact resources can currently be downloaded directly from the atlas repository:

  • Genome metadata and quality estimates: subclade_master.tsv contains the complete metadata for 542 SAR11 genomes and 20 phylogenetic outgroups. SAR11_Atlas_542_CheckM2_v1.0.2_quality_report.tsv contains the CheckM2 quality estimates for the 542 released SAR11 genomes.
  • Orthogroup assignments and statistics: SAR11_Orthogroup_Assignments_542.tar.gz contains the core OrthoFinder 3 assignment, count, overlap, hierarchical-orthogroup, species-tree, and run-information files.
  • Orthogroup annotations and chart data: og_suggest.tsv contains representative annotations for all 4,577 orthogroups. The KO and COG count tables drive the pie charts, and the Pfam count table drives the bar chart in the OG Information Viewer.
  • OG representative protein sequences: OG_representative_sequences.faa contains one observed representative protein for each of the 4,577 orthogroups. representative_sequences.tsv records the selected sequence ID, length, alignment gap count, identity and difference counts, and percent change from the multiple-alignment consensus.
  • Resolved orthogroup gene trees: Resolved_Gene_Trees.txt.tar.gz contains 3,411 resolved gene trees.
  • Species phylogenies: the default rooted SAR11_165 IQ-TREE 2 phylogeny used for topology-based taxonomy assignment, plus the original SAR11_165 tree with outgroups. The bac120 IQ-TREE 2 tree and both FastTree results are retained only as comparison trees and are not used for taxonomy assignment.
  • All-vs-all ANI and AAI results: the directional FastANI v1.34 output, a symmetric 542-genome ANI matrix, and the CompareM v0.1.2 pairwise AAI summary.
  • Neighboring-gene network: neighbor_network.tsv is the directed edge table used by the Neighboring Network page.
  • CORGIAS results: corgias_network.tsv contains the significant phylogenetically informed OG associations calculated with the rooted SAR11_165 phylogeny and used by the network and result table. bac120_stat_sev.tsv contains the corresponding significant associations calculated with the rooted bac120 phylogeny.
  • High-similarity UniProt matches: gene-level, OG-representative, and AlphaFoldDB-confirmed match tables generated against UniProtKB release 2026_01 at a minimum of 85% identity.
  • Broader UniProt similarity results: uniprot_pid30_cov80_filtered_hits.tsv.gz contains the filtered 30%-identity search results used to identify homologous structure references.
  • Tara Oceans metatranscriptomic quantification: SAR11_merged_metaT.tsv.gz contains the gene-level quantification table for 509 runs used to derive OG Expression Scores.
  • SAR11 literature table: 250729_SAR11_paper_list.tsv, updated on 2025-07-29, contains the publication metadata displayed on the Literature page.

The Download page marks every link by source. Direct atlas files are labelled Available; Zenodo files under the current embargo are labelled Embargoed; cards containing both types are labelled Partially available. OG-specific neighborhood tables are downloaded separately from the Neighboring Genes page.

Zenodo Resources (Embargoed)

The following larger resources are hosted in the versioned Zenodo dataset and are currently marked Embargoed on the Download page:

  • Genome assemblies, predicted proteins, and GFF annotations for all 542 SAR11 genomes.
  • The complete protein-level annotation table for 675,669 predicted proteins.
  • The combined HMM library for 4,577 orthogroups.
  • Complete gene-coordinate, CORGIAS, UniProt similarity-search, and Tara Oceans metatranscriptome resources.

Genome Metadata

subclade_master.tsv is the complete metadata table with harmonized taxonomic labels. It records assignment evidence, confidence, and the phylogenetic or ANI/AAI reference used where applicable. Two derived forms are maintained for web components:

  • subclade.txt emphasizes numeric sampling depth and coordinates for the searchable table and map.
  • subclade_cat.tsv emphasizes categorical metadata for Taxonium coloring.

All three forms contain 542 SAR11 genomes and 20 phylogenetic outgroups. Marine Longhurst codes and descriptions are retained where applicable. Freshwater and other nonmarine records keep Longhurst fields as NA and are represented through habitat and waterbody fields instead. Use subclade_master.tsv unless a web-component-specific input is required.

SAR11_Atlas_542_CheckM2_v1.0.2_quality_report.tsv contains CheckM2 v1.0.2 completeness and contamination estimates for the 542 SAR11 genomes; the 20 phylogenetic outgroups are not included.

Orthogroup Assignments And Trees

The OrthoFinder archive includes Orthogroups.tsv, Orthogroups.GeneCount.tsv, Orthogroups_UnassignedGenes.tsv, overall and per-species statistics, species-overlap counts, the root-level hierarchical orthogroup table, the rooted node-labeled species tree, and a README recording the 542-genome analysis.

The resolved gene-tree archive contains trees for the 3,411 orthogroups with at least four protein sequences. The archive README records the analysis workflow and software versions.

The Zenodo-hosted SAR11_Orthogroups_4577.hmm.tar.gz archive provides one profile HMM for each orthogroup in a combined HMM file; individual profiles can be extracted with hmmfetch.

OG Representative Sequences

The compact files used by the browser-based OG Representative Similarity Search are available directly:

For each OG, the representative is an observed member sequence rather than a synthetic consensus. These representatives support rapid exploratory assignment but do not capture all within-OG diversity. For more sensitive assignment, use the combined OG HMM profiles with HMMER.

Species Phylogenies

SAR11_542_SCG165_iqtree_rooted_SAR11_only.tree is the default tree displayed in the Genome Information and OG Information Taxonium views. The original SAR11_165 IQ-TREE 2 tree was rooted using 20 alphaproteobacterial outgroups before those outgroups were pruned, leaving the 542 SAR11 tips.

The original 562-tip SAR11_165 and bac120 IQ-TREE 2 trees, the rooted 542-tip bac120 IQ-TREE 2 tree, and the two FastTree trees are retained for comparison.

Functional Annotations

og_suggest.tsv summarizes representative KO, COG, and Pfam evidence for browsing. Representative values describe the most supported annotations within an orthogroup and are not necessarily shared by every member. The accompanying og_ko_counts.tsv, og_cog_counts.tsv, and og_pfam_counts.tsv files preserve the within-OG counts used by the interactive charts.

The full protein-coding gene annotation resource is a separate, larger table for all 675,669 proteins. It combines orthogroup IDs with outputs from COGclassifier, KOfamScan, PfamScan, quickARSC, and related protein-level fields. The complete all_prot_annotations.tsv table is distributed through Zenodo.

UniProt And AlphaFoldDB Results

All SAR11 proteins were searched against UniProtKB release 2026_01 Swiss-Prot and TrEMBL with DIAMOND v2.1.10.164 using a minimum amino-acid identity of 85% and --max-target-seqs 1. Candidates for the AlphaFoldDB-linked web table were additionally required to cover at least 80% of both query and subject sequences.

  • uniprot_pid85_gene_hits.tsv contains 226,536 matched SAR11 proteins across 1,912 orthogroups.
  • uniprot_pid85_og_representative_hits.tsv contains one representative UniProt hit for each of those 1,912 orthogroups.
  • uniprot_pid85_afdb_matches.tsv contains 3,622 matches with confirmed AlphaFoldDB models across 1,686 orthogroups. Of these, 3,082 links across 1,409 orthogroups are exact sequence matches (100% identity and 100% query and subject coverage); the remaining 540 links represent close sequence matches.
  • uniprot_pid30_homologous_afdb_matches.tsv contains 11,051 confirmed distant homologous matches across 2,993 orthogroups after applying at least 30% identity, at least 80% query and subject coverage, and an E-value no greater than 1e-5.
  • uniprot_structure_references.tsv is the downloadable compact Web table containing 8,057 AlphaFoldDB structure references. It prioritizes exact and close sequence matches and uses distant homologous matches only for the 1,308 additional orthogroups without either class, covering 2,994 orthogroups in total.

A UniProt hit is a sequence-similarity result, not proof that every orthogroup member has the same sequence or structure. Use identity, coverage, domain annotations, and AlphaFold confidence together when interpreting a match.

A broader DIAMOND search using a minimum identity of 30% and --max-target-seqs 10 was filtered at 80% query coverage, 80% subject coverage, and an E-value of 1e-5. For the combined structure-reference table, exact matches require 100% identity and 100% query and subject coverage; close matches require identity ≥85% and both coverage values ≥80%, excluding exact matches; and distant homologous matches require identity ≥30%, both coverage values ≥80%, and E-value ≤1e-5. The complete 8,300,702-row filtered table is distributed as the 199.50 MB compressed uniprot_pid30_cov80_filtered_hits.tsv.gz Zenodo file; compact AFDB-confirmed tables are available directly from the atlas.

Network And Expression Resources

The directly available network files include the complete edge tables used by the Neighboring Network and CORGIAS Network interfaces. The CORGIAS viewer edge table (corgias_network.tsv) was calculated using the rooted SAR11_165 IQ-TREE 2 phylogeny. The directly downloadable bac120_stat_sev.tsv table contains 18,606 significant associations from the parallel CORGIAS analysis using the rooted bac120 IQ-TREE 2 phylogeny. Larger supporting files are distributed through the Zenodo release:

  • gene_coordinates_with_og.tsv, which underlies the neighborhood, operon, and neighboring-network analyses.
  • CORGIAS_result.csv, the complete CORGIAS analysis output beyond the compact significant-edge table.
  • SAR11_merged_metaT.tsv.gz contains gene-level quantification results for 509 Tara Oceans metatranscriptomic runs. Join its feature identifiers to the protein and orthogroup assignments, then sum TPM by sample and orthogroup, to reproduce the aggregate Expression Scores used by the Metatranscriptome Viewer.

All-vs-all ANI And AAI

fastani_SAR11_542.out is the original directional five-column FastANI v1.34 output: query genome, reference genome, ANI, matched fragments, and total query fragments. fastani_SAR11_542_symmetric_matrix.tsv is a 542 x 542 matrix derived from that output. Reciprocal ANI estimates are averaged when both directions are reported; a single reported direction is retained when its reciprocal comparison is absent; comparisons absent in both directions are shown as NA; and the diagonal is set to 100. Genome names in the matrix omit the .fna suffix.

comparem_aaiwf_SAR11_542_out_summary.tsv is the all-vs-all amino-acid identity summary generated with CompareM v0.1.2 aai_wf from the predicted proteins of the same 542 genomes. It reports the protein-coding gene counts for each genome, number of detected orthologs, mean and standard deviation of AAI, and orthologous fraction for each genome pair.

The ANI and AAI files are comparative-genome measurements and are not presented as a formal SAR11 species classification.

Distribution And Versioning

Small interactive tables are downloaded from the atlas repository. Large sequence collections, complete annotations, HMM profiles, CORGIAS output, UniProt similarity results, gene-coordinate data, and metatranscriptomic quantification are downloaded from Zenodo. The OrthoFinder assignment and resolved gene-tree archives currently remain direct atlas downloads.

Choosing A File

Use SAR11 Genome Information to inspect genome metadata before downloading genome resources. Use All OG List or OG Information Viewer to identify orthogroups before downloading assignments, annotations, trees, or HMM profiles.

Large collections are distributed as compressed, versioned archives. Check each archive's README and column definitions before combining it with results from another atlas release.

Clone this wiki locally