Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

17 Commits
 
 
 
 

Repository files navigation

South African Blood Regulatory (SABR) Resource

The South African Blood Regulatory (SABR) Resource generated whole genome sequencing (WGS) and blood RNA-seq data from over 600 individuals spanning three South Eastern Bantu-speaking groups. This was a collaboration between Variant Bio and Michele Ramsay's group at Wits University, and includes individuals from the AWI-Gen cohort. These data were used to map genetic variants that impact gene expression, splicing, and cell type levels. A full description of the resource can be found in our manuscript.

A map of blood regulatory variation in South Africans enables GWAS interpretation. Castel et al. 2025. Nature Genetics.

Functional genomics resources are critical for interpreting human genetic studies, but currently they are predominantly from European-ancestry individuals. Here we present the South African Blood Regulatory (SABR) resource, a map of blood regulatory variation that includes three South Eastern Bantu-speaking groups. Using paired whole genome and blood transcriptome data from over 600 individuals, we map the genetic architecture of 40 blood cell traits derived from deconvolution analysis, as well as expression, splice, and cell type interaction quantitative trait loci. We comprehensively compare SABR to the Genotype Expression (GTEx) Project and characterize thousands of regulatory variants only observed in African-ancestry individuals. Finally, we demonstrate the increased utility of SABR for interpreting African-ancestry association studies by identifying putatively causal genes and molecular mechanisms through colocalization analysis of blood-relevant traits from the Pan-UK Biobank. Importantly, we make full SABR summary statistics publicly available to support the African genomics community.

Table of Contents

  1. Cohort and Data Overview
  2. Summary Statistics
  3. Analysis Results
  4. Zenodo
  5. Controlled Access Data
  6. Methods

Cohort and Data Overview

The SABR cohort consists of 754 individuals who participanted in the AWI-Gen cohort and were recontacted and reconsented for this study. Venous whole blood samples were taken and used to carry out WGS at a median depth of 5.1x and paired-end, stranded RNA-sequencing with globin and rRNA depletion at a median depth of 30M mapped read pairs. All data are made available for non-commercial use only.

All coordinates are provided in GRCh38

  • Individual-level sequencing data quality control metrics and inclusion in downstream analyses - GitHub

Summary Statistics

Variant-level summary statics from WGS genotyping and full summary statistics from xCell GWAS and QTL mapping are publicly available via the links listed below. All full summary statistics files (.txt.bgz) are provided with a corresponding index file (.tbi). Note, genome-wide summary files are provided through an AWS S3 bucket and will require the AWS CLI tool to download.

You can list all files availabe in the S3 bucket using the following command:

aws s3 ls --no-sign-request s3://public.us-prod.variantbio.com/SABR/

Genotype Data

  • South African enriched, putatively functional alleles - GitHub
  • Variant-level summary statistics from imputed mid-pass WGS including allele frequencies and functional annotations - s3://public.us-prod.variantbio.com/SABR/VARS/SABR_variant_summary.txt.bgz
  • Description of variant-level summary statistics fields - GitHub

GWAS

  • List of xCell types included in analyses - GitHub
  • xCell codes - GitHub
  • Summary of genome-wide significant loci (p < 5e-8) identified - GitHub
  • Full summary statistics outputted by Hail for each GWAS, including p-value, beta (alt allele), standard error, minor allele frequency - s3://public.us-prod.variantbio.com/SABR/XCELL_GWAS/

QTL Mapping

Gene-level results and full variant level summary statistics outputted by fastQTL are provided for all QTL mapping runs.

  1. Expression QTLs (eQTLs)
    • Gene-level results - GitHub
    • eVariant annotations - GitHub
    • Conditionally independent eQTLs - GitHub
    • Conditionally independent eVariant annotations - GitHub
    • Nominally significant structural variant eQTLs - GitHub
    • Full summary statistics - s3://public.us-prod.variantbio.com/SABR/EQTL/SABR_eQTL_allpairs.txt.bgz
    • Conditionally independent eQTL summary statistics - s3://public.us-prod.variantbio.com/SABR/EQTL/SABR_eQTL_conditional_variants.txt.gz
  2. Splice QTLs (sQTLs)
    • Gene-level results - GitHub
    • sVariant annotations - GitHub
    • Full summary statistics - s3://public.us-prod.variantbio.com/SABR/SQTL/SABR_sQTL_allpairs.txt.bgz
  3. Cell-type Interaction eQTLs (ieQTLs)
    • Gene-level results - GitHub
    • ieVariant annotations - GitHub
    • Full summary statistics - s3://public.us-prod.variantbio.com/SABR/IEQTL/

Analysis Results

xCell Disease Modeling

  • Modeling results for each xCell type by disease - GitHub

Colocalization Analyses

  1. SABR QTLs x PAN-UKBB African GWAS
    • List of African GWAS included in analysis - GitHub
    • Colocalization results - GitHub
    • Colocalization lead variant annotations - GitHub
  2. SABR eQTLs x PAN-UKBB Multi-ancestry GWAS
    • List of multi-ancestrty (MA) GWAS included in analysis - GitHub
    • Colocalization results - GitHub

Zenodo

In addition to being avilable here and on AWS, eQTL and sQTL summary statistics have been uploaded to a Zenodo repository.

Controlled Access Data

Indvidual-level data are availabe to authorized users purusing reasearch in line with informed consent and ethical approvals. Data are provided through the European Phenome Genome Archive (EGA) project EGAS50000001008. The following is a list of data availble via controlled access.

  • Joint called and imputed genotype VCF
    • HAIL-ZAAG-WGS-MP-1-2S.metadata.txt
    • HAIL-ZAAG-WGS-MP-1-2S.vcf.bgz
    • HAIL-ZAAG-WGS-MP-1-2S.vcf.bgz.tbi
  • Extended indivividual-level metadata, including disease status
    • SABR_subject_phenotypes.txt
  • xCell enrichment scores
    • SABR_xcell_quantifications.txt
  • Expression quantifications (counts, TMM, TPM)
    • SABR_gene_reads.gct.gz
    • SABR_gene_tmm.bed.gz
    • SABR_gene_tpm.gct.gz
  • Splice quantifications (junction counts, clusters)
    • SABR_leafcutter_junctions.bed.gz
    • SABR_leafcutter_clusters.bed.gz
  • xCell GWAS input files (covariates)
    • SABR_gwas_covariates.txt
  • QTL mapping input files (covariates, normalized quantifications)
    • SABR_EQTL_combined_covariates.txt
    • SABR_EQTL_expression.bed.gz
    • SABR_SQTL_combined_covariates.txt
    • SABR_SQTL_leafcutter.bed.gz
  • Whole-genome sequencing data (FASTQs)
  • RNA-sequencing data (FASTQS)

Methods

All methods are described in detail in the Supplementary Materials provided with our manuscript. Below, we briefly describe the methods used and link to relevant software and pipeline pages.

Genotype Calling and Imputation

Genotype calling and imputation from mid-pass WGS data was carried out as described in Emde et al.. Code for mid-pass genotype calling and imputation is available on GitHub.

Expression Quantification, Splice Quantification, and QTL Mapping

The GTEx/TOPMed v8 pipeline was used for read mapping, quantifying expression levels, normalizing data, and mapping eQTLs with fastQTL. RNA-SeQC (v2.3.6) was used for gene quantification of collapsed genes with GENCODE v34 annotations. Splicing quantification and sQTL mapping was carried out using the approach described by the GTEx Consortium. Cell type estimation and interaction expression QTL mapping was carried using the approach described by Kim-Hellmuth et al.. For eQTL and ieQTL mapping, 60 PEER factors and 20 genotype PCs were used as covariates in addition to age, sex, and mean WGS depth. For sQTL mapping, 15 PEER factors and 20 genotype PCs were used as covariates in addition to age, sex, and mean WGS depth.

xCell GWAS

All cell types with enrichment scores > 0 in > 50% of participants were used for GWAS. Cell type enrichment scores were inverse normal transformed and GWAS were run using the linear_regression_rows() function in Hail v0.2 with the following covariates: age, sex, sex*age, sex*age^2, mean WGS depth, and 20 genotype PCs.

Colocalization

A subset of African and multi-ancestry GWAS from the Pan-UKBB were used for colocalization analysis. Colocalization analysis was carried out using coloc v5 for all significant QTLs (FDR < 5%) within 500kb of a significant (p<5e-8, multi-ancesrtry GWAS) or suggestive (p<5e-6, African ancestry GWAS) association signal. Minor allele frequencies, p-values, and a prior12 value of 1e-5 were used when running coloc.

About

South African Blood Regulatory Resource

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors