Skip to content

Binary traits and PCGC

Moonseong Jeong bronsonj98@g.ucla.edu edited this page Sep 27, 2026 · 5 revisions

Binary traits and PCGC

PCGC estimates genetic variance on a liability scale while accounting for case–control sampling and individual disease risks. Select it with --binary-method.

Preparation needs individual genotypes and binary phenotypes. It produces compatible score and reference moments together; ordinary logistic GWAS beta/SE files and ordinary LD scores are not interchangeable with these inputs.

Inputs

  • BED or biallelic diploid PGEN genotypes, with their companion files.
  • A table with FID IID Y, where Y is 0 for controls and 1 for cases. Include numeric risk covariates or a column of supplied population risks if needed.
  • Population prevalence K. The sample case fraction P is computed from Y.
  • Population genotype means and inverse standard deviations, on the same SNP order and counted alleles as the genotypes.
  • Optional nonnegative SNP annotation weights, including overlapping columns.

The population-scale TSV has exactly SNP A1 A2 MEAN INV_SD. MEAN refers to the reader's counted allele: BED A1 or PGEN REF. A saved SUMMIT genotype-scale directory is also accepted; if it was prepared with a build label, supply the same --genome-build. Build labels are otherwise optional. Do not estimate these scales from an ascertained study and label them population scales.

The sample table selects the analyzed people. IDs and SNPs must be unique; missing phenotypes, risks, or selected covariates are rejected. Encode categorical covariates numerically and omit an intercept column; the risk fit includes one.

Prepare and fit

For age and another exogenous risk factor:

summit --binary-method pcgc \
  --make-binary-sumstats people.tsv --geno study.bed \
  --binary-prevalence 0.1 \
  --binary-scale population_scale.tsv \
  --binary-covariates AGE,RISK_FACTOR \
  --nvecs 256 --memory-gib 4 \
  --num-threads 4 --out results/trait_pcgc

summit --binary-method pcgc \
  --h2 results/trait_pcgc.binary.npz --out results/trait_pcgc_fit

The first command fits an ascertainment-aware population-probit risk model. For supplied individual population risks, replace --binary-covariates with --binary-risk-column RISK. --binary-covariate-variance optionally supplies the population variance of the covariate liability predictor.

Add --annot annotations.tsv at preparation for partitioned estimates. This table contains SNP followed by weight columns, with every genotype SNP present exactly once. Without it, SUMMIT fits one component over all SNPs.

Preparation writes .binary.npz; inference writes .binary.json. The output contains conditional and marginal component estimates and SEs, totals, risk diagnostics, and reference diagnostics. Existing output files are not overwritten. Use the same method at both stages. Changing risks, prevalence, or the sample selection requires preparing new moments.

Choose a method

Flag value Use Main condition
pcgc Individual risk weights in scores and the study reference Exogenous risk covariates and population genotype scaling
liability Constant population risk; scalar liability conversion Constant population risk
pcgc-basis Supplied basis and coefficients that exactly reproduce the risk sensitivity Exact sensitivity span
pcgc-inverse Inverse risk weighting with an unweighted genotype reference Inverse weights must have finite variance
pcgc-ld Transfer from an independent population LD reference Matched reference and valid risk/genotype factorization

pcgc-basis takes --binary-basis-columns and --binary-basis-coefficients. An exact span gives the same moments as pcgc; arbitrary risk binning is not an exact span. pcgc-ld additionally needs --binary-reference-geno. Its factorization can fail when risk weights and genotypes are dependent after ascertainment. Inverse weighting can be unstable with strong continuous risk factors. Neither method is automatically selected from the data.

Standard errors and interpretation

All five methods report SNP-block jackknife SEs for annotation components and totals by default. Binary inference uses 200 contiguous SNP blocks. Set --njack to another integer of at least two when needed for the SNP count and LD structure; it cannot exceed the number of SNPs. The HE/LDSC default is chromosome deletion (chr).

For example, to use 100 blocks:

summit --binary-method pcgc \
  --h2 results/trait_pcgc.binary.npz --njack 100 \
  --out results/trait_pcgc_100_blocks

The jackknife removes target SNP blocks and retains the full source reference. An integer block count divides the saved variant order into contiguous groups. Use genotypes ordered by chromosome and position, with blocks large enough to contain local LD. It holds the risk fit, prevalence, population scales, and reference probes fixed. It therefore excludes uncertainty in those quantities.

Conditional genetic variance is measured relative to liability variance after the covariate predictor, fixed at one. Marginal estimates divide by 1 + V_cov. Negative moment estimates are retained. With overlapping annotations, components are conditional contributions, not heritability of each annotation's SNP set in isolation.

The model assumes case-status sampling, exogenous risk covariates, and a compatible population genotype scale. It does not implement ancestry adjustment by ordinary covariate projection. See the PCGC methods for the estimating equations and assumptions of each method.

Computation and cross-trait work

Study-specific preparation uses two genotype traversals and SUMMIT's existing genotype readers and matrix operations. Exact basis contraction is performed before reference calculation. The external-LD method uses one study scoring traversal and two reference traversals.

--memory-gib budgets reference workspace; other process allocations need additional memory. Reducing --block-size changes tiling, while reducing --nvecs also changes numerical precision. Annotation count increases both sketch memory and computation.

Binary–binary and binary–quantitative covariance are available through summit.pcgc.cross.prepare_pair and fit_pair under an explicit marginal case-status sampling assumption. Shared controls require aligned sample IDs and a compatible selection design. This remains a research Python API; binary --rg and general cross-study binary artifacts are not exposed.

See the PCGC methods for the equations and API limits.

Clone this wiki locally