Repository navigation
Generalized GxE PCGC
PCGC estimates the covariance of additive and context-dependent SNP effects
for a binary trait, accounting for case–control sampling. A context basis
[1, E1, E2] gives six parameters per annotation: additive variance, two
interaction variances, and their three covariances. Continuous exposures,
categorical variables encoded numerically, and supplied nonlinear basis
columns use the same interface.
For population-scaled genotypes X and context vector φ, the genetic model is
Each Ω is a symmetric effect-covariance matrix. Annotation weights must be nonnegative; annotation sets may overlap. Overlapping annotation coefficients describe their joint model and require identifiable kernels.
The binary outcome follows a liability threshold. Specify either unit total liability variance conditional on covariates and contexts, or known positive conditional liability SDs on a common scale. Disease risks alone cannot identify an arbitrary liability-variance function. Under the unit-variance model, changes in genetic variance are accompanied by changes in residual variance that keep their sum equal to one.
PCGC uses the first-order relationship between distinct-person binary covariance and liability covariance. It requires correctly specified disease risks, case-status sampling, exogenous contexts, and compatible population genotype scaling. Higher-order correlations, fitted nuisance parameters, and a finite randomized reference can affect finite-sample estimates.
Use the genotype, population-scale, and sample-table formats described in
Binary traits and PCGC. The sample table
must include FID IID Y E1 E2, with Y coded 0/1, plus any risk covariates.
Context columns are used as supplied; SUMMIT adds an intercept. Choose the
exposure centers, scales, and categorical reference levels before preparation.
Provide population prevalence with --binary-prevalence. A fitted probit
risk model includes the context columns and --binary-covariates.
Alternatively, --binary-risk-column RISK supplies individual population
risks. Ordinary logistic GWAS beta/SE files cannot replace these inputs.
Use --binary-unit-liability for the unit conditional variance model.
For known heterogeneous SDs, use --binary-risk-column RISK together with
--binary-liability-sd-column SD.
--make-binary-sumstats computes both the PCGC directional LD scores and
the trait score moments:
summit --binary-method pcgc \
--make-binary-sumstats people.tsv --geno study.bed \
--binary-scale population_scale.tsv --binary-prevalence 0.05 \
--binary-context-columns E1,E2 --binary-unit-liability \
--binary-covariates AGE,PC1,PC2 \
--binary-genotype-covariates PC1,PC2 \
--binary-sampling-partners 128 --binary-architecture-probes 32 \
--nvecs 256 --memory-gib 8 --num-threads 2 \
--out results/disease_gxePreparation writes two NumPy archives:
| File | Contents |
|---|---|
results/disease_gxe.binary.ldscores.npz |
PCGC directional reference rows, annotations, and any reference uncertainty arrays |
results/disease_gxe.binary.sumstats.npz |
Per-SNP context-pair trait score products, population moments, and any sampling or SNP-effect covariance summaries |
Both the reference rows (ldscores) and trait score products (rhs_rows) have
same-person terms removed. The NPZ format retains their multiple context and
annotation dimensions. The files record variant order, alleles, context names,
and the sample, risk, genotype-scale, and liability-scale definitions.
Add --annot annotations.tsv for partitioned estimates. Changes to risks,
prevalence, genotype scaling, context coding, or sample selection require new
preparation. Choose a new output prefix for each preparation.
summit --binary-method pcgc \
--h2 results/disease_gxe.binary.sumstats.npz \
--ldscores results/disease_gxe.binary.ldscores.npz \
--njack 100 --out results/disease_gxe_fitUse the same PCGC method as in preparation. Fitting writes
results/disease_gxe_fit.binary.json and requires no genotype input. You can
reuse the two files with a different --njack value and a new output prefix.
The summary records the expected reference identity. Fitting checks the stored sample, risk, scaling, context, annotation, and variant definitions, including alleles and any genome-build label. It also checks the reference array contents and both files' checksums. A mismatch stops the fit.
Use the reference written with the summary. The checks require that particular reference realization, including any uncertainty arrays, because the saved sampling calculations can depend on it. Matching SNP names or dimensions alone is insufficient. You can move or rename the files; matching uses their contents. Standard PCGC's risk-weighted reference is generally specific to the disease and analyzed sample.
Existing combined .binary.npz files remain accepted through --h2 alone.
To prepare that format, add --binary-output-format combined. Separate files
are the default for all binary methods. The input guide
compares the formats across workflows.
Genotype-PC adjustment removes the supplied PC directions from X before context and risk weighting. These PCs also enter fitted risks. The binary response and weighted context features receive no additional least-squares covariate projection. Add PC-by-context columns to the risk covariates when the mean model requires them; SUMMIT does not create those columns.
Use pcgc for the primary analysis. Its study-specific directional LD scores
use features F_q = diag(d/s) diag(phi_q) X, where d is the PCGC risk
sensitivity and s is the supplied conditional liability SD. The sensitivity
depends on population risk and case–control sampling, as defined in
PCGC methods. Trait scores
use the same features, and preparation removes same-person terms from both
the trait and reference moments. Ordinary quantitative-trait directional LD
scores generally cannot be reused for this fit.
| Method | Reference and weighting | Additional inputs or assumptions |
|---|---|---|
pcgc |
Study-specific risk-weighted context reference | Population risks and genotype scale |
pcgc-basis |
Exact contraction of a supplied risk-sensitivity basis |
--binary-basis-columns and --binary-basis-coefficients must reproduce sensitivity exactly |
pcgc-inverse |
Inverse risk weighting with an unweighted context reference | Inverse weights need finite variance; strong risk differences can cause instability |
pcgc-ld |
Independent population LD reference |
--binary-reference-geno and --binary-ld-factorization; requires risk/context/genotype factorization |
The risk-sensitivity basis and the genetic context basis have separate roles.
An exact sensitivity basis reproduces standard PCGC. External LD is an
approximation whose validity depends on the stated factorization. With
genotype-PC adjustment, external LD also needs a matching PC table through
--binary-reference-covariates.
omega contains one matrix per annotation; omega_total is their sum.
For contexts x and z, x.T @ omega_total @ z describes genetic covariance
under the standardized kernel model. Estimates remain signed, and fitted
matrices can have negative eigenvalues. Inspect the conditioning and
eigenvalue diagnostics before interpreting individual components.
population_heritability is total SNP heritability, including G×E and its
covariances with additive effects. Its numerator averages the genetic kernel
diagonal over the population, using the actual adjusted genotype diagonals.
Its denominator is the marginal liability variance. When conditional genotype
variances are one, the numerator reduces to
Population moments use inverse ascertainment weights by default. The Python API also accepts supplied population moments. Context-specific heritability requires the corresponding context-specific liability variance.
--binary-sampling-partners enables participant-sampling covariance, including
a same-study fitted risk model, estimated population moments, and randomized
reference error. --binary-architecture-probes additionally models variation
in Gaussian SNP effects. Outputs include component SEs, their full covariance,
normal intervals, and score confidence sets. A score set can be unbounded.
The counts in the example are starting values; assess simulation precision
for the study size and model being analyzed.
These intervals remain experimental. Simulations have shown undercoverage for some rare-disease inverse-weighted and overlapping-LD models. Supplied prevalence, genotype scaling, liability SDs, genotype adjustment, and supplied risks are treated as fixed. Uncertainty from estimating those external inputs is not included.
With zero sampling partners, inference uses the SNP-block jackknife. It holds
risks, population moments, and reference probes fixed. When both calculations
are available, snp_block_standard_errors preserves the jackknife result and
standard_errors reports the sampling result.
Study-specific preparation reads genotypes twice. External-LD preparation uses
one study pass and two reference passes, with a second study pass when sampling
covariance is requested. Computation uses streamed variant probes and avoids
participant-by-participant matrices. More contexts and annotations increase
memory and work; --memory-gib budgets workspace, so allow additional process
memory when sizing jobs.
summit.context.binary exposes prepare_gxe_moments, prepare_gxe_external,
fit_gxe, evaluate_contexts, and plan_gxe_reference. File-backed
preparation is available through
summit.pcgc.gxe_io.prepare_gxe_from_source. Use
summit.pcgc.split_io.write_split_artifact and load_split_artifact to save
and reload the separate files. Quantitative G×E summaries and binary contextual
summaries have separate formats. Cross-trait contextual
PCGC is not currently exposed.
Start here
Analyses
- LD scores
- h² and rg
- Batch analyses
- G×E models
- Multiple environments
- Quantitative epistasis
- Cross-trait response models
- Binary traits and PCGC
- Generalized G×E PCGC
- Polygenic scores
- Mixture priors
Results and reference