Skip to content

JLinAlg v0.3.6

Latest

Choose a tag to compare

@robbyjo robbyjo released this 17 Sep 18:27

JLinAlg 0.3.6

Differential, regional, multiplicity, and imputation inference

  • Added continuous empirical-Bayes and count-aware voom differential models,
    plus a separate median-ratio normalized negative-binomial estimator with
    cross-feature dispersion shrinkage and fixed-dispersion Wald inference.
  • Added coordinate/build-aware EWAS regions using signed Stouffer aggregation
    with an explicit exponential spatial-correlation matrix, coverage metadata,
    duplicate-coordinate rejection, and complete-region BH adjustment.
  • Added cross-fitted IHW-style bin weights with weighted BH and a separate
    hierarchy-respecting leaf-weighted Bonferroni procedure. Failed hypotheses
    remain in the prespecified family with p=1; hierarchical output controls
    FWER rather than being mislabeled as hierarchical FDR.
  • Added mixed-type chained-equations imputation with reproducible independent
    streams, predictive mean matching, binary/categorical probability draws,
    chain diagnostics, Rubin pooling, and Barnard-Rubin finite-sample degrees of
    freedom.
  • Added differential, ewas-regions, multiple-test,
    multiple-impute, and mi-pool commands, a worked vignette, a dedicated
    validation report, and website/citation navigation.
  • Frozen R 4.6.1/Bioconductor 3.23 fixtures validate effects against limma
    3.68.5, voom, edgeR 4.10.5, and DESeq2 1.52.0. Independent checks cover the
    NB mean ratio, spatial covariance, fold isolation, hierarchy accounting,
    reproducible imputations, and Rubin variance arithmetic.
  • Updated the generated scientific references page with a dedicated v0.3.6
    index linking the primary sources for differential, regional EWAS,
    multiple-testing, and multiple-imputation inference.

Latent confounders, batch adjustment, and CLI completion

  • Added feature-by-sample PCA, Bioconductor-style iteratively reweighted SVA,
    deterministic AutoSVA, official-source dense PEER VBFA, and parametric or
    nonparametric continuous-data ComBat APIs.
  • Added confounders and batch-adjust commands with sample-ID intersection,
    protected/null designs, factor/loadings/weight outputs, adjusted matrices,
    factor-selection paths, convergence metadata, and batch-confounding checks.
  • Pinned Bioconductor sva 3.60.0 and PMBio PEER 1.3 source revisions. Frozen
    R fixtures cover PCA, standard SVA, AutoSVA and ComBat; PEER retains an
    explicit native-runtime validation boundary on the Windows host.
  • Added a tall-matrix sample-Gram/backend path so global factor extraction never
    allocates a feature-by-feature covariance matrix.
  • Wired the preceding feature tranche through glm-predict, iv-regression,
    existing conditional-score, arima-regression --smooth, and the existing
    general --family probit path, with end-to-end CLI tests.
  • Added a practical latent-confounder vignette and source/validation audit.

Probit, predictions, IV, conditional scores, and scalable ARIMA

  • Added a stable binomial probit GLM through GlmFamilies.probit() and the
    --family probit / binomial-probit CLI aliases. Predictor-aware normal-tail
    likelihood, deviance, IRLS weights, and inference are checked against base R.
  • Added ModelPredictions, a shared GLM/GEE API for expected responses,
    scenario differences, risk ratios, and average marginal effects. It carries
    joint scenario covariance through analytic delta-method inference, separates
    population averages from average-covariate predictions, and reports mean
    uncertainty rather than future-outcome intervals.
  • Added individual-level linear IV/2SLS with QR projections, SVD rank and
    conditioning checks, homoskedastic/HC0/HC1/cluster covariance, structural-
    residual inference, first-stage partial R-squared/F diagnostics, and base-R
    fixtures. Weak-identification-robust intervals and exogeneity tests remain
    outside this API.
  • Added conditional-score to import complete version-1 conditional-GWAS score
    blocks, align exact alleles (and REF/ALT swaps only for Cox or models with an
    intercept), validate model contracts and matrix coverage, condition each
    cohort with a Schur complement, and pool independent cohorts. Unknown cross-
    block covariance is rejected; nonlinear
    cohort refits and the existing mr-estimate path remain distinct.
  • Replaced the historical ARIMA smoother's dense observation conditioning and
    1024-date cap with exact diffuse backward information recursions. ARIMA
    regression now exposes joint regression/dynamic parameter covariance,
    explicitly conditional forecasts, and delta-method forecasts that add fitted-
    parameter uncertainty while excluding innovation-variance estimation.
  • Added three focused vignettes and updated the GLM, conditional-GWAS, time-
    series, estimator-extension, README, TODO, and website feature guidance.

xWAS genetic architecture, molecular prediction and scores

  • Added ldsc for unpartitioned observed-scale heritability, genetic covariance
    and correlation with free intercepts, shared block-jackknife uncertainty,
    and reusable S/V matrix exports.
  • Added allele-aware twas / pwas summary association from raw-dosage model
    weights and reference LD, with optional full-rank joint model inference.
  • Added genomic-factor for a full-WLS single genetic factor and optional SNP
    GLS effects/heterogeneity conditional on fitted loadings.
  • Added score-train / score-apply for Gaussian penalized training, portable
    coefficients, genetic dosage alignment, and held-out score evaluation.
  • Added independent R fixtures, executable examples, CLI regression checks,
    and four web/Markdown vignettes. Recorded the remaining five analysis
    families in TODO.md. See validation and limitations.

Rare-variant model metadata and Raremetal2 assessment

  • Added version-1 quantitative score metadata, rare-score --trait-id and
    --trait-units, and rare-meta --model-metadata strict. Legacy quantitative
    imports retain logged assumptions; declared unsupported models/calibrations
    and incompatible cohort or score/covariance metadata fail before results.
  • Validate covariance column headers explicitly, including rejection of
    unsupported Raremetal2 compressed/allele-aware formats. Quantitative score
    arithmetic and historical supported RMW/rvtests layouts are unchanged.
  • Completed the Raremetal2 assessment and trait-model contract,
    with a deterministic R rare-case counterexample and separate implementation
    prerequisites for multiallelic groups, --useExact, and additional traits.

Conditional GWAS aggregate exports and local refits

  • Added --conditional-gwas-summary with required --score-genome-build to
    genotype Gaussian, binary logistic, Poisson and model-based Cox scans. It
    appends efficient scores, variances, counts, one-step estimates and explicit
    normal-only calibration status, and writes linked covariance/metadata files.
  • Added --condition-on for cohort-side null refitting with requested lead
    genotypes, without exporting participant records. --score-block-size bounds
    covariance blocks; unavailable cross-block covariance is never assumed zero.
  • Added streaming Cox genotype scans with offsets, reference-coded categorical
    covariates, right censoring/left truncation and Efron/Breslow ties. Related,
    repeated-subject and other unsupported covariance models fail explicitly.
  • Non-Gaussian genotype scans print and log the conditional-GWAS limitation;
    the ordinary output schema is unchanged unless the switch is used. Existing
    mr-estimate remains the Gaussian summary approximation. Rare-event tail
    calibration remains a separate extension; the summary-only importer above
    consumes only complete, compatible score/covariance blocks.
  • See the schema, CLI workflow and independent R validation.

Quantile covariance, missing ordinal SEM, and diffuse ARIMA inference

  • Added QuantileRegressionInference for exact quantile fits with supplied
    conditional-density sandwich covariance or a pooled iid Gaussian residual
    kernel at a caller-set bandwidth. Inference requires continuous responses,
    independent observations, regular densities, and a converged exact fit.
  • Added SemOrdinal.fitPairwiseMissing with explicit -1 missing categories,
    available-pair likelihood and case/cluster covariance. This requires MCAR or
    a justified pair-specific mechanism; general MAR ordinal FIML is not implied.
    Ordinal nonconvergence now suppresses covariance and all parameter inference.
  • Added diffuse ARIMA coefficient covariance and standard errors, including
    joint drift/dynamic cross-covariance and seasonal drift units. Unresolved
    information, transform-bound solutions, and nonconvergence suppress inference.
  • Independent R fixtures, analytic checks, and reproduction commands are in
    the validation report. Other estimator
    extensions and mathematical limits remain tracked in TODO.

GRM construction CLI and simpler getting started

  • Added grm --genotypes FILE --out matrix.tsv for VCF/BCF, BGEN, and dosage
    tables, with MAF/call-rate filters, bounded genotype blocks, backend selection,
    labeled CSV/TSV output compatible with --grm, and timestamped run logs.
  • Added GenomicRelationshipMatrix.fromSource, sharing the existing array
    estimator's normalization. Dense sample-matrix storage remains quadratic.
  • Simplified the README to requirements, build/run instructions, website links,
    numerical evidence, and license; detailed runtime setup is in the documentation.

Rare-variant score summaries and meta-analysis

  • Added rare-score for unrelated-sample quantitative-trait Gaussian scores,
    original-unit score covariance, and BGZF/tabix export from aligned VCF/BCF data.
  • Added rare-meta for independent cohort single-variant score pooling,
    equal/weighted burden, SKAT, and SKAT-O, with explicit test selection,
    test-specific files, optional cohort results, group files or sliding windows,
    cohort-count filters, allele alignment, and input/covariance checks.
  • Added primitive-array ScoreMetaAnalysis and summary-only SummarySetTests
    APIs. SKAT-O simulation now uses bounded memory, with an exact rank-one
    reduction in the summary API and explicit approximation/Monte Carlo limits.
  • Validated reference single-variant, burden, weighted burden, and SKAT results
    against RAREMETAL 4.15.1 fixtures; added independent R integration checks.
    See tutorial and
    validation/timing evidence for scope and limits.

CLI run timing

  • CLI logs now record UTC started and finished timestamps, numeric
    elapsed_ms, and a readable elapsed duration such as 2 h 3 min 4.125 s.
    Elapsed time uses a monotonic clock, independent of wall-clock corrections.
  • Meta-analysis, mediation, SuSiE, and colocalization flush their start log
    before computation and retain timing/status on handled failures.
  • Other executable subcommands now get a run log beside --out/--output,
    or a unique log in logs/ for console-only runs. They support --log and
    --no-log; help/version requests do not create logs.

Meta-analysis CLI and array workflows

  • Added meta-analysis and meta-regression commands for separate cohort
    TSV/CSV files (including gzip), matched by feature ID with external sorting.
  • Fixed/random effects support REML, DL, and PM, numeric moderators, coefficient
    inference controls, cohort direction strings, and --min-cohorts (default 1).
    Single-cohort pooling passes through explicitly; insufficient or rank-deficient
    regression rows are marked as unestimable.
  • Extended primitive-array batches to support missing cohorts, directions, and
    minimum counts; added primitive-array meta-regression. Heterogeneity stopping
    rules respect effect units and fail explicitly on iteration exhaustion.
  • Added independent R/metafor CLI fixtures, runnable six-cohort examples, and a
    reproducible synthetic allocation/time comparison with the study-list API.

Statistical correctness fixes

  • Stabilized the Lugannani–Rice kernel-tail fallback at and around the mean,
    including quantile inversion. It remains an explicitly labeled approximation;
    invalid numerical probabilities now fail instead of being clamped.
  • Made sparse REML and Paule–Mandel heterogeneity optimization independent of
    effect units and refined REML convergence near a flat numerical optimum.
  • Enforced the initial multidimensional quadrature node budget before allocating
    dense random-effect matrices or fitting, including node-count overflow.
  • Added scale-aware shape, symmetry, diagonal, and positive-semidefinite checks
    for conditional MR exposure covariances. Valid singular covariance matrices
    remain supported when conditional residual variances are positive.
  • Applied the ordinary profile-likelihood input checks to boundary profiles,
    with boundary confidence levels restricted to (0.5, 1).

Regression validation: 667 tests, 664 passed, zero failures, and three optional
native CHOLMOD tests skipped; website and Javadoc checks passed. Analytic
equal-variance meta-analysis and independent R integration checks were also run.