Skip to content

Releases: robbyjo/JLinAlg

JLinAlg v0.3.6

Choose a tag to compare

@robbyjo robbyjo released this 17 Sep 18:27

JLinAlg 0.3.6

Differential, regional, multiplicity, and imputation inference

  • Added continuous empirical-Bayes and count-aware voom differential models,
    plus a separate median-ratio normalized negative-binomial estimator with
    cross-feature dispersion shrinkage and fixed-dispersion Wald inference.
  • Added coordinate/build-aware EWAS regions using signed Stouffer aggregation
    with an explicit exponential spatial-correlation matrix, coverage metadata,
    duplicate-coordinate rejection, and complete-region BH adjustment.
  • Added cross-fitted IHW-style bin weights with weighted BH and a separate
    hierarchy-respecting leaf-weighted Bonferroni procedure. Failed hypotheses
    remain in the prespecified family with p=1; hierarchical output controls
    FWER rather than being mislabeled as hierarchical FDR.
  • Added mixed-type chained-equations imputation with reproducible independent
    streams, predictive mean matching, binary/categorical probability draws,
    chain diagnostics, Rubin pooling, and Barnard-Rubin finite-sample degrees of
    freedom.
  • Added differential, ewas-regions, multiple-test,
    multiple-impute, and mi-pool commands, a worked vignette, a dedicated
    validation report, and website/citation navigation.
  • Frozen R 4.6.1/Bioconductor 3.23 fixtures validate effects against limma
    3.68.5, voom, edgeR 4.10.5, and DESeq2 1.52.0. Independent checks cover the
    NB mean ratio, spatial covariance, fold isolation, hierarchy accounting,
    reproducible imputations, and Rubin variance arithmetic.
  • Updated the generated scientific references page with a dedicated v0.3.6
    index linking the primary sources for differential, regional EWAS,
    multiple-testing, and multiple-imputation inference.

Latent confounders, batch adjustment, and CLI completion

  • Added feature-by-sample PCA, Bioconductor-style iteratively reweighted SVA,
    deterministic AutoSVA, official-source dense PEER VBFA, and parametric or
    nonparametric continuous-data ComBat APIs.
  • Added confounders and batch-adjust commands with sample-ID intersection,
    protected/null designs, factor/loadings/weight outputs, adjusted matrices,
    factor-selection paths, convergence metadata, and batch-confounding checks.
  • Pinned Bioconductor sva 3.60.0 and PMBio PEER 1.3 source revisions. Frozen
    R fixtures cover PCA, standard SVA, AutoSVA and ComBat; PEER retains an
    explicit native-runtime validation boundary on the Windows host.
  • Added a tall-matrix sample-Gram/backend path so global factor extraction never
    allocates a feature-by-feature covariance matrix.
  • Wired the preceding feature tranche through glm-predict, iv-regression,
    existing conditional-score, arima-regression --smooth, and the existing
    general --family probit path, with end-to-end CLI tests.
  • Added a practical latent-confounder vignette and source/validation audit.

Probit, predictions, IV, conditional scores, and scalable ARIMA

  • Added a stable binomial probit GLM through GlmFamilies.probit() and the
    --family probit / binomial-probit CLI aliases. Predictor-aware normal-tail
    likelihood, deviance, IRLS weights, and inference are checked against base R.
  • Added ModelPredictions, a shared GLM/GEE API for expected responses,
    scenario differences, risk ratios, and average marginal effects. It carries
    joint scenario covariance through analytic delta-method inference, separates
    population averages from average-covariate predictions, and reports mean
    uncertainty rather than future-outcome intervals.
  • Added individual-level linear IV/2SLS with QR projections, SVD rank and
    conditioning checks, homoskedastic/HC0/HC1/cluster covariance, structural-
    residual inference, first-stage partial R-squared/F diagnostics, and base-R
    fixtures. Weak-identification-robust intervals and exogeneity tests remain
    outside this API.
  • Added conditional-score to import complete version-1 conditional-GWAS score
    blocks, align exact alleles (and REF/ALT swaps only for Cox or models with an
    intercept), validate model contracts and matrix coverage, condition each
    cohort with a Schur complement, and pool independent cohorts. Unknown cross-
    block covariance is rejected; nonlinear
    cohort refits and the existing mr-estimate path remain distinct.
  • Replaced the historical ARIMA smoother's dense observation conditioning and
    1024-date cap with exact diffuse backward information recursions. ARIMA
    regression now exposes joint regression/dynamic parameter covariance,
    explicitly conditional forecasts, and delta-method forecasts that add fitted-
    parameter uncertainty while excluding innovation-variance estimation.
  • Added three focused vignettes and updated the GLM, conditional-GWAS, time-
    series, estimator-extension, README, TODO, and website feature guidance.

xWAS genetic architecture, molecular prediction and scores

  • Added ldsc for unpartitioned observed-scale heritability, genetic covariance
    and correlation with free intercepts, shared block-jackknife uncertainty,
    and reusable S/V matrix exports.
  • Added allele-aware twas / pwas summary association from raw-dosage model
    weights and reference LD, with optional full-rank joint model inference.
  • Added genomic-factor for a full-WLS single genetic factor and optional SNP
    GLS effects/heterogeneity conditional on fitted loadings.
  • Added score-train / score-apply for Gaussian penalized training, portable
    coefficients, genetic dosage alignment, and held-out score evaluation.
  • Added independent R fixtures, executable examples, CLI regression checks,
    and four web/Markdown vignettes. Recorded the remaining five analysis
    families in TODO.md. See validation and limitations.

Rare-variant model metadata and Raremetal2 assessment

  • Added version-1 quantitative score metadata, rare-score --trait-id and
    --trait-units, and rare-meta --model-metadata strict. Legacy quantitative
    imports retain logged assumptions; declared unsupported models/calibrations
    and incompatible cohort or score/covariance metadata fail before results.
  • Validate covariance column headers explicitly, including rejection of
    unsupported Raremetal2 compressed/allele-aware formats. Quantitative score
    arithmetic and historical supported RMW/rvtests layouts are unchanged.
  • Completed the Raremetal2 assessment and trait-model contract,
    with a deterministic R rare-case counterexample and separate implementation
    prerequisites for multiallelic groups, --useExact, and additional traits.

Conditional GWAS aggregate exports and local refits

  • Added --conditional-gwas-summary with required --score-genome-build to
    genotype Gaussian, binary logistic, Poisson and model-based Cox scans. It
    appends efficient scores, variances, counts, one-step estimates and explicit
    normal-only calibration status, and writes linked covariance/metadata files.
  • Added --condition-on for cohort-side null refitting with requested lead
    genotypes, without exporting participant records. --score-block-size bounds
    covariance blocks; unavailable cross-block covariance is never assumed zero.
  • Added streaming Cox genotype scans with offsets, reference-coded categorical
    covariates, right censoring/left truncation and Efron/Breslow ties. Related,
    repeated-subject and other unsupported covariance models fail explicitly.
  • Non-Gaussian genotype scans print and log the conditional-GWAS limitation;
    the ordinary output schema is unchanged unless the switch is used. Existing
    mr-estimate remains the Gaussian summary approximation. Rare-event tail
    calibration remains a separate extension; the summary-only importer above
    consumes only complete, compatible score/covariance blocks.
  • See the schema, CLI workflow and independent R validation.

Quantile covariance, missing ordinal SEM, and diffuse ARIMA inference

  • Added QuantileRegressionInference for exact quantile fits with supplied
    conditional-density sandwich covariance or a pooled iid Gaussian residual
    kernel at a caller-set bandwidth. Inference requires continuous responses,
    independent observations, regular densities, and a converged exact fit.
  • Added SemOrdinal.fitPairwiseMissing with explicit -1 missing categories,
    available-pair likelihood and case/cluster covariance. This requires MCAR or
    a justified pair-specific mechanism; general MAR ordinal FIML is not implied.
    Ordinal nonconvergence now suppresses covariance and all parameter inference.
  • Added diffuse ARIMA coefficient covariance and standard errors, including
    joint drift/dynamic cross-covariance and seasonal drift units. Unresolved
    information, transform-bound solutions, and nonconvergence suppress inference.
  • Independent R fixtures, analytic checks, and reproduction commands are in
    the validation report. Other estimator
    extensions and mathematical limits remain tracked in TODO.

GRM construction CLI and simpler getting started

  • Added grm --genotypes FILE --out matrix.tsv for VCF/BCF, BGEN, and dosage
    tables, with MAF/call-rate filters, bounded genotype blocks, backend selection,
    labeled CSV/TSV output compatible with --grm, and timestamped run logs.
  • Added GenomicRelationshipMatrix.fromSource, sharing the existing array
    estimator's normalization. Dense sample-matrix storage remains quadratic.
  • Simplified the README to requirements, build/run instructions, website links,
    numerical evidence, and license; detailed runtime setup is in the documentation.

Rare-variant score summaries and meta-analysis

  • Added rare-score for unrelated-sample quantitative-trait Gaussian scores,
    original-unit score covariance, and BGZF/...
Read more

JLinAlg v0.3.5

Choose a tag to compare

@github-actions github-actions released this 11 Sep 21:41

JLinAlg 0.3.5

Executable mediation and fine-mapping workflows

  • The executable JAR now exposes Gaussian mediation workflows for ordinary,
    grouped-REML, and pedigree-REML models, with common complete-case alignment
    and explicit component-model convergence output.
  • New susie and coloc commands run summary-statistic fine mapping from
    labeled LD matrices and multi-signal colocalization from the resulting
    effect tables.
  • Progressive CLI tutorials now cover association, Mendelian randomization,
    mediation, SuSiE, and colocalization from input files through interpretation.

Association preflight and option validation

  • --max-maf and --max-mac complement the existing lower-bound variant
    filters and are applied after sample alignment.
  • Phenotype-only --dry-run and --explain now compile and validate the
    selected route before fitting. Genotype transforms, unsupported family/link
    combinations, and the common --explains typo fail with specific guidance.

Pedigree identity and hierarchy

  • Family columns now disambiguate duplicate member IDs without partitioning
    the ancestry graph. Globally unique raw observation and parent IDs resolve
    automatically across family labels, while exact family:individual IDs
    remain accepted and ambiguous raw IDs fail with a qualification request.
  • Parent-child links recursively define the full pedigree hierarchy, including
    shared ancestors and consanguineous matings. A direct R pedigreemm 0.3.5
    fixture now verifies the additive relationship matrix, inbreeding
    coefficients, and Henderson sparse inverse through a cousin-mating example.
  • Runs with zero aligned source-pedigree matches now fail rather than silently
    constructing an all-singleton model. Logs and manifests distinguish matched
    source observations, automatically resolved aliases, and true singletons.

Release downloads: jlinalg-0.3.5.jar is the self-contained executable;
JLinAlg-0.3.5-library.jar is the thin library. Sources and Javadoc JARs and
SHA256SUMS.txt are also included.

JLinAlg v0.3.4

Choose a tag to compare

@github-actions github-actions released this 11 Sep 18:25

JLinAlg 0.3.4

High-core adaptive omics scheduling

  • Automatic feature blocks are sized from both current JVM heap headroom and
    the requested thread count.
  • When memory permits, each block queues at least two complete 256-feature
    chunks per scan worker, giving work stealing enough queued work to absorb
    uneven feature-fit times without leaving cores idle at block barriers.
  • Under tighter memory, JLinAlg selects the largest complete worker wave that
    fits or safely caps worker capacity. An explicit --block-size remains
    authoritative.

Pedigree singleton handling

  • Phenotype individuals absent from the pedigree file are retained as
    unrelated, noninbred singleton founders instead of stopping the analysis.
  • Repeated observations with the same missing pedigree ID remain grouped in
    one singleton family. The console, log, and manifest report singleton
    families and observations.

CLI documentation and transform semantics

  • Added a comprehensive CLI-only guide and copy-pasteable command-line examples
    to the model vignettes, including grouped REML, pedigrees, GRMs, GLMMs,
    genotype scans, Cox models, and specialized subcommands.
  • Omics row transforms intentionally run after phenotype complete-case
    filtering, so winsor_mad and other row-statistic transforms use the final
    analysis sample set.

Release downloads: jlinalg-0.3.4.jar is the self-contained executable;
JLinAlg-0.3.4-library.jar is the thin library. Sources and Javadoc JARs and
SHA256SUMS.txt are also included.

JLinAlg v0.3.3

Choose a tag to compare

@github-actions github-actions released this 10 Sep 22:54

JLinAlg 0.3.3

Complete-case sample handling

  • Phenotype rows with missing values in any model column are omitted by
    default, and the corresponding samples are removed from streamed omics
    inputs before fitting.
  • Console, log, and manifest output distinguish aligned samples from the final
    analysis sample count and report how many incomplete phenotype rows were
    omitted.

Exact streamed mixed-model scans

  • Numeric Gaussian omics scans now refit variance components by REML for every
    feature through the lmer-like pathway instead of using P3D or EMMAX.
  • Non-Gaussian numeric omics scans now refit a first-order Laplace GLMM for
    every feature through the glmer-like pathway instead of PQL.
  • Pedigree-enabled scans use direct sparse Henderson relationship precision
    with ancestry-graph inbreeding, matching the intended pedigreemm model
    structure without materializing a dense relationship matrix.
  • The selected mixed-fit strategy is recorded in the console preflight, run
    log, and manifest. Genotype LMM scans retain their documented null-model
    P3D pathway.

Delimited output and compact schemas

  • A .csv output path now produces correctly quoted comma-separated output;
    .tsv and other existing outputs remain tab-separated.
  • Generic numeric omics results omit genotype-only allele, frequency, call,
    imputation, Hardy-Weinberg, and filter columns.
  • Per-row error fields are consolidated as failure_reason. Invariant
    omics_type, statistic_type, df_method, partial_r2_method, and
    output_format metadata now live in the run log and manifest instead of
    being repeated in every result row.

Release downloads: jlinalg-0.3.3.jar is the self-contained executable;
JLinAlg-0.3.3-library.jar is the thin library. Sources and Javadoc JARs and
SHA256SUMS.txt are also included.