Releases: robbyjo/JLinAlg
Releases · robbyjo/JLinAlg
Release list
JLinAlg v0.3.6
JLinAlg 0.3.6
Differential, regional, multiplicity, and imputation inference
- Added continuous empirical-Bayes and count-aware voom differential models,
plus a separate median-ratio normalized negative-binomial estimator with
cross-feature dispersion shrinkage and fixed-dispersion Wald inference. - Added coordinate/build-aware EWAS regions using signed Stouffer aggregation
with an explicit exponential spatial-correlation matrix, coverage metadata,
duplicate-coordinate rejection, and complete-region BH adjustment. - Added cross-fitted IHW-style bin weights with weighted BH and a separate
hierarchy-respecting leaf-weighted Bonferroni procedure. Failed hypotheses
remain in the prespecified family with p=1; hierarchical output controls
FWER rather than being mislabeled as hierarchical FDR. - Added mixed-type chained-equations imputation with reproducible independent
streams, predictive mean matching, binary/categorical probability draws,
chain diagnostics, Rubin pooling, and Barnard-Rubin finite-sample degrees of
freedom. - Added
differential,ewas-regions,multiple-test,
multiple-impute, andmi-poolcommands, a worked vignette, a dedicated
validation report, and website/citation navigation. - Frozen R 4.6.1/Bioconductor 3.23 fixtures validate effects against limma
3.68.5, voom, edgeR 4.10.5, and DESeq2 1.52.0. Independent checks cover the
NB mean ratio, spatial covariance, fold isolation, hierarchy accounting,
reproducible imputations, and Rubin variance arithmetic. - Updated the generated scientific references page with a dedicated v0.3.6
index linking the primary sources for differential, regional EWAS,
multiple-testing, and multiple-imputation inference.
Latent confounders, batch adjustment, and CLI completion
- Added feature-by-sample PCA, Bioconductor-style iteratively reweighted SVA,
deterministic AutoSVA, official-source dense PEER VBFA, and parametric or
nonparametric continuous-data ComBat APIs. - Added
confoundersandbatch-adjustcommands with sample-ID intersection,
protected/null designs, factor/loadings/weight outputs, adjusted matrices,
factor-selection paths, convergence metadata, and batch-confounding checks. - Pinned Bioconductor
sva3.60.0 and PMBio PEER 1.3 source revisions. Frozen
R fixtures cover PCA, standard SVA, AutoSVA and ComBat; PEER retains an
explicit native-runtime validation boundary on the Windows host. - Added a tall-matrix sample-Gram/backend path so global factor extraction never
allocates a feature-by-feature covariance matrix. - Wired the preceding feature tranche through
glm-predict,iv-regression,
existingconditional-score,arima-regression --smooth, and the existing
general--family probitpath, with end-to-end CLI tests. - Added a practical latent-confounder vignette and source/validation audit.
Probit, predictions, IV, conditional scores, and scalable ARIMA
- Added a stable binomial probit GLM through
GlmFamilies.probit()and the
--family probit/binomial-probitCLI aliases. Predictor-aware normal-tail
likelihood, deviance, IRLS weights, and inference are checked against base R. - Added
ModelPredictions, a shared GLM/GEE API for expected responses,
scenario differences, risk ratios, and average marginal effects. It carries
joint scenario covariance through analytic delta-method inference, separates
population averages from average-covariate predictions, and reports mean
uncertainty rather than future-outcome intervals. - Added individual-level linear IV/2SLS with QR projections, SVD rank and
conditioning checks, homoskedastic/HC0/HC1/cluster covariance, structural-
residual inference, first-stage partial R-squared/F diagnostics, and base-R
fixtures. Weak-identification-robust intervals and exogeneity tests remain
outside this API. - Added
conditional-scoreto import complete version-1 conditional-GWAS score
blocks, align exact alleles (and REF/ALT swaps only for Cox or models with an
intercept), validate model contracts and matrix coverage, condition each
cohort with a Schur complement, and pool independent cohorts. Unknown cross-
block covariance is rejected; nonlinear
cohort refits and the existingmr-estimatepath remain distinct. - Replaced the historical ARIMA smoother's dense observation conditioning and
1024-date cap with exact diffuse backward information recursions. ARIMA
regression now exposes joint regression/dynamic parameter covariance,
explicitly conditional forecasts, and delta-method forecasts that add fitted-
parameter uncertainty while excluding innovation-variance estimation. - Added three focused vignettes and updated the GLM, conditional-GWAS, time-
series, estimator-extension, README, TODO, and website feature guidance.
xWAS genetic architecture, molecular prediction and scores
- Added
ldscfor unpartitioned observed-scale heritability, genetic covariance
and correlation with free intercepts, shared block-jackknife uncertainty,
and reusable S/V matrix exports. - Added allele-aware
twas/pwassummary association from raw-dosage model
weights and reference LD, with optional full-rank joint model inference. - Added
genomic-factorfor a full-WLS single genetic factor and optional SNP
GLS effects/heterogeneity conditional on fitted loadings. - Added
score-train/score-applyfor Gaussian penalized training, portable
coefficients, genetic dosage alignment, and held-out score evaluation. - Added independent R fixtures, executable examples, CLI regression checks,
and four web/Markdown vignettes. Recorded the remaining five analysis
families in TODO.md. See validation and limitations.
Rare-variant model metadata and Raremetal2 assessment
- Added version-1 quantitative score metadata,
rare-score --trait-idand
--trait-units, andrare-meta --model-metadata strict. Legacy quantitative
imports retain logged assumptions; declared unsupported models/calibrations
and incompatible cohort or score/covariance metadata fail before results. - Validate covariance column headers explicitly, including rejection of
unsupported Raremetal2 compressed/allele-aware formats. Quantitative score
arithmetic and historical supported RMW/rvtests layouts are unchanged. - Completed the Raremetal2 assessment and trait-model contract,
with a deterministic R rare-case counterexample and separate implementation
prerequisites for multiallelic groups,--useExact, and additional traits.
Conditional GWAS aggregate exports and local refits
- Added
--conditional-gwas-summarywith required--score-genome-buildto
genotype Gaussian, binary logistic, Poisson and model-based Cox scans. It
appends efficient scores, variances, counts, one-step estimates and explicit
normal-only calibration status, and writes linked covariance/metadata files. - Added
--condition-onfor cohort-side null refitting with requested lead
genotypes, without exporting participant records.--score-block-sizebounds
covariance blocks; unavailable cross-block covariance is never assumed zero. - Added streaming Cox genotype scans with offsets, reference-coded categorical
covariates, right censoring/left truncation and Efron/Breslow ties. Related,
repeated-subject and other unsupported covariance models fail explicitly. - Non-Gaussian genotype scans print and log the conditional-GWAS limitation;
the ordinary output schema is unchanged unless the switch is used. Existing
mr-estimateremains the Gaussian summary approximation. Rare-event tail
calibration remains a separate extension; the summary-only importer above
consumes only complete, compatible score/covariance blocks. - See the schema, CLI workflow and independent R validation.
Quantile covariance, missing ordinal SEM, and diffuse ARIMA inference
- Added
QuantileRegressionInferencefor exact quantile fits with supplied
conditional-density sandwich covariance or a pooled iid Gaussian residual
kernel at a caller-set bandwidth. Inference requires continuous responses,
independent observations, regular densities, and a converged exact fit. - Added
SemOrdinal.fitPairwiseMissingwith explicit-1missing categories,
available-pair likelihood and case/cluster covariance. This requires MCAR or
a justified pair-specific mechanism; general MAR ordinal FIML is not implied.
Ordinal nonconvergence now suppresses covariance and all parameter inference. - Added diffuse ARIMA coefficient covariance and standard errors, including
joint drift/dynamic cross-covariance and seasonal drift units. Unresolved
information, transform-bound solutions, and nonconvergence suppress inference. - Independent R fixtures, analytic checks, and reproduction commands are in
the validation report. Other estimator
extensions and mathematical limits remain tracked in TODO.
GRM construction CLI and simpler getting started
- Added
grm --genotypes FILE --out matrix.tsvfor VCF/BCF, BGEN, and dosage
tables, with MAF/call-rate filters, bounded genotype blocks, backend selection,
labeled CSV/TSV output compatible with--grm, and timestamped run logs. - Added
GenomicRelationshipMatrix.fromSource, sharing the existing array
estimator's normalization. Dense sample-matrix storage remains quadratic. - Simplified the README to requirements, build/run instructions, website links,
numerical evidence, and license; detailed runtime setup is in the documentation.
Rare-variant score summaries and meta-analysis
- Added
rare-scorefor unrelated-sample quantitative-trait Gaussian scores,
original-unit score covariance, and BGZF/...
JLinAlg v0.3.5
JLinAlg 0.3.5
Executable mediation and fine-mapping workflows
- The executable JAR now exposes Gaussian
mediationworkflows for ordinary,
grouped-REML, and pedigree-REML models, with common complete-case alignment
and explicit component-model convergence output. - New
susieandcoloccommands run summary-statistic fine mapping from
labeled LD matrices and multi-signal colocalization from the resulting
effect tables. - Progressive CLI tutorials now cover association, Mendelian randomization,
mediation, SuSiE, and colocalization from input files through interpretation.
Association preflight and option validation
--max-mafand--max-maccomplement the existing lower-bound variant
filters and are applied after sample alignment.- Phenotype-only
--dry-runand--explainnow compile and validate the
selected route before fitting. Genotype transforms, unsupported family/link
combinations, and the common--explainstypo fail with specific guidance.
Pedigree identity and hierarchy
- Family columns now disambiguate duplicate member IDs without partitioning
the ancestry graph. Globally unique raw observation and parent IDs resolve
automatically across family labels, while exactfamily:individualIDs
remain accepted and ambiguous raw IDs fail with a qualification request. - Parent-child links recursively define the full pedigree hierarchy, including
shared ancestors and consanguineous matings. A direct Rpedigreemm0.3.5
fixture now verifies the additive relationship matrix, inbreeding
coefficients, and Henderson sparse inverse through a cousin-mating example. - Runs with zero aligned source-pedigree matches now fail rather than silently
constructing an all-singleton model. Logs and manifests distinguish matched
source observations, automatically resolved aliases, and true singletons.
Release downloads: jlinalg-0.3.5.jar is the self-contained executable;
JLinAlg-0.3.5-library.jar is the thin library. Sources and Javadoc JARs and
SHA256SUMS.txt are also included.
JLinAlg v0.3.4
JLinAlg 0.3.4
High-core adaptive omics scheduling
- Automatic feature blocks are sized from both current JVM heap headroom and
the requested thread count. - When memory permits, each block queues at least two complete 256-feature
chunks per scan worker, giving work stealing enough queued work to absorb
uneven feature-fit times without leaving cores idle at block barriers. - Under tighter memory, JLinAlg selects the largest complete worker wave that
fits or safely caps worker capacity. An explicit--block-sizeremains
authoritative.
Pedigree singleton handling
- Phenotype individuals absent from the pedigree file are retained as
unrelated, noninbred singleton founders instead of stopping the analysis. - Repeated observations with the same missing pedigree ID remain grouped in
one singleton family. The console, log, and manifest report singleton
families and observations.
CLI documentation and transform semantics
- Added a comprehensive CLI-only guide and copy-pasteable command-line examples
to the model vignettes, including grouped REML, pedigrees, GRMs, GLMMs,
genotype scans, Cox models, and specialized subcommands. - Omics row transforms intentionally run after phenotype complete-case
filtering, sowinsor_madand other row-statistic transforms use the final
analysis sample set.
Release downloads: jlinalg-0.3.4.jar is the self-contained executable;
JLinAlg-0.3.4-library.jar is the thin library. Sources and Javadoc JARs and
SHA256SUMS.txt are also included.
JLinAlg v0.3.3
JLinAlg 0.3.3
Complete-case sample handling
- Phenotype rows with missing values in any model column are omitted by
default, and the corresponding samples are removed from streamed omics
inputs before fitting. - Console, log, and manifest output distinguish aligned samples from the final
analysis sample count and report how many incomplete phenotype rows were
omitted.
Exact streamed mixed-model scans
- Numeric Gaussian omics scans now refit variance components by REML for every
feature through thelmer-like pathway instead of using P3D or EMMAX. - Non-Gaussian numeric omics scans now refit a first-order Laplace GLMM for
every feature through theglmer-like pathway instead of PQL. - Pedigree-enabled scans use direct sparse Henderson relationship precision
with ancestry-graph inbreeding, matching the intendedpedigreemmmodel
structure without materializing a dense relationship matrix. - The selected mixed-fit strategy is recorded in the console preflight, run
log, and manifest. Genotype LMM scans retain their documented null-model
P3D pathway.
Delimited output and compact schemas
- A
.csvoutput path now produces correctly quoted comma-separated output;
.tsvand other existing outputs remain tab-separated. - Generic numeric omics results omit genotype-only allele, frequency, call,
imputation, Hardy-Weinberg, and filter columns. - Per-row error fields are consolidated as
failure_reason. Invariant
omics_type,statistic_type,df_method,partial_r2_method, and
output_formatmetadata now live in the run log and manifest instead of
being repeated in every result row.
Release downloads: jlinalg-0.3.3.jar is the self-contained executable;
JLinAlg-0.3.3-library.jar is the thin library. Sources and Javadoc JARs and
SHA256SUMS.txt are also included.