Skip to content

Releases: Veronica0206/Gtheory4LLM

Gtheory4LLM 0.2.0

Choose a tag to compare

@Veronica0206 Veronica0206 released this 23 Sep 13:07
Immutable release. Only release title and notes can be modified.

Generalizability Theory for LLM subjective tasks: exact balanced Gaussian
G-theory models, and a bounded experimental Laplace engine for small discrete
G-theory models used in LLM annotation and evaluation studies.

This is a research beta. Passing the package's numerical acceptance checks
does not establish parameter recovery, interval coverage, Laplace approximation
quality, or a scientifically sufficient number of evaluators. Read
docs/LIMITATIONS.md
before quoting a coefficient, and
docs/VALIDATION_SCOPE.md
for what has actually been validated.

Install

install.packages("OpenMx")
install.packages("Gtheory4LLM_0.2.0.tar.gz", repos = NULL, type = "source")

R 4.5.0 or later is required. Nothing in this package's own R code needs it; the
floor is inherited from OpenMx 2.22.11, which calls a C entry point introduced in
R 4.5.0 while declaring only R (>= 3.5.0) itself.

Verifying these files

The archive is the exact candidate the readiness workflow built from source
commit a4dd59c and checked with R-devel, adopted unchanged, never rebuilt.
manifest.json records that source commit, the archive's origin and the R
versions that built and checked it, together with the sizes and SHA-256 digests:

Gtheory4LLM_0.2.0.tar.gz   295761 bytes
  910286884cbbac91ec9ea466ba1997e0a0c9a9b0d8779852d5c158ed6bc5eb78
Gtheory4LLM-manual.pdf     196344 bytes
  4aa081ad81e8a5d363b50d658ebeecd877f823a0b3d35c22cc002de75c1fbb9c

From a checkout of the v0.2.0 tag, this re-derives the whole chain: archive
integrity, correspondence to the recorded source commit, release identity, then
a fresh-library install and smoke test:

python3 scripts/check_committed_artifact.py --check-release-identity --release-tag v0.2.0

What changed since 0.1.0

Interior-block standard errors say what they condition on: a source whose
fitted covariance is entirely zero is held at zero, and a singular but nonzero
source is held fixed at its fitted covariance. logLik(), AIC() and BIC()
refuse a numerically rejected fit, as reliability and D studies already did,
and the REML criteria are one documented contract. Dropping the data from a fit
drops every copy of the observations, the retained model's metadata and the
recorded call included. Gaussian facet levels are matched by value, and
outcomes are centred before the factorial contrasts, so large identifiers and
large common offsets no longer disturb accepted fits. Staged numerical
diagnostics summarise what a fit passed through. The README links that CRAN's
incoming check rejected in 0.0.6 are repository URLs, guarded by a test. The
full change list is in
NEWS.md.

What is not in it

Unbalanced inference; discrete standard errors or intervals; observed-score
discrete reliability; a scalar coefficient for unordered categorical outcomes;
joint Gaussian-discrete fitting; a public sparse discrete backend, which ships
only as a private prototype behind an evaluator seam that no public entry point
selects. The dense discrete engine is bounded to small models by row, dimension,
parameter and working-memory limits, and the bundled 21,600-row panels exceed
them by more than an order of magnitude: they are data resources, not
full-panel fitting examples. Planned work is in
docs/ROADMAP.md.

Publishing here establishes no CRAN status. Annotation data retain CC BY 4.0;
package code is GPL-3.

Gtheory4LLM 0.1.0

Choose a tag to compare

@Veronica0206 Veronica0206 released this 14 Sep 01:45

Generalizability Theory for LLM subjective tasks: exact balanced Gaussian
G-theory models, and a bounded experimental Laplace engine for small discrete
G-theory models used in LLM annotation and evaluation studies.

This is a research beta. Passing the package's numerical acceptance checks
does not establish parameter recovery, interval coverage, Laplace approximation
quality, or a scientifically sufficient number of evaluators. Read
docs/LIMITATIONS.md
before quoting a coefficient, and
docs/VALIDATION_SCOPE.md
for what has actually been validated.

Install

install.packages("OpenMx")
install.packages("Gtheory4LLM_0.1.0.tar.gz", repos = NULL, type = "source")

R 4.5.0 or later is required. Nothing in this package's own R code needs it; the
floor is inherited from OpenMx 2.22.11, which calls a C entry point introduced in
R 4.5.0 while declaring only R (>= 3.5.0) itself.

Verifying these files

manifest.json records the source commit these bytes were built from, together
with their sizes and SHA-256 digests:

Gtheory4LLM_0.1.0.tar.gz   268403 bytes
  401a10b0a56848d2d4d60ef48548ba483ec247920534594c2d573357d1a14f88
Gtheory4LLM-manual.pdf     194456 bytes
  5b9054f9adff3e32bea73b45248cc778f44376a31848c5d8607b95481352c4ee

From a checkout of the v0.1.0 tag, this re-derives the whole chain — archive
integrity, correspondence to the recorded source commit, release identity, then
an isolated install and the archive's own smoke test:

python3 scripts/check_committed_artifact.py --check-release-identity --release-tag v0.1.0

What is in it

Exact balanced Gaussian ML/REML for univariate and jointly multivariate
outcomes, with Wald standard errors for variance components and delta-method
coefficient intervals. Crossed, nested and mixed designs with explicit source
selection; Brennan (2001) fixed facets; balanced decision studies. A dense
first-order Laplace engine for small binary, ordinal and unordered categorical
models. Three publicly archived LLM annotation panels.

The full change list is in
NEWS.md.

What is not in it

Unbalanced designs; discrete standard errors or intervals; observed-score
discrete reliability; a scalar coefficient for unordered categorical outcomes;
joint Gaussian-discrete fitting; a sparse discrete backend. The dense discrete
engine is bounded to small models by row, dimension, parameter and
working-memory limits, and the bundled 21,600-row panels exceed them by more
than an order of magnitude: they are data resources, not full-panel fitting
examples. Planned work is in
docs/ROADMAP.md.

Publishing here establishes no CRAN status. Annotation data retain CC BY 4.0;
package code is GPL-3.

Gtheory4LLM 0.0.7

Choose a tag to compare

@Veronica0206 Veronica0206 released this 13 Sep 04:21

Adds coefficient uncertainty and mixed-model fixed facets.

Uncertainty. Gaussian fits report asymptotic Wald standard errors for every source variance, and gt_reliability() / gt_dstudy() report delta-method intervals for G and Phi. Where a zero-variance component makes the joint Hessian indefinite, standard errors are computed on the interior block conditional on those components being held at zero, and the fit says so. The discrete Laplace engine computes no observed information and reports point estimates only.

Fixed facets. fixed= applies the mixed model of Brennan (2001): object-by-fixed-facet variance is averaged over that facet's levels and joins universe-score variance, and a source built only from fixed facets leaves the model. A fixed facet's count cannot be changed and a decision study cannot project over it.

Designs and docs. gt_design(full_cell = FALSE) removes exactly the object-by-all-facets source, which is how a binary or ordinal study with one observation per cell declares the default crossed design. The tutorial is now a real vignette (browseVignettes("Gtheory4LLM")). Dataset help records that temperature is a chosen setting rather than a sampled level, and that seed agreement falls with temperature, so one pooled seed variance is misspecified for these panels.

See NEWS.md for the full list.

Verify this archive against its recorded source commit with:

python3 scripts/run_validation.py --scope artifact

Gtheory4LLM 0.0.6

Choose a tag to compare

@Veronica0206 Veronica0206 released this 13 Sep 04:23

First public distribution under GPL-3.

Bundles eight outcome sets and fifteen codings of three real, publicly archived LLM annotation panels, preserving public-source checksums and CC BY 4.0 data attribution. Adds gt_preflight(), an installed synthetic LLM tutorial, and a concise classed fit summary.

Scope at this release is exact balanced Gaussian likelihood and small-model first-order Laplace discrete likelihood. Uncertainty intervals of any kind are not implemented here; see v0.0.7 for standard errors and coefficient intervals.

Archive and manual checksums are recorded in manifest.json, which pins them to source commit 6a7667e. Verify with:

python3 scripts/run_validation.py --scope artifact