Skip to content

Releases: ContextLab/llm-stylometry

v2.0 — Paper Revision

Choose a tag to compare

@jeremymanning jeremymanning released this 28 Mar 04:35
1411b89

What's New

This release accompanies the revised manuscript submitted to Computational Linguistics, addressing reviewer feedback on the original submission.

New Analyses

  • Dataset-size sweep: Systematically varied training tokens per author (2,500–643,041) across 1,520 newly trained models to characterize data requirements. A sigmoid fit (R²=0.979) estimates that ≥95% attribution accuracy requires ~51,000 tokens per author.
  • Embedding comparison: Evaluated three pre-trained text embedding models from the MTEB leaderboard (nomic-embed-text-v1.5, bge-m3, Qwen3-Embedding-4B) against our predictive comparison approach. Best embedding accuracy: 81.0% vs our 100%.

Paper Updates

  • New methods and results sections for both analyses
  • Expanded discussion of relationship to Huang et al. (2025), benchmark feasibility, cross-domain robustness
  • Supplementary materials with embedding purity/confusion figures
  • Point-by-point response letter addressing editor and 3 reviewers
  • Regenerated Oz attribution figure (previously had empty panels)

Code & Infrastructure

  • code/fit_sigmoid.py: Sigmoid fit with bootstrap confidence intervals
  • code/embedding_comparison.py: Chunk-level nearest-neighbor attribution pipeline with per-book checkpointing
  • New figure types (flags 6, 7) in run_llm_stylometry.sh and generate_figures.py
  • 3 new remote scripts for dataset-size sweep on GPU clusters
  • paper/compile.sh: Builds main paper, supplement, response letter, and latexdiff
  • 15 new tests covering sigmoid fit and embedding comparison
  • Black formatting applied across codebase

Data

  • data/model_results_ntokens.pkl.gz: Pre-computed results for the dataset-size sweep (98MB)
  • Embedding results cached in data/embedding_results/ (gitignored; regenerable via code/embedding_comparison.py)

v1.0 - Public Release

Choose a tag to compare

@jeremymanning jeremymanning released this 10 Dec 01:29

LLM Stylometry v1.0 - Public Release

Paper: A Stylometric Application of Large Language Models (Stropkay et al., 2025)

This release accompanies the arXiv preprint and makes all code, data, models, and analyses publicly available.

Key Features

📊 Reproducible Analysis

  • 320 trained models (8 authors × 10 seeds × 4 conditions)
  • Pre-computed results included for all figures
  • One-line figure generation from pre-computed data
  • Complete analysis pipeline from raw data to publication figures

🤖 HuggingFace Models (NEW!)

All 8 author-specific GPT-2 models publicly available:

Each model trained for 50,000 epochs (final loss ~1.2-1.5).

📚 HuggingFace Datasets (NEW!)

All 8 author text corpora with verified book titles:

📦 Pre-trained Model Weights

  • Dropbox distribution for all 320 paper models
  • Automatic download script with checksum verification
  • ~26GB compressed archives
  • See models/README.md for download instructions

🎨 Visualization & Analysis

  • 7 main figures (baseline condition)
  • 32 supplemental figures (3 linguistic variants)
  • Statistical analyses (t-tests, cross-variant comparisons)
  • Text classification experiments

Quick Start

# Clone repository
git clone https://github.com/ContextLab/llm-stylometry.git
cd llm-stylometry

# Generate all figures (from pre-computed results)
./run_llm_stylometry.sh

# Or generate specific figure
./run_llm_stylometry.sh -f 1a

What's Included

  • ✅ Complete training pipeline for GPT-2 models
  • ✅ Visualization tools for all paper figures
  • ✅ Statistical analysis scripts
  • ✅ Text classification experiments
  • ✅ Comprehensive documentation
  • ✅ Full test suite (pytest)
  • ✅ CI/CD integration (GitHub Actions)

Installation

One-line setup:

./run_llm_stylometry.sh

This automatically creates conda environment, installs dependencies, and generates all figures.

Citation

@article{StroEtal25,
  title={A Stylometric Application of Large Language Models},
  author={Stropkay, Harrison F. and Chen, Jiayi and Jabelli, Mohammad J. L. and Rockmore, Daniel N. and Manning, Jeremy R.},
  journal={arXiv preprint arXiv:2510.21958},
  year={2025}
}

Contact


Major contributors: Harrison Stropkay, Jiayi Chen, Mohammad Jabelli, Daniel Rockmore, Jeremy Manning

License: MIT