Releases: ContextLab/llm-stylometry
Releases · ContextLab/llm-stylometry
Release list
v2.0 — Paper Revision
What's New
This release accompanies the revised manuscript submitted to Computational Linguistics, addressing reviewer feedback on the original submission.
New Analyses
- Dataset-size sweep: Systematically varied training tokens per author (2,500–643,041) across 1,520 newly trained models to characterize data requirements. A sigmoid fit (R²=0.979) estimates that ≥95% attribution accuracy requires ~51,000 tokens per author.
- Embedding comparison: Evaluated three pre-trained text embedding models from the MTEB leaderboard (nomic-embed-text-v1.5, bge-m3, Qwen3-Embedding-4B) against our predictive comparison approach. Best embedding accuracy: 81.0% vs our 100%.
Paper Updates
- New methods and results sections for both analyses
- Expanded discussion of relationship to Huang et al. (2025), benchmark feasibility, cross-domain robustness
- Supplementary materials with embedding purity/confusion figures
- Point-by-point response letter addressing editor and 3 reviewers
- Regenerated Oz attribution figure (previously had empty panels)
Code & Infrastructure
code/fit_sigmoid.py: Sigmoid fit with bootstrap confidence intervalscode/embedding_comparison.py: Chunk-level nearest-neighbor attribution pipeline with per-book checkpointing- New figure types (flags 6, 7) in
run_llm_stylometry.shandgenerate_figures.py - 3 new remote scripts for dataset-size sweep on GPU clusters
paper/compile.sh: Builds main paper, supplement, response letter, and latexdiff- 15 new tests covering sigmoid fit and embedding comparison
- Black formatting applied across codebase
Data
data/model_results_ntokens.pkl.gz: Pre-computed results for the dataset-size sweep (98MB)- Embedding results cached in
data/embedding_results/(gitignored; regenerable viacode/embedding_comparison.py)
v1.0 - Public Release
LLM Stylometry v1.0 - Public Release
Paper: A Stylometric Application of Large Language Models (Stropkay et al., 2025)
This release accompanies the arXiv preprint and makes all code, data, models, and analyses publicly available.
Key Features
📊 Reproducible Analysis
- 320 trained models (8 authors × 10 seeds × 4 conditions)
- Pre-computed results included for all figures
- One-line figure generation from pre-computed data
- Complete analysis pipeline from raw data to publication figures
🤖 HuggingFace Models (NEW!)
All 8 author-specific GPT-2 models publicly available:
- Jane Austen
- L. Frank Baum
- Charles Dickens
- F. Scott Fitzgerald
- Herman Melville
- Ruth Plumly Thompson
- Mark Twain
- H.G. Wells
Each model trained for 50,000 epochs (final loss ~1.2-1.5).
📚 HuggingFace Datasets (NEW!)
All 8 author text corpora with verified book titles:
- 84 books total from Project Gutenberg
- Cleaned and preprocessed for stylometry
- Professionally documented dataset cards
- Browse at: https://huggingface.co/contextlab
📦 Pre-trained Model Weights
- Dropbox distribution for all 320 paper models
- Automatic download script with checksum verification
- ~26GB compressed archives
- See
models/README.mdfor download instructions
🎨 Visualization & Analysis
- 7 main figures (baseline condition)
- 32 supplemental figures (3 linguistic variants)
- Statistical analyses (t-tests, cross-variant comparisons)
- Text classification experiments
Quick Start
# Clone repository
git clone https://github.com/ContextLab/llm-stylometry.git
cd llm-stylometry
# Generate all figures (from pre-computed results)
./run_llm_stylometry.sh
# Or generate specific figure
./run_llm_stylometry.sh -f 1aWhat's Included
- ✅ Complete training pipeline for GPT-2 models
- ✅ Visualization tools for all paper figures
- ✅ Statistical analysis scripts
- ✅ Text classification experiments
- ✅ Comprehensive documentation
- ✅ Full test suite (pytest)
- ✅ CI/CD integration (GitHub Actions)
Installation
One-line setup:
./run_llm_stylometry.shThis automatically creates conda environment, installs dependencies, and generates all figures.
Citation
@article{StroEtal25,
title={A Stylometric Application of Large Language Models},
author={Stropkay, Harrison F. and Chen, Jiayi and Jabelli, Mohammad J. L. and Rockmore, Daniel N. and Manning, Jeremy R.},
journal={arXiv preprint arXiv:2510.21958},
year={2025}
}Contact
- Paper: https://arxiv.org/abs/2510.21958
- Code: https://github.com/ContextLab/llm-stylometry
- Issues: https://github.com/ContextLab/llm-stylometry/issues
- ContextLab: https://www.context-lab.com/
Major contributors: Harrison Stropkay, Jiayi Chen, Mohammad Jabelli, Daniel Rockmore, Jeremy Manning
License: MIT