What's New
This release accompanies the revised manuscript submitted to Computational Linguistics, addressing reviewer feedback on the original submission.
New Analyses
- Dataset-size sweep: Systematically varied training tokens per author (2,500–643,041) across 1,520 newly trained models to characterize data requirements. A sigmoid fit (R²=0.979) estimates that ≥95% attribution accuracy requires ~51,000 tokens per author.
- Embedding comparison: Evaluated three pre-trained text embedding models from the MTEB leaderboard (nomic-embed-text-v1.5, bge-m3, Qwen3-Embedding-4B) against our predictive comparison approach. Best embedding accuracy: 81.0% vs our 100%.
Paper Updates
- New methods and results sections for both analyses
- Expanded discussion of relationship to Huang et al. (2025), benchmark feasibility, cross-domain robustness
- Supplementary materials with embedding purity/confusion figures
- Point-by-point response letter addressing editor and 3 reviewers
- Regenerated Oz attribution figure (previously had empty panels)
Code & Infrastructure
code/fit_sigmoid.py: Sigmoid fit with bootstrap confidence intervalscode/embedding_comparison.py: Chunk-level nearest-neighbor attribution pipeline with per-book checkpointing- New figure types (flags 6, 7) in
run_llm_stylometry.shandgenerate_figures.py - 3 new remote scripts for dataset-size sweep on GPU clusters
paper/compile.sh: Builds main paper, supplement, response letter, and latexdiff- 15 new tests covering sigmoid fit and embedding comparison
- Black formatting applied across codebase
Data
data/model_results_ntokens.pkl.gz: Pre-computed results for the dataset-size sweep (98MB)- Embedding results cached in
data/embedding_results/(gitignored; regenerable viacode/embedding_comparison.py)