Skip to content

v2.0 — Paper Revision

Latest

Choose a tag to compare

@jeremymanning jeremymanning released this 28 Mar 04:35
· 3 commits to main since this release
1411b89

What's New

This release accompanies the revised manuscript submitted to Computational Linguistics, addressing reviewer feedback on the original submission.

New Analyses

  • Dataset-size sweep: Systematically varied training tokens per author (2,500–643,041) across 1,520 newly trained models to characterize data requirements. A sigmoid fit (R²=0.979) estimates that ≥95% attribution accuracy requires ~51,000 tokens per author.
  • Embedding comparison: Evaluated three pre-trained text embedding models from the MTEB leaderboard (nomic-embed-text-v1.5, bge-m3, Qwen3-Embedding-4B) against our predictive comparison approach. Best embedding accuracy: 81.0% vs our 100%.

Paper Updates

  • New methods and results sections for both analyses
  • Expanded discussion of relationship to Huang et al. (2025), benchmark feasibility, cross-domain robustness
  • Supplementary materials with embedding purity/confusion figures
  • Point-by-point response letter addressing editor and 3 reviewers
  • Regenerated Oz attribution figure (previously had empty panels)

Code & Infrastructure

  • code/fit_sigmoid.py: Sigmoid fit with bootstrap confidence intervals
  • code/embedding_comparison.py: Chunk-level nearest-neighbor attribution pipeline with per-book checkpointing
  • New figure types (flags 6, 7) in run_llm_stylometry.sh and generate_figures.py
  • 3 new remote scripts for dataset-size sweep on GPU clusters
  • paper/compile.sh: Builds main paper, supplement, response letter, and latexdiff
  • 15 new tests covering sigmoid fit and embedding comparison
  • Black formatting applied across codebase

Data

  • data/model_results_ntokens.pkl.gz: Pre-computed results for the dataset-size sweep (98MB)
  • Embedding results cached in data/embedding_results/ (gitignored; regenerable via code/embedding_comparison.py)