Skip to content

v3.1.0 — Workshop-competitive research release

Choose a tag to compare

@arnavd371 arnavd371 released this 04 Aug 11:24

Summary

Workshop-facing research release on top of the v3 distilled production model.

  • Multi-seed advanced suite (42/43/44) with bootstrap 95% CIs — robustness, calibration, hierarchical heads, SSL
  • Hierarchical fine-tune (strict gate): 97.15% vs 96.84% baseline (seed 42)
  • CCAI @ NeurIPS 2026 Papers-track draft in workshop/
  • LIMITATIONS.md — IRRI ≠ phone-on-bag; soft 84.51% reference
  • Apache-2.0 — filled LICENSE appendix + detailed NOTICE; web/ aligned from MIT

Copyright © 2026 Arnav Dhiman (arnavd371@gmail.com). Sole GitHub contributor.

Key numbers

Metric Value
Baseline acc (3-seed mean) 96.73% ± 0.37%
SSL acc (3-seed mean) 96.62% ± 0.18%
Phone-band robustness 60.9% ± 8.4%
SNR ≤10 / hard combos ~7.6%
Hier fine-tune (seed 42) 97.15%

Paths

  • Aggregate: experiments/results/advanced_multiseed/
  • Hier FT: experiments/results/hier_finetune/
  • Paper: workshop/kaan_ccai_neurips2026.pdf
  • Changelog: CHANGELOG.md