Releases: ActiveInferenceInstitute/on_policy_distillation
Release list
On-Policy Distillation as Active Inference in Finite Variational Models (v1.0.2)
On-Policy Distillation as Active Inference in Finite Variational Models — v1.0.2
DOI (version): https://doi.org/10.5281/zenodo.20749817
DOI (concept, always latest): https://doi.org/10.5281/zenodo.20747834
Zenodo record: https://zenodo.org/records/20749817
Changes since v1.0.1:
- Added a cover-art title page (graphical abstract) to the rendered/deposited PDF.
PDF SHA-256: c6b5ec494915e6e046f24cf723f8dbbf93a5b168544daed3cca14c089d4087aa
On-Policy Distillation as Active Inference in Finite Variational Models (v1.0.1)
On-Policy Distillation as Active Inference in Finite Variational Models — v1.0.1
DOI (version): https://doi.org/10.5281/zenodo.20748663
DOI (concept, always latest): https://doi.org/10.5281/zenodo.20747834
Zenodo record: https://zenodo.org/records/20748663
Changes since v1.0.0:
- New Zenodo version with the DOI baked into the manuscript metadata and PDF.
- Dropped a duplicate "Appendix" header in the reproducibility supplement.
- Figure-source audit now allows the transmission-bookend figures.
PDF SHA-256: 4f7040bccf04ccc9f50a984e371ef927753b7d9699c06f2dd6d2960b23752ba4
On-Policy Distillation as Active Inference in Finite Variational Models (v1.0.0)
Release v1.0.0 for working/active_inference_on_policy_distillation.
Publication
- Version: 1.0.0
- GitHub release: https://github.com/ActiveInferenceInstitute/on_policy_distillation/releases/tag/v1.0.0
- DOI: https://doi.org/10.5281/zenodo.20747834
- Zenodo: https://zenodo.org/records/20747834
- PDF SHA-256:
db0f2e4f193efa4dd65058d8d8094659a7c9200454acb2da14e88c304f29819e
Abstract
Abstract
This paper formulates on-policy distillation as active inference in finite variational models, with exact claims only for declared objects and interpretive claims explicitly bounded outside them. In the construction, the intractable teacher policy plays the role of the generative model
The title's "as" is therefore a scoped mathematical correspondence rather than the slogan OPD = Active Inference. Variational free energy names the realized-rollout distillation loss; expected free energy remains the planning-side objective by which the pymdp agent selects actions . On-policy student rollouts generate the observations on which the posterior is scored, connecting the construction to induced-distribution mismatch in imitation learning and exposure-bias analyses while preserving their different objectives, empirical regimes, and contested severity . Privileged traces and feedback play the role that train-time-only information plays in the LUPI/distillation lineage .
Four deterministic witnesses instantiate the correspondence. A Bernoulli-Ising oracle couples a teacher's privileged variable to the answer through
These are toy, generated findings, not production-LLM measurements. Recent privileged-context, context-distillation, adaptive-teacher, freshness-aware OPD, RLHF/instruction-tuning, self-generated reasoning, Qwen OPD-vs-RL, and Thinking Machines replication reports remain external context rather than reproduced results . The supplemental sheaf/provenance layer keeps that boundary operational: every reported number is hydrated from a generated artifact, every figure is source-bound, and 16 / 16 invariant checks pass before rendering.