Skip to content

Releases: ActiveInferenceInstitute/on_policy_distillation

On-Policy Distillation as Active Inference in Finite Variational Models (v1.0.2)

Choose a tag to compare

@docxology docxology released this 18 Jun 15:54

On-Policy Distillation as Active Inference in Finite Variational Models — v1.0.2

DOI (version): https://doi.org/10.5281/zenodo.20749817
DOI (concept, always latest): https://doi.org/10.5281/zenodo.20747834
Zenodo record: https://zenodo.org/records/20749817

Changes since v1.0.1:

  • Added a cover-art title page (graphical abstract) to the rendered/deposited PDF.

PDF SHA-256: c6b5ec494915e6e046f24cf723f8dbbf93a5b168544daed3cca14c089d4087aa

On-Policy Distillation as Active Inference in Finite Variational Models (v1.0.1)

Choose a tag to compare

@docxology docxology released this 18 Jun 15:18

On-Policy Distillation as Active Inference in Finite Variational Models — v1.0.1

DOI (version): https://doi.org/10.5281/zenodo.20748663
DOI (concept, always latest): https://doi.org/10.5281/zenodo.20747834
Zenodo record: https://zenodo.org/records/20748663

Changes since v1.0.0:

  • New Zenodo version with the DOI baked into the manuscript metadata and PDF.
  • Dropped a duplicate "Appendix" header in the reproducibility supplement.
  • Figure-source audit now allows the transmission-bookend figures.

PDF SHA-256: 4f7040bccf04ccc9f50a984e371ef927753b7d9699c06f2dd6d2960b23752ba4

On-Policy Distillation as Active Inference in Finite Variational Models (v1.0.0)

Choose a tag to compare

@docxology docxology released this 18 Jun 13:42

Release v1.0.0 for working/active_inference_on_policy_distillation.

Publication

Abstract

Abstract

This paper formulates on-policy distillation as active inference in finite variational models, with exact claims only for declared objects and interpretive claims explicitly bounded outside them. In the construction, the intractable teacher policy plays the role of the generative model $p(o,s)$, the tractable student policy is the approximate posterior $q(s)$, and the per-token reverse-KL distillation loss is variational free energy up to the evidence constant, $F = D_{\mathrm{KL}}(q,|,p(s\mid o)) - \log p(o)$, whose KL target is the teacher-induced posterior $p(s\mid o)\propto p(o,s)$ .

The title's "as" is therefore a scoped mathematical correspondence rather than the slogan OPD = Active Inference. Variational free energy names the realized-rollout distillation loss; expected free energy remains the planning-side objective by which the pymdp agent selects actions . On-policy student rollouts generate the observations on which the posterior is scored, connecting the construction to induced-distribution mismatch in imitation learning and exposure-bias analyses while preserving their different objectives, empirical regimes, and contested severity . Privileged traces and feedback play the role that train-time-only information plays in the LUPI/distillation lineage .

Four deterministic witnesses instantiate the correspondence. A Bernoulli-Ising oracle couples a teacher's privileged variable to the answer through $\lambda$, making $I(\lambda)$ the teacher-student mutual information and the finite free-energy gap the toy distillation objective; the closed-form and independently recomputed mutual-information sweeps agree to machine precision (RMSE 2.1e-16 nats). A pymdp T-maze rollout supplies the on-policy student that samples its own observations under a privileged cue . A two-agent classroom pits a privileged teacher (cue validity 0.98) against an on-policy student (cue validity 0.5), measuring teacher belief entropy 0.247 nats versus student 0.347 nats and a mean reverse-KL distillation signal of 6.28 nats. A four-state/two-action sequential-shift witness shows teacher-forced train loss 0.333 nats underestimating student-induced test loss 0.409 nats, with deterministic on-policy correction reducing it to 0.096 nats.

These are toy, generated findings, not production-LLM measurements. Recent privileged-context, context-distillation, adaptive-teacher, freshness-aware OPD, RLHF/instruction-tuning, self-generated reasoning, Qwen OPD-vs-RL, and Thinking Machines replication reports remain external context rather than reproduced results . The supplemental sheaf/provenance layer keeps that boundary operational: every reported number is hydrated from a generated artifact, every figure is source-bound, and 16 / 16 invariant checks pass before rendering.