(WIP) Exploring adaptive computation-type algorithms.
Started with Continuous Thought Machines' loss (and a "dense" version) on the toy parity task:
Methods to try:
- Heuristic halting (convergence, confidence, etc.)
- Halting heads (PonderNet, HRM (Q learning), TRM, etc)
- "Emergent" halting (CTM's loss, etc)