Course papers and projects
- Policy Gradient Methods for Continuous Control Tasks. Will one Mujoco Hopper rule them all? (probably, and it's the SAC)
- Value Function Approximation with Temporal Difference Learning. Why do we know so little about the theoretical properties of TD(λ)?
- Error Bounds for Approximate Dynammic Programming. What are they, anyways?
- A Probabilistic Interpretation of Data Augmentations. A solo venture to characterize data augmentations based on their information about the learning problem.
- Computing Entanglement via Tensor Network Optimization. Can we use tensor networks to speed up quantum entanglement calculations? (yes we can!)