PyTorch implementation of TRPO

Try my implementation of PPO (aka newer better variant of TRPO), unless you need to you TRPO for some specific reasons.

This is a PyTorch implementation of "Trust Region Policy Optimization (TRPO)".

This is code mostly ported from original implementation by John Schulman and forked from here. In contrast to another implementation of TRPO in PyTorch, this implementation uses exact Hessian-vector product instead of finite differences approximation.

Contributions

We are presenting a new method for optimization, called 'Adaptive regularized cubics using L-SR1 hessian approximations'. We present the comparitive results with TRPO, forked from here. All implementations are in pytorch.

Usage

python main.py --env-name "Reacher-v1"

Recommended hyper parameters

InvertedPendulum-v1: 5000

Reacher-v1, InvertedDoublePendulum-v1: 15000

HalfCheetah-v1, Hopper-v1, Swimmer-v1, Walker2d-v1: 25000

Ant-v1, Humanoid-v1: 50000

Results

More or less similar to the original code. Coming soon.

Name		Name	Last commit message	Last commit date
Latest commit History 55 Commits
.DS_Store		.DS_Store
.gitignore		.gitignore
Contour-plotter.ipynb		Contour-plotter.ipynb
InteriorPointMethod.py		InteriorPointMethod.py
LICENSE.md		LICENSE.md
README.md		README.md
RL_ARCLSR1.py		RL_ARCLSR1.py
conjugate_gradients.py		conjugate_gradients.py
constrainedexample.py		constrainedexample.py
main.py		main.py
models.py		models.py
plotting.ipynb		plotting.ipynb
replay_memory.py		replay_memory.py
requirements.txt		requirements.txt
running_state.py		running_state.py
trpo.py		trpo.py
utils.py		utils.py

License

aranganath/pytorch-trpo

Folders and files

Latest commit

History

Repository files navigation

PyTorch implementation of TRPO

Contributions

Usage

Recommended hyper parameters

Results

About

Resources

License

Stars

Watchers

Forks

Languages