PyTorch implementation of TRPO

Overview

This is the code for this video on Youtube by Siraj Raval. This is an implementation of the Trust Region Policy Optimization algorithm that was used by the researchers in the video. They did not, however, make their full code public. So here is the technique applied to game environments. Someone can use it as a starting point to recreate their code. Meanwhile -- hey researchers :) go ahead and release it the community would appreciate it.

PyTorch implementation of TRPO

Try this implementation of PPO (aka newer better variant of TRPO), unless you need to you TRPO for some specific reasons.

This is a PyTorch implementation of "Trust Region Policy Optimization (TRPO)".

This is code mostly ported from original implementation by John Schulman. In contrast to another implementation of TRPO in PyTorch, this implementation uses exact Hessian-vector product instead of finite differences approximation.

Contributions

Contributions are very welcome. If you know how to make this code better, don't hesitate to send a pull request.

Usage

python main.py --env-name "Reacher-v1"

Recommended hyper parameters

InvertedPendulum-v1: 5000

Reacher-v1, InvertedDoublePendulum-v1: 15000

HalfCheetah-v1, Hopper-v1, Swimmer-v1, Walker2d-v1: 25000

Ant-v1, Humanoid-v1: 50000

Credits

Credits for this code go to ikostrikov. I've merely created a wrapper to get people started.

Name		Name	Last commit message	Last commit date
Latest commit History 4 Commits
LICENSE.md		LICENSE.md
README.md		README.md
conjugate_gradients.py		conjugate_gradients.py
main.py		main.py
models.py		models.py
replay_memory.py		replay_memory.py
running_state.py		running_state.py
trpo.py		trpo.py
utils.py		utils.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Overview

PyTorch implementation of TRPO

Contributions

Usage

Recommended hyper parameters

Credits

About

Releases

Packages

Languages

License

llSourcell/AI_Dresses_Itself

Folders and files

Latest commit

History

Repository files navigation

Overview

PyTorch implementation of TRPO

Contributions

Usage

Recommended hyper parameters

Credits

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages