Theano-based implementation of Deep Q-learning
Python Shell
Switch branches/tags
Nothing to show
Clone or download


This package provides a Lasagne/Theano-based implementation of the deep Q-learning algorithm described in:

Playing Atari with Deep Reinforcement Learning Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, Martin Riedmiller


Mnih, Volodymyr, et al. "Human-level control through deep reinforcement learning." Nature 518.7540 (2015): 529-533.

Here is a video showing a trained network playing breakout (using an earlier version of the code):


The script can be used to install all dependencies under Ubuntu.


Use the scripts or to start all the necessary processes:

$ ./ --rom breakout

$ ./ --rom breakout

The script uses parameters consistent with the original NIPS workshop paper. This code should take 2-4 days to complete. The script uses parameters consistent with the Nature paper. The final policies should be better, but it will take 6-10 days to finish training.

Either script will store output files in a folder prefixed with the name of the ROM. Pickled version of the network objects are stored after every epoch. The file results.csv will contain the testing output. You can plot the progress by executing

$ python breakout_05-28-17-09_0p00025_0p99/results.csv

After training completes, you can watch the network play using the script:

$ python breakout_05-28-17-09_0p00025_0p99/network_file_99.pkl

Performance Tuning

Theano Configuration

Setting allow_gc=False in THEANO_FLAGS or in the .theanorc file significantly improves performance at the expense of a slight increase in memory usage on the GPU.

Getting Help

The deep Q-learning web-forum can be used for discussion and advice related to deep Q-learning in general and this package in particular.

See Also