# Navigation

---

In this notebook, you will learn how to use the Unity ML-Agents environment for the first project of the [Deep Reinforcement Learning Nanodegree](https://www.udacity.com/course/deep-reinforcement-learning-nanodegree--nd893).

### 1. Start the Environment

We begin by importing some necessary packages.  If the code cell below returns an error, please revisit the project instructions to double-check that you have installed [Unity ML-Agents](https://github.com/Unity-Technologies/ml-agents/blob/master/docs/Installation.md) and [NumPy](http://www.numpy.org/).

In [49]:
from unityagents import UnityEnvironment
import numpy as np

Next, we will start the environment!  **_Before running the code cell below_**, change the `file_name` parameter to match the location of the Unity environment that you downloaded.

- **Mac**: `"path/to/Banana.app"`
- **Windows** (x86): `"path/to/Banana_Windows_x86/Banana.exe"`
- **Windows** (x86_64): `"path/to/Banana_Windows_x86_64/Banana.exe"`
- **Linux** (x86): `"path/to/Banana_Linux/Banana.x86"`
- **Linux** (x86_64): `"path/to/Banana_Linux/Banana.x86_64"`
- **Linux** (x86, headless): `"path/to/Banana_Linux_NoVis/Banana.x86"`
- **Linux** (x86_64, headless): `"path/to/Banana_Linux_NoVis/Banana.x86_64"`

For instance, if you are using a Mac, then you downloaded `Banana.app`.  If this file is in the same folder as the notebook, then the line below should appear as follows:
```
env = UnityEnvironment(file_name="Banana.app")
```

In [2]:
# env = UnityEnvironment(file_name="/home/arasdar/VisualBanana_Linux/Banana.x86")
env = UnityEnvironment(file_name="/home/arasdar/unity-envs/Banana_Linux/Banana.x86_64")

INFO:unityagents:
'Academy' started successfully!
Unity Academy name: Academy
        Number of Brains: 1
        Number of External Brains : 1
        Lesson number : 0
        Reset Parameters :
		
Unity brain name: BananaBrain
        Number of Visual Observations (per agent): 0
        Vector Observation space type: continuous
        Vector Observation space size (per agent): 37
        Number of stacked Vector Observation: 1
        Vector Action space type: discrete
        Vector Action space size (per agent): 4
        Vector Action descriptions: , , , 


Environments contain **_brains_** which are responsible for deciding the actions of their associated agents. Here we check for the first brain available, and set it as the default brain we will be controlling from Python.

In [3]:
# get the default brain
brain_name = env.brain_names[0]
brain = env.brains[brain_name]

### 2. Examine the State and Action Spaces

The simulation contains a single agent that navigates a large environment.  At each time step, it has four actions at its disposal:
- `0` - walk forward 
- `1` - walk backward
- `2` - turn left
- `3` - turn right

The state space has `37` dimensions and contains the agent's velocity, along with ray-based perception of objects around agent's forward direction.  A reward of `+1` is provided for collecting a yellow banana, and a reward of `-1` is provided for collecting a blue banana. 

Run the code cell below to print some information about the environment.

In [4]:
# reset the environment
env_info = env.reset(train_mode=True)[brain_name]

# number of agents in the environment
print('Number of agents:', len(env_info.agents))

# number of actions
action_size = brain.vector_action_space_size
print('Number of actions:', action_size)

# examine the state space 
state = env_info.vector_observations[0]
# print('States look like:', state)
state_size = len(state)
print('States have length:', state_size)
# print(state.shape, len(env_info.vector_observations), env_info.vector_observations.shape)

Number of agents: 1
Number of actions: 4
States have length: 37


### 3. Take Random Actions in the Environment

In the next code cell, you will learn how to use the Python API to control the agent and receive feedback from the environment.

Once this cell is executed, you will watch the agent's performance, if it selects an action (uniformly) at random with each time step.  A window should pop up that allows you to observe the agent, as it moves through the environment.  

Of course, as part of the project, you'll have to change the code so that the agent is able to use its experience to gradually choose better actions when interacting with the environment!

In [6]:
env_info = env.reset(train_mode=False)[brain_name] # reset the environment
state = env_info.vector_observations[0]            # get the current state
score = 0                                          # initialize the score
num_steps = 0
while True:
    num_steps += 1
    action = np.random.randint(action_size)        # select an action
    env_info = env.step(action)[brain_name]        # send the action to the environment
    next_state = env_info.vector_observations[0]   # get the next state
    reward = env_info.rewards[0]                   # get the reward
    done = env_info.local_done[0]                  # see if episode has finished
    score += reward                                # update the score
    state = next_state                             # roll over the state to next time step
    if done:                                       # exit loop if episode finished
        print(state.shape)
        break
    
print("Score: {}".format(score))
num_steps

(37,)
Score: -2.0


300

When finished, you can close the environment.

In [7]:
# env.close()

### 4. It's Your Turn!

Now it's your turn to train your own agent to solve the environment!  When training the environment, set `train_mode=True`, so that the line for resetting the environment looks like the following:
```python
env_info = env.reset(train_mode=True)[brain_name]
```

In [8]:
env_info = env.reset(train_mode=True)[brain_name] # reset the environment
state = env_info.vector_observations[0]            # get the current state
score = 0                                          # initialize the score
num_steps = 0
while True:
    num_steps += 1
    action = np.random.randint(action_size)        # select an action
    env_info = env.step(action)[brain_name]        # send the action to the environment
    next_state = env_info.vector_observations[0]   # get the next state
    reward = env_info.rewards[0]                   # get the reward
    done = env_info.local_done[0]                  # see if episode has finished
    score += reward                                # update the score
    state = next_state                             # roll over the state to next time step
    #print(state)
    if done:                                       # exit loop if episode finished
        break
    
print("Score: {}".format(score))
num_steps

Score: 0.0


300

In [30]:
# In this one we should define and detect GPUs for tensorflow
# GPUs or CPU
import tensorflow as tf

# Check TensorFlow Version
print('TensorFlow Version: {}'.format(tf.__version__))

# Check for a GPU
print('Default GPU Device: {}'.format(tf.test.gpu_device_name()))

TensorFlow Version: 1.7.1
Default GPU Device: 


In [31]:
env_info = env.reset(train_mode=True)[brain_name] # reset the environment
state = env_info.vector_observations[0]            # get the current state
score = 0                                          # initialize the score
batch = []
num_steps = 0
while True: # infinite number of steps
    num_steps += 1
    action = np.random.randint(action_size)        # select an action
    env_info = env.step(action)[brain_name]        # send the action to the environment
    next_state = env_info.vector_observations[0]   # get the next state
    reward = env_info.rewards[0]                   # get the reward
    done = env_info.local_done[0]                  # see if episode has finished
    score += reward                                # update the score
    #print(state, action, reward, done)
    batch.append([state, action, next_state, reward, float(done)])
    state = next_state                             # roll over the state to next time step
    if done:                                       # exit loop if episode finished
        break
    
print("Score: {}".format(score))
num_steps

Score: -2.0


300

In [32]:
batch[0], batch[0][1]

([array([0.        , 0.        , 1.        , 0.        , 0.04087751,
         0.        , 0.        , 1.        , 0.        , 0.04246911,
         1.        , 0.        , 0.        , 0.        , 0.27552733,
         0.        , 0.        , 1.        , 0.        , 0.05028759,
         0.        , 0.        , 1.        , 0.        , 0.05254348,
         0.        , 0.        , 1.        , 0.        , 0.05888711,
         0.        , 0.        , 1.        , 0.        , 0.03669675,
         0.        , 0.        ]),
  0,
  array([0.00000000e+00, 0.00000000e+00, 1.00000000e+00, 0.00000000e+00,
         3.64827104e-02, 0.00000000e+00, 0.00000000e+00, 0.00000000e+00,
         1.00000000e+00, 0.00000000e+00, 0.00000000e+00, 0.00000000e+00,
         1.00000000e+00, 0.00000000e+00, 4.61156890e-02, 0.00000000e+00,
         0.00000000e+00, 1.00000000e+00, 0.00000000e+00, 4.66352813e-02,
         0.00000000e+00, 0.00000000e+00, 1.00000000e+00, 0.00000000e+00,
         4.72340956e-02, 0.00000000e+00

In [33]:
batch[0]

[array([0.        , 0.        , 1.        , 0.        , 0.04087751,
        0.        , 0.        , 1.        , 0.        , 0.04246911,
        1.        , 0.        , 0.        , 0.        , 0.27552733,
        0.        , 0.        , 1.        , 0.        , 0.05028759,
        0.        , 0.        , 1.        , 0.        , 0.05254348,
        0.        , 0.        , 1.        , 0.        , 0.05888711,
        0.        , 0.        , 1.        , 0.        , 0.03669675,
        0.        , 0.        ]),
 0,
 array([0.00000000e+00, 0.00000000e+00, 1.00000000e+00, 0.00000000e+00,
        3.64827104e-02, 0.00000000e+00, 0.00000000e+00, 0.00000000e+00,
        1.00000000e+00, 0.00000000e+00, 0.00000000e+00, 0.00000000e+00,
        1.00000000e+00, 0.00000000e+00, 4.61156890e-02, 0.00000000e+00,
        0.00000000e+00, 1.00000000e+00, 0.00000000e+00, 4.66352813e-02,
        0.00000000e+00, 0.00000000e+00, 1.00000000e+00, 0.00000000e+00,
        4.72340956e-02, 0.00000000e+00, 0.00000000e+00

In [34]:
states = np.array([each[1] for each in batch])
actions = np.array([each[0] for each in batch])
next_states = np.array([each[1] for each in batch])
rewards = np.array([each[2] for each in batch])
dones = np.array([each[3] for each in batch])
# infos = np.array([each[4] for each in batch])

In [35]:
# print(rewards[:])
print(np.array(rewards).shape, np.array(states).shape, np.array(actions).shape, np.array(dones).shape)
print(np.array(rewards).dtype, np.array(states).dtype, np.array(actions).dtype, np.array(dones).dtype)
print(np.max(np.array(actions)), np.min(np.array(actions)), 
      (np.max(np.array(actions)) - np.min(np.array(actions)))+1)
print(np.max(np.array(rewards)), np.min(np.array(rewards)))
print(np.max(np.array(states)), np.min(np.array(states)))

(300, 37) (300,) (300, 37) (300,)
float64 int64 float64 float64
10.520530700683594 -10.738459587097168 22.25899028778076
10.520530700683594 -10.738459587097168
3 0


In [36]:
def model_input(state_size, lstm_size, batch_size=1):
    states = tf.placeholder(tf.float32, [None, state_size], name='states')
    actions = tf.placeholder(tf.int32, [None], name='actions')
    targetQs = tf.placeholder(tf.float32, [None], name='targetQs')
    # RNN
    gru = tf.nn.rnn_cell.GRUCell(lstm_size)
    cell = tf.nn.rnn_cell.MultiRNNCell([gru], state_is_tuple=False)
    initial_state = cell.zero_state(batch_size, tf.float32)
    return states, actions, targetQs, cell, initial_state

In [37]:
# RNN generator or sequence generator
def generator(states, initial_state, cell, lstm_size, num_classes, reuse=False): 
    with tf.variable_scope('generator', reuse=reuse):
        # First fully connected layer
        inputs = tf.layers.dense(inputs=states, units=lstm_size)
        print(states.shape, inputs.shape)
        
        # with tf.variable_scope('dynamic_rnn_', reuse=tf.AUTO_REUSE):
        # dynamic means adapt to the batch_size
        inputs_rnn = tf.reshape(inputs, [1, -1, lstm_size]) # NxH -> 1xNxH
        print(inputs_rnn.shape, initial_state.shape)
        outputs_rnn, final_state = tf.nn.dynamic_rnn(cell=cell, inputs=inputs_rnn, initial_state=initial_state)
        print(outputs_rnn.shape, final_state.shape)
        outputs = tf.reshape(outputs_rnn, [-1, lstm_size]) # 1xNxH -> NxH
        print(outputs.shape)

        # Last fully connected layer
        logits = tf.layers.dense(inputs=outputs, units=num_classes)
        print(logits.shape)
        #predictions = tf.nn.softmax(logits)
        
        # logits are the action logits
        return logits, final_state

In [38]:
def model_loss(action_size, hidden_size, states, cell, initial_state, actions, targetQs):
    actions_logits, final_state = generator(states=states, cell=cell, initial_state=initial_state, 
                                            lstm_size=hidden_size, num_classes=action_size)
    actions_labels = tf.one_hot(indices=actions, depth=action_size, dtype=actions_logits.dtype)
    Qs = tf.reduce_max(actions_logits*actions_labels, axis=1)
    loss = tf.reduce_mean(tf.square(Qs - targetQs))
    return actions_logits, final_state, loss

In [39]:
def model_opt(loss, learning_rate):
    # Get weights and bias to update
    t_vars = tf.trainable_variables()
    g_vars = [var for var in t_vars if var.name.startswith('generator')]

    # # Optimize
    # with tf.control_dependencies(tf.get_collection(tf.GraphKeys.UPDATE_OPS)): # Required for batchnorm (BN)
    # #opt = tf.train.AdamOptimizer(learning_rate).minimize(loss, var_list=g_vars)

    #grads, _ = tf.clip_by_global_norm(t_list=tf.gradients(loss, g_vars), clip_norm=5) # usually around 1-5
    grads = tf.gradients(loss, g_vars)
    opt = tf.train.AdamOptimizer(learning_rate).apply_gradients(grads_and_vars=zip(grads, g_vars))

    return opt

In [40]:
class Model:
    def __init__(self, state_size, action_size, hidden_size, learning_rate):

        # Data of the Model: make the data available inside the framework
        self.states, self.actions, self.targetQs, cell, self.initial_state = model_input(
            state_size=state_size, lstm_size=hidden_size)
        
        # Create the Model: calculating the loss and forwad pass
        self.actions_logits, self.final_state, self.loss = model_loss(
            action_size=action_size, hidden_size=hidden_size, 
            states=self.states, actions=self.actions, 
            targetQs=self.targetQs, cell=cell, initial_state=self.initial_state)

        # Update the model: backward pass and backprop
        self.opt = model_opt(loss=self.loss, learning_rate=learning_rate)

In [41]:
from collections import deque

class Memory():    
    def __init__(self, max_size = 1000):
        self.buffer = deque(maxlen=max_size)
        self.states = deque(maxlen=max_size)

In [42]:
# Network parameters
action_size = 4
state_size = 37
hidden_size = 37*2             # number of units in each Q-network hidden layer
learning_rate = 0.0001         # Q-network learning rate

# Memory parameters
memory_size = 300            # memory capacity
batch_size = 300             # experience mini-batch size
gamma = 0.99                 # future reward discount

In [43]:
# Reset/init the graph/session
graph = tf.reset_default_graph()

# Init the model
model = Model(action_size=action_size, hidden_size=hidden_size, state_size=state_size, learning_rate=learning_rate)

# Init the memory
memory = Memory(max_size=memory_size)

(?, 37) (?, 74)
(1, ?, 74) (1, 74)
(1, ?, 74) (1, 74)
(?, 74)
(?, 4)


In [45]:
# state = env.reset()
# for _ in range(batch_size):
#     action = env.action_space.sample()
#     next_state, reward, done, _ = env.step(action)
#     memory.buffer.append([state, action, next_state, reward, float(done)])
#     state = next_state
#     if done is True:
#         state = env.reset()

In [46]:
env_info = env.reset(train_mode=True)[brain_name] # reset the environment
state = env_info.vector_observations[0]   # get the state
for _ in range(memory_size):
    action = np.random.randint(action_size)        # select an action
    env_info = env.step(action)[brain_name]        # send the action to the environment
    next_state = env_info.vector_observations[0]   # get the next state
    reward = env_info.rewards[0]                   # get the reward
    done = env_info.local_done[0]                  # see if episode has finished
    memory.buffer.append([state, action, next_state, reward, float(done)])
    memory.states.append(np.zeros([1, hidden_size])) # initial_states for rnn/mem
    state = next_state
    if done:                                       # exit loop if episode finished
        env_info = env.reset(train_mode=True)[brain_name] # reset the environment
        state = env_info.vector_observations[0]   # get the state
        break

In [47]:
# initial_states = memory.states
memory.states[0].shape

(1, 74)

In [48]:
# Save/load the model and save for plotting
saver = tf.train.Saver()
episode_rewards_list, rewards_list, loss_list = [], [], []

# TF session for training
with tf.Session(graph=graph) as sess:
    sess.run(tf.global_variables_initializer())
    #saver.restore(sess, 'checkpoints/model.ckpt')    
    #saver.restore(sess, tf.train.latest_checkpoint('checkpoints'))
    total_step = 0 # Explore or exploit parameter
    episode_reward = deque(maxlen=100) # 100 episodes average/running average/running mean/window
    
    # Training episodes/epochs
    for ep in range(11111):
        total_reward = 0
        loss_batch = []
        #state = env.reset()
        env_info = env.reset(train_mode=True)[brain_name] # reset the environment
        state = env_info.vector_observations[0]   # get the current state
        initial_state = sess.run(model.initial_state)

        # Training steps/batches
        for num_steps in range(11111111111):
            action_logits, final_state = sess.run([model.actions_logits, model.final_state],
                                                  feed_dict = {model.states: state.reshape([1, -1]), 
                                                               model.initial_state: initial_state})
            action = np.argmax(action_logits)
            #state, reward, done, _ = env.step(action)
            env_info = env.step(action)[brain_name]        # send the action to the environment
            next_state = env_info.vector_observations[0]   # get the next state
            reward = env_info.rewards[0]                   # get the reward
            done = env_info.local_done[0]                  # see if episode has finished
            memory.buffer.append([state, action, next_state, reward, float(done)])
            memory.states.append(initial_state)
            total_reward += reward
            initial_state = final_state
            state = next_state
            
            # Training
            #batch, rnn_states = memory.sample(batch_size)
            batch = memory.buffer
            states = np.array([each[0] for each in batch])
            actions = np.array([each[1] for each in batch])
            next_states = np.array([each[2] for each in batch])
            rewards = np.array([each[3] for each in batch])
            dones = np.array([each[4] for each in batch])
            initial_states = memory.states
            next_actions_logits = sess.run(model.actions_logits,
                                           feed_dict = {model.states: next_states, 
                                                        model.initial_state: initial_states[1]})
            nextQs = np.max(next_actions_logits, axis=1) * (1-dones)
            targetQs = rewards + (gamma * nextQs)
            loss, _ = sess.run([model.loss, model.opt], feed_dict = {model.states: states, 
                                                                     model.actions: actions,
                                                                     model.targetQs: targetQs,
                                                                     model.initial_state: initial_states[0]})
            loss_batch.append(loss)
            if done is True:
                break
                
        episode_reward.append(total_reward)
        print('Episode:{}'.format(ep),
              'meanR:{:.4f}'.format(np.mean(episode_reward)),
              'R:{}'.format(total_reward),
              'loss:{:.4f}'.format(np.mean(loss_batch)))
        # Ploting out
        episode_rewards_list.append([ep, np.mean(episode_reward)])
        rewards_list.append([ep, total_reward])
        loss_list.append([ep, np.mean(loss_batch)])
        # Break episode/epoch loop
        if np.mean(episode_reward) >= +13:
            break
            
    # At the end of all training episodes/epochs
    saver.save(sess, 'checkpoints/model.ckpt')

Episode:0 meanR:0.0000 R:0.0 loss:0.6707
Episode:1 meanR:0.5000 R:1.0 loss:0.0315
Episode:2 meanR:0.0000 R:-1.0 loss:0.0442
Episode:3 meanR:1.2500 R:5.0 loss:0.0304
Episode:4 meanR:1.4000 R:2.0 loss:0.0518
Episode:5 meanR:1.0000 R:-1.0 loss:0.0390
Episode:6 meanR:0.8571 R:0.0 loss:0.0241
Episode:7 meanR:1.0000 R:2.0 loss:0.0217
Episode:8 meanR:1.0000 R:1.0 loss:0.0270
Episode:9 meanR:1.0000 R:1.0 loss:0.0183
Episode:10 meanR:0.9091 R:0.0 loss:0.0253
Episode:11 meanR:0.7500 R:-1.0 loss:0.0221
Episode:12 meanR:0.5385 R:-2.0 loss:0.0151
Episode:13 meanR:0.5714 R:1.0 loss:0.0109
Episode:14 meanR:0.5333 R:0.0 loss:0.0267
Episode:15 meanR:0.5625 R:1.0 loss:0.0330
Episode:16 meanR:0.5294 R:0.0 loss:0.0284
Episode:17 meanR:0.5000 R:0.0 loss:0.0123
Episode:18 meanR:0.4737 R:0.0 loss:0.0139
Episode:19 meanR:0.4500 R:0.0 loss:0.0252
Episode:20 meanR:0.5238 R:2.0 loss:0.0133
Episode:21 meanR:0.5000 R:0.0 loss:0.0200
Episode:22 meanR:0.4783 R:0.0 loss:0.0408
Episode:23 meanR:0.4583 R:0.0 loss:0.011

Episode:193 meanR:2.7800 R:2.0 loss:0.0581
Episode:194 meanR:2.8600 R:8.0 loss:0.1225
Episode:195 meanR:2.8800 R:3.0 loss:0.0765
Episode:196 meanR:2.9300 R:7.0 loss:0.0931
Episode:197 meanR:2.9200 R:2.0 loss:0.0874
Episode:198 meanR:2.9400 R:3.0 loss:0.0605
Episode:199 meanR:2.9800 R:5.0 loss:0.0381
Episode:200 meanR:2.9900 R:6.0 loss:0.0862
Episode:201 meanR:2.9800 R:6.0 loss:0.0796
Episode:202 meanR:3.0400 R:10.0 loss:0.1346
Episode:203 meanR:3.1200 R:9.0 loss:0.0862
Episode:204 meanR:3.2300 R:13.0 loss:0.0832
Episode:205 meanR:3.3000 R:8.0 loss:0.0829
Episode:206 meanR:3.3800 R:8.0 loss:0.0723
Episode:207 meanR:3.4600 R:8.0 loss:0.0811
Episode:208 meanR:3.5300 R:8.0 loss:0.0558
Episode:209 meanR:3.5900 R:6.0 loss:0.1203
Episode:210 meanR:3.6700 R:7.0 loss:0.0433
Episode:211 meanR:3.7200 R:6.0 loss:0.0582
Episode:212 meanR:3.7600 R:4.0 loss:0.0473
Episode:213 meanR:3.8700 R:12.0 loss:0.0469
Episode:214 meanR:3.9300 R:7.0 loss:0.0427
Episode:215 meanR:4.0000 R:4.0 loss:0.0457
Episode:

Episode:383 meanR:2.2900 R:3.0 loss:0.0610
Episode:384 meanR:2.3600 R:7.0 loss:0.0345
Episode:385 meanR:2.4300 R:7.0 loss:0.0481
Episode:386 meanR:2.4500 R:3.0 loss:0.0377
Episode:387 meanR:2.4800 R:3.0 loss:0.0471
Episode:388 meanR:2.4500 R:2.0 loss:0.0444
Episode:389 meanR:2.4900 R:3.0 loss:0.0358
Episode:390 meanR:2.4600 R:3.0 loss:0.0400
Episode:391 meanR:2.4600 R:4.0 loss:0.0757
Episode:392 meanR:2.4300 R:0.0 loss:0.0277
Episode:393 meanR:2.4300 R:0.0 loss:0.0239
Episode:394 meanR:2.4400 R:1.0 loss:0.0380
Episode:395 meanR:2.4000 R:0.0 loss:0.0379
Episode:396 meanR:2.3600 R:1.0 loss:0.0262
Episode:397 meanR:2.3300 R:0.0 loss:0.0465
Episode:398 meanR:2.2500 R:-2.0 loss:0.0396
Episode:399 meanR:2.2400 R:1.0 loss:0.0295
Episode:400 meanR:2.2100 R:0.0 loss:0.0397
Episode:401 meanR:2.1800 R:0.0 loss:0.0388
Episode:402 meanR:2.1500 R:1.0 loss:0.0322
Episode:403 meanR:2.1300 R:0.0 loss:0.0472
Episode:404 meanR:2.1500 R:1.0 loss:0.0432
Episode:405 meanR:2.1300 R:2.0 loss:0.0308
Episode:40

Episode:574 meanR:3.7400 R:5.0 loss:0.0709
Episode:575 meanR:3.7800 R:7.0 loss:0.0604
Episode:576 meanR:3.7600 R:4.0 loss:0.0356
Episode:577 meanR:3.7000 R:7.0 loss:0.0388
Episode:578 meanR:3.7300 R:2.0 loss:0.0523
Episode:579 meanR:3.6700 R:5.0 loss:0.0449
Episode:580 meanR:3.6300 R:4.0 loss:0.0354
Episode:581 meanR:3.6300 R:5.0 loss:0.0296
Episode:582 meanR:3.6600 R:3.0 loss:0.0367
Episode:583 meanR:3.6200 R:0.0 loss:0.0219
Episode:584 meanR:3.6400 R:5.0 loss:0.0365
Episode:585 meanR:3.6700 R:3.0 loss:0.0427
Episode:586 meanR:3.6900 R:7.0 loss:0.0498
Episode:587 meanR:3.6500 R:1.0 loss:0.1001
Episode:588 meanR:3.6200 R:3.0 loss:0.0484
Episode:589 meanR:3.5900 R:1.0 loss:0.0510
Episode:590 meanR:3.5900 R:0.0 loss:0.0249
Episode:591 meanR:3.5900 R:1.0 loss:0.0416
Episode:592 meanR:3.5300 R:1.0 loss:0.0292
Episode:593 meanR:3.5300 R:1.0 loss:0.0295
Episode:594 meanR:3.4900 R:2.0 loss:0.0401
Episode:595 meanR:3.4600 R:2.0 loss:0.0575
Episode:596 meanR:3.4400 R:5.0 loss:0.0296
Episode:597

Episode:765 meanR:3.1600 R:1.0 loss:0.0616
Episode:766 meanR:3.1800 R:2.0 loss:0.0448
Episode:767 meanR:3.2500 R:9.0 loss:0.0782
Episode:768 meanR:3.2500 R:0.0 loss:0.0962
Episode:769 meanR:3.2900 R:5.0 loss:0.0579
Episode:770 meanR:3.2500 R:0.0 loss:0.0360
Episode:771 meanR:3.2000 R:0.0 loss:0.0194
Episode:772 meanR:3.1100 R:-1.0 loss:0.0432
Episode:773 meanR:3.0800 R:0.0 loss:0.0383
Episode:774 meanR:3.0500 R:1.0 loss:0.0295
Episode:775 meanR:3.0800 R:3.0 loss:0.0283
Episode:776 meanR:3.0400 R:2.0 loss:0.0421
Episode:777 meanR:3.0400 R:5.0 loss:0.0483
Episode:778 meanR:3.0600 R:3.0 loss:0.0302
Episode:779 meanR:3.0800 R:0.0 loss:0.0405
Episode:780 meanR:3.1800 R:6.0 loss:0.0292
Episode:781 meanR:3.2500 R:7.0 loss:0.0659
Episode:782 meanR:3.2900 R:7.0 loss:0.0896
Episode:783 meanR:3.3000 R:4.0 loss:0.0691
Episode:784 meanR:3.2300 R:2.0 loss:0.0634
Episode:785 meanR:3.1900 R:1.0 loss:0.0262
Episode:786 meanR:3.2100 R:3.0 loss:0.0649
Episode:787 meanR:3.1700 R:2.0 loss:0.0566
Episode:78

Episode:956 meanR:3.8400 R:1.0 loss:0.0326
Episode:957 meanR:3.8800 R:5.0 loss:0.0258
Episode:958 meanR:3.8900 R:2.0 loss:0.0465
Episode:959 meanR:3.9300 R:5.0 loss:0.0370
Episode:960 meanR:4.0400 R:11.0 loss:0.0870
Episode:961 meanR:4.1500 R:10.0 loss:0.0448
Episode:962 meanR:4.2200 R:6.0 loss:0.0521
Episode:963 meanR:4.2700 R:7.0 loss:0.0336
Episode:964 meanR:4.3200 R:5.0 loss:0.0375
Episode:965 meanR:4.3300 R:1.0 loss:0.0449
Episode:966 meanR:4.3400 R:4.0 loss:0.0318
Episode:967 meanR:4.3500 R:6.0 loss:0.0517
Episode:968 meanR:4.3400 R:4.0 loss:0.0234
Episode:969 meanR:4.3000 R:0.0 loss:0.0167
Episode:970 meanR:4.3000 R:5.0 loss:0.0226
Episode:971 meanR:4.2100 R:-1.0 loss:0.0199
Episode:972 meanR:4.1500 R:3.0 loss:0.0267
Episode:973 meanR:4.1200 R:-1.0 loss:0.0279
Episode:974 meanR:4.0200 R:2.0 loss:0.0385
Episode:975 meanR:3.9900 R:0.0 loss:0.0600
Episode:976 meanR:4.0200 R:2.0 loss:0.0279
Episode:977 meanR:4.0200 R:5.0 loss:0.0253
Episode:978 meanR:4.0500 R:3.0 loss:0.0359
Episode

ERROR:root:Exception calling application: Ran out of input
Traceback (most recent call last):
  File "/home/arasdar/anaconda3/envs/env3/lib/python3.6/site-packages/grpc/_server.py", line 385, in _call_behavior
    return behavior(argument, context), True
  File "/home/arasdar/anaconda3/envs/env3/lib/python3.6/site-packages/unityagents/rpc_communicator.py", line 26, in Exchange
    return self.child_conn.recv()
  File "/home/arasdar/anaconda3/envs/env3/lib/python3.6/multiprocessing/connection.py", line 251, in recv
    return _ForkingPickler.loads(buf.getbuffer())
EOFError: Ran out of input


KeyError: 'BananaBrain'

In [None]:
%matplotlib inline
import matplotlib.pyplot as plt

def running_mean(x, N):
    cumsum = np.cumsum(np.insert(x, 0, 0)) 
    return (cumsum[N:] - cumsum[:-N]) / N 

In [None]:
eps, arr = np.array(episode_rewards_list).T
smoothed_arr = running_mean(arr, 10)
plt.plot(eps[-len(smoothed_arr):], smoothed_arr)
plt.plot(eps, arr, color='grey', alpha=0.3)
plt.xlabel('Episode')
plt.ylabel('Episode rewards')

In [None]:
eps, arr = np.array(rewards_list).T
smoothed_arr = running_mean(arr, 10)
plt.plot(eps[-len(smoothed_arr):], smoothed_arr)
plt.plot(eps, arr, color='grey', alpha=0.3)
plt.xlabel('Episode')
plt.ylabel('Total rewards')

In [None]:
eps, arr = np.array(loss_list).T
smoothed_arr = running_mean(arr, 10)
plt.plot(eps[-len(smoothed_arr):], smoothed_arr)
plt.plot(eps, arr, color='grey', alpha=0.3)
plt.xlabel('Episode')
plt.ylabel('Average losses')

In [37]:
# TF session for training
with tf.Session(graph=graph) as sess:
    sess.run(tf.global_variables_initializer())
    #saver.restore(sess, 'checkpoints/model.ckpt')    
    saver.restore(sess, tf.train.latest_checkpoint('checkpoints'))
    
    # Testing episodes/epochs
    for _ in range(1):
        total_reward = 0
        #state = env.reset()
        env_info = env.reset(train_mode=False)[brain_name] # reset the environment
        state = env_info.vector_observations[0]   # get the current state

        # Testing steps/batches
        while True:
            action_logits = sess.run(model.actions_logits, feed_dict={model.states: state.reshape([1, -1])})
            action = np.argmax(action_logits)
            #state, reward, done, _ = env.step(action)
            env_info = env.step(action)[brain_name]        # send the action to the environment
            state = env_info.vector_observations[0]   # get the next state
            reward = env_info.rewards[0]                   # get the reward
            done = env_info.local_done[0]                  # see if episode has finished
            total_reward += reward
            if done:
                break
                
        print('total_reward: {:.2f}'.format(total_reward))

INFO:tensorflow:Restoring parameters from checkpoints/model-nav.ckpt


total_reward: 14.00


In [None]:
# Be careful!!!!!!!!!!!!!!!!
# Closing the env
env.close()