This is a brief outline of the approach. Please feel free to comment and/or add stuff.
State space:
Possible actions/subgoals:
Now we try to do Q-learning on these state-action pairs.
There was an error while loading. Please reload this page.