Why the action parameter is optimized using the advantage_svo instead of advantage_action?
Why the action parameter is optimized using the advantage_svo instead of advantage_action?