RL agents win by arriving, not deciding
A new evaluation method splits RL agent gains into two parts: reaching good states versus solving from them.
A new evaluation method splits RL agent gains into two parts: reaching good states versus solving from them.