Curiosity Killed the Mario
The ideal reward is the infinite sum:
The model must be fixed for these surprisals to be well-defined, since the states depend on the actions taken. Inverse Dynamics concatenates each state’s feature vector with the subsquent state’s feature vector, feeds that through two Dense layers, and uses a Softmax Cross Entropy to decide which action was taken. The rollout contains 129 states, surprisals, and rewards estimated from the Policy network.
Source: www.michaelburge.us