Info Bound on World Models in Optimal Policies
New ArXiv paper quantifies information optimal policies encode about environments. Proves mutual information of exactly n log m bits in Controlled Markov Processes. Bound holds for finite-horizon, discounted, and average reward maximization.