Typical adults can track reward probabilities across trials to estimate the volatility of the environment and use this information to modify their learning rate (Behrens et al ., 2007). In a stable environment, it is advantageous to take account of outcomes over many trials, whereas in a volatile…