Exploring Replay-Reference-Cited by-同舟云学术

Exploring Replay

Published:2023-01-28 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Antonov Georgy^ORCID,Dayan Peter^ORCID

Abstract

Exploration is vital for animals and artificial agents who face uncertainty about their environments due to initial ignorance or subsequent changes. Their choices need to balance exploitation of the knowledge already acquired, with exploration to resolve uncertainty [1, 2]. However, the exact algorithmic structure of exploratory choices in the brain still remains largely elusive. A venerable idea in reinforcement learning is that agents can plan appropriate exploratory choices offline, during the equivalent of quiet wakefulness or sleep. Although offline processing in humans and other animals, in the form of hippocampal replay and preplay, has recently been the subject of highly successful modelling [3–5], existing methods only apply to known environments. Thus, they cannot predict exploratory replay choices during learning and/or behaviour in dynamic environments. Here, we extend the theory of Mattar & Daw [3] to examine the potential role of replay in approximately optimal exploration, deriving testable predictions for the patterns of exploratory replay choices in a paradigmatic spatial navigation task. Our modelling provides a normative interpretation of the available experimental data suggestive of exploratory replay. Furthermore, we highlight the importance of sequence replay, and license a range of new experimental paradigms that should further our understanding of offline processing.

Publisher

Cold Spring Harbor Laboratory

Reference39 articles.

1. Planning and Acting in Partially Observable Stochastic Domains;Artificial Intelligence,2021

2. Michael O’Gordon Duff. Optimal Learning: Computational Procedures for Bayes-adaptive Markov Decision Processes. PhD Thesis. https://scholarworks.umass.edu/dissertations/AAI3039353/ (Feb. 2002).

3. Prioritized Memory Access Explains Planning and Hippocampal Replay;Nature Neuroscience,2022

4. Experience Replay Is Associated with Efficient Nonlocal Learning;Science,2022

5. Optimism and Pessimism in Optimised Replay;PLOS Computational Biology,2022

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Risking your Tail: Modeling Individual Differences in Risk-sensitive Exploration using Bayes Adaptive Markov Decision Processes;2024-01-08