» Articles » PMID: 33372237

Meta-control of the Exploration-exploitation Dilemma Emerges from Probabilistic Inference over a Hierarchy of Time Scales

Overview
Publisher Springer
Date 2020 Dec 29
PMID 33372237
Citations 5
Authors
Affiliations
Soon will be listed here.
Abstract

Cognitive control is typically understood as a set of mechanisms that enable humans to reach goals that require integrating the consequences of actions over longer time scales. Importantly, using routine behaviour or making choices beneficial only at short time scales would prevent one from attaining these goals. During the past two decades, researchers have proposed various computational cognitive models that successfully account for behaviour related to cognitive control in a wide range of laboratory tasks. As humans operate in a dynamic and uncertain environment, making elaborate plans and integrating experience over multiple time scales is computationally expensive. Importantly, it remains poorly understood how uncertain consequences at different time scales are integrated into adaptive decisions. Here, we pursue the idea that cognitive control can be cast as active inference over a hierarchy of time scales, where inference, i.e., planning, at higher levels of the hierarchy controls inference at lower levels. We introduce the novel concept of meta-control states, which link higher-level beliefs with lower-level policy inference. Specifically, we conceptualize cognitive control as inference over these meta-control states, where solutions to cognitive control dilemmas emerge through surprisal minimisation at different hierarchy levels. We illustrate this concept using the exploration-exploitation dilemma based on a variant of a restless multi-armed bandit task. We demonstrate that beliefs about contexts and meta-control states at a higher level dynamically modulate the balance of exploration and exploitation at the lower level of a single action. Finally, we discuss the generalisation of this meta-control concept to other control dilemmas.

Citing Articles

Post-injury pain and behaviour: a control theory perspective.

Seymour B, Crook R, Chen Z Nat Rev Neurosci. 2023; 24(6):378-392.

PMID: 37165018 PMC: 10465160. DOI: 10.1038/s41583-023-00699-5.


Cognitive effort and active inference.

Parr T, Holmes E, Friston K, Pezzulo G Neuropsychologia. 2023; 184:108562.

PMID: 37080424 PMC: 10636588. DOI: 10.1016/j.neuropsychologia.2023.108562.


The Willpower Paradox: Possible and Impossible Conceptions of Self-Control.

Goschke T, Job V Perspect Psychol Sci. 2023; 18(6):1339-1367.

PMID: 36791675 PMC: 10623621. DOI: 10.1177/17456916221146158.


The exploration-exploitation trade-off in a foraging task is affected by mood-related arousal and valence.

van Dooren R, de Kleijn R, Hommel B, Sjoerds Z Cogn Affect Behav Neurosci. 2021; 21(3):549-560.

PMID: 34086199 PMC: 8208924. DOI: 10.3758/s13415-021-00917-6.


Neural Dynamics under Active Inference: Plausibility and Efficiency of Information Processing.

Da Costa L, Parr T, Sengupta B, Friston K Entropy (Basel). 2021; 23(4).

PMID: 33921298 PMC: 8069154. DOI: 10.3390/e23040454.

References
1.
Daw N, Doya K . The computational neurobiology of learning and reward. Curr Opin Neurobiol. 2006; 16(2):199-204. DOI: 10.1016/j.conb.2006.03.006. View

2.
Garbusow M, Schad D, Sommer C, Junger E, Sebold M, Friedel E . Pavlovian-to-instrumental transfer in alcohol dependence: a pilot study. Neuropsychobiology. 2014; 70(2):111-21. DOI: 10.1159/000363507. View

3.
Friston K . The free-energy principle: a unified brain theory?. Nat Rev Neurosci. 2010; 11(2):127-38. DOI: 10.1038/nrn2787. View

4.
Doya K . Metalearning and neuromodulation. Neural Netw. 2002; 15(4-6):495-506. DOI: 10.1016/s0893-6080(02)00044-8. View

5.
Nassar M, Wilson R, Heasly B, Gold J . An approximately Bayesian delta-rule model explains the dynamics of belief updating in a changing environment. J Neurosci. 2010; 30(37):12366-78. PMC: 2945906. DOI: 10.1523/JNEUROSCI.0822-10.2010. View