Multiple Visual Objects Are Represented Differently in the Human Brain and Convolutional Neural Networks

Overview

Journal Sci Rep

Specialty Science

Date 2023 Jun 5

PMID 37277406

Authors

Viola Mocz

Su Keun Jeong

Marvin Chun

Yaoda Xu

Affiliations

Soon will be listed here.

Abstract

Objects in the real world usually appear with other objects. To form object representations independent of whether or not other objects are encoded concurrently, in the primate brain, responses to an object pair are well approximated by the average responses to each constituent object shown alone. This is found at the single unit level in the slope of response amplitudes of macaque IT neurons to paired and single objects, and at the population level in fMRI voxel response patterns in human ventral object processing regions (e.g., LO). Here, we compare how the human brain and convolutional neural networks (CNNs) represent paired objects. In human LO, we show that averaging exists in both single fMRI voxels and voxel population responses. However, in the higher layers of five CNNs pretrained for object classification varying in architecture, depth and recurrent processing, slope distribution across units and, consequently, averaging at the population level both deviated significantly from the brain data. Object representations thus interact with each other in CNNs when objects are shown together and differ from when objects are shown individually. Such distortions could significantly limit CNNs' ability to generalize object representations formed in different contexts.

Citing Articles

The human posterior parietal cortices orthogonalize the representation of different streams of information concurrently coded in visual working memory.

Xu Y PLoS Biol. 2024; 22(11):e3002915.

PMID: 39570984 PMC: 11620661. DOI: 10.1371/journal.pbio.3002915.

Integrative processing in artificial and biological vision predicts the perceived beauty of natural images.

Nara S, Kaiser D Sci Adv. 2024; 10(9):eadi9294.

PMID: 38427730 PMC: 10906925. DOI: 10.1126/sciadv.adi9294.

References

Baeck A, Wagemans J, Op de Beeck H . The distributed representation of random and meaningful object pairs in human occipitotemporal cortex: the weighted average as a general rule. Neuroimage. 2012; 70:37-47. DOI: 10.1016/j.neuroimage.2012.12.023. View

Kay K . Principles for models of neural information processing. Neuroimage. 2017; 180(Pt A):101-109. DOI: 10.1016/j.neuroimage.2017.08.016. View

Reddy L, Kanwisher N . Category selectivity in the ventral visual pathway confers robustness to clutter and diverted attention. Curr Biol. 2007; 17(23):2067-72. PMC: 2744456. DOI: 10.1016/j.cub.2007.10.043. View

Taylor J, Xu Y . Joint representation of color and form in convolutional neural networks: A stimulus-rich network perspective. PLoS One. 2021; 16(6):e0253442. PMC: 8244861. DOI: 10.1371/journal.pone.0253442. View

Khaligh-Razavi S, Kriegeskorte N . Deep supervised, but not unsupervised, models may explain IT cortical representation. PLoS Comput Biol. 2014; 10(11):e1003915. PMC: 4222664. DOI: 10.1371/journal.pcbi.1003915. View

Yamins D, Hong H, Cadieu C, Solomon E, Seibert D, DiCarlo J . Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proc Natl Acad Sci U S A. 2014; 111(23):8619-24. PMC: 4060707. DOI: 10.1073/pnas.1403112111. View

Tang K, Chin M, Chun M, Xu Y . The contribution of object identity and configuration to scene representation in convolutional neural networks. PLoS One. 2022; 17(6):e0270667. PMC: 9239439. DOI: 10.1371/journal.pone.0270667. View

MacEvoy S, Epstein R . Decoding the representation of multiple simultaneous objects in human occipitotemporal cortex. Curr Biol. 2009; 19(11):943-7. PMC: 2875119. DOI: 10.1016/j.cub.2009.04.020. View

Tacchetti A, Isik L, Poggio T . Invariant Recognition Shapes Neural Representations of Visual Input. Annu Rev Vis Sci. 2018; 4:403-422. DOI: 10.1146/annurev-vision-091517-034103. View

10.

Kamitani Y, Tong F . Decoding the visual and subjective contents of the human brain. Nat Neurosci. 2005; 8(5):679-85. PMC: 1808230. DOI: 10.1038/nn1444. View

11.

Carandini M, Heeger D . Normalization as a canonical neural computation. Nat Rev Neurosci. 2011; 13(1):51-62. PMC: 3273486. DOI: 10.1038/nrn3136. View

12.

Heeger D . Normalization of cell responses in cat striate cortex. Vis Neurosci. 1992; 9(2):181-97. DOI: 10.1017/s0952523800009640. View

13.

Bao P, She L, McGill M, Tsao D . A map of object space in primate inferotemporal cortex. Nature. 2020; 583(7814):103-108. PMC: 8088388. DOI: 10.1038/s41586-020-2350-5. View

14.

Tarhan L, Konkle T . Reliability-based voxel selection. Neuroimage. 2019; 207:116350. DOI: 10.1016/j.neuroimage.2019.116350. View

15.

Dale A, Fischl B, Sereno M . Cortical surface-based analysis. I. Segmentation and surface reconstruction. Neuroimage. 1999; 9(2):179-94. DOI: 10.1006/nimg.1998.0395. View

16.

Jacob G, Pramod R, Katti H, Arun S . Qualitative similarities and differences in visual object representations between brains and deep networks. Nat Commun. 2021; 12(1):1872. PMC: 7994307. DOI: 10.1038/s41467-021-22078-3. View

17.

Epstein R, Kanwisher N . A cortical representation of the local visual environment. Nature. 1998; 392(6676):598-601. DOI: 10.1038/33402. View

18.

DiCarlo J, Cox D . Untangling invariant object recognition. Trends Cogn Sci. 2007; 11(8):333-41. DOI: 10.1016/j.tics.2007.06.010. View

19.

DiCarlo J, Zoccolan D, Rust N . How does the brain solve visual object recognition?. Neuron. 2012; 73(3):415-34. PMC: 3306444. DOI: 10.1016/j.neuron.2012.01.010. View

20.

Baker C, Behrmann M, Olson C . Impact of learning on representation of parts and wholes in monkey inferotemporal cortex. Nat Neurosci. 2002; 5(11):1210-6. DOI: 10.1038/nn960. View