Psychology

Toddlers' Active Gaze Behavior Supports Self-Supervised Object Learning.

Yu Z, Aubret A, Raabe MC, Yang J, Yu C, Triesch J. Published July 1, 2026 CC-BY

Toddlers learn to recognize objects from different viewpoints with almost no supervision. During this learning, they execute frequent eye and head movements that shape their visual experience. It is presently unclear if and how these behaviors contribute to toddlers' emerging object recognition abilities. To answer this question, we here combine head-mounted eye tracking during dyadic play with unsupervised machine learning. We approximate toddlers' central visual field experience by cropping image regions from a head-mounted camera centered on the current gaze location estimated via eye tracking. This visual stream feeds a neural network model, which uses a biologically plausible unsupervised learning objective. Our experiments demonstrate that a few minutes of such first-person experience suffice to learn strong object representations permitting invariant object recognition. Importantly, by simulating alternative gaze behaviors we show that toddlers' eye movement patterns play a crucial role in this. Our analysis also reveals that the limited size of the central visual field where visual acuity is high plays an important role for successful learning. Together, this highlights the benefits of temporally structured visual experience arising from toddlers' natural interactions with objects. SUMMARY: We combine recordings of toddlers' first-person central visual field experience with biologically inspired self-supervised learning algorithms to model toddlers' development of invariant object recognition. Just a few minutes of toddlers' central visual field experience captured with head-mounted eye tracking suffice to learn strong object representations. Simulated alternative gaze behaviors produce weaker representations, demonstrating the importance of toddlers' active gaze strategies for learning. Our results emphasize the importance of toddlers' eye movements for learning object representations.

Introduction

Within their first year of life, toddlers learn to robustly recognize objects despite variations in viewpoints, lighting, and so forth (Kraebel and Gerhardstein2006; Ayzenberg and Behrmann2024). This early emergence of invariant object recognition and the ease with which adults perform this skill hide the complexities of acquiring it. For example, retinal images vary drastically when objects are rotated in depth and even state‐of‐the‐art machine learning methods make surprising recognition mistakes when faced with unusual viewpoints of objects (Dong et al.2022; Abbas and Deny2023; Ruan et al.2023).

One of the main theories for how toddlers acquire invariant object recognition posits that their brains exploit a mechanism to construct visual representations that slowly change over time (Földiák1991; Li and DiCarlo2008; Miyashita1988).

The computational principle of slowness learning, particularly as implemented in Slow Feature Analysis (SFA), is based on the assumption that although primary sensory inputs such as retinal pixel intensities change rapidly, the underlying semantic properties of the environment, such as object identity, vary much more slowly. By optimizing a slowness objective, the learning system is encouraged to discard information related to rapidly varying factors such as local illumination or object pose, while retaining temporally invariant features that remain stable over time. This temporal stability provides an unsupervised heuristic that enables the brain to associate different views of the same object into a coherent and viewpoint invariant representation (Franzius et al.2011; Wiskott and Sejnowski2002).

In the context of early childhood development, this computational framework aligns with the maturation of toddlers’ visual, motor, and cognitive systems. Developing toddlers actively structure their sensory environment through eye movements, manual object manipulation, and locomotion. These self‐generated physical interactions shape the temporal structure of visual experiences, and may provide the temporal stability of semantic information required for a slowness‐based learning mechanism to function.

Toddlers abundantly manipulate (or move around) objects while watching them, which gives them access to diverse views of a single object over a short period of time. By learning slowly changing representations, a toddler may be able to associate these different views, allowing them to form viewpoint‐invariant representations of objects (Wiskott and Sejnowski2002; Schneider et al.2021; Aubret et al.2022a).

When processing visual input, humans typically select only a limited portion of the central visual field of a few degrees of visual angle for detailed analysis (Quaia and Krauzlis2024; Yu et al.2015; Zhaoping2024). One reason for this is that receptor densities in the retina decline sharply towards the periphery (Jonas et al.1992; Provis et al.1985). As humans make on average around three saccades per second, the contents of the central visual field may be semantically unstable, that is, the central visual field may contain frequent transitions between different objects. This might interfere with a learning mechanism based on slowness. However, toddlers’ gaze behavior during interaction may stabilize the semantic content in their central visual field and thereby support learning via a slowness objective. For example, a learning toddler may not move their gaze randomly within a scene, but watch an object they are manipulating for an extended period of time before saccading to a different object. Previous models of visual representation learning in infants and toddlers have neglected the importance of eye gaze for learning (Orhan et al.2020; Orhan and Lake2024; Orhan et al.2024). An exception to this is the work of Bambach et al.2018, who, however, used a biologically implausiblesupervisedlearning approach.

Here, we investigate whether toddlers’ gaze behaviors may support theunsupervisedlearning of view‐invariant object representations. To this end, we leverage a dataset of head‐camera video recordings and eye gaze tracking from toddlers and adults during play sessions (Bambach et al.2018). To simulate participants’ central visual field experience, we extract image patches centered on tracked gaze locations. This data feeds a computational model of a toddler's visual representation learning, which constructs representations that slowly change over time (Schneider et al.2021; Aubret et al.2022a). Our results show that toddlers’ gaze strategies boost visual learning in comparison to several baselines. Furthermore, we demonstrate that restricting learning to input from the central visual field improves the emerging object representations. Finally, we show that the visual input from toddlers permits learning better representations than that from adults. This finding is explained by toddlers looking longer at individual objects while manipulating them. Overall, by leveraging head‐mounted eye tracking and a neural network architecture with a biologically motivated unsupervised learning mechanism, our work represents the most advanced and precise model to date of the development of children's invariant object recognition abilities. It reveals how gaze behaviors and the principle of temporal slowness may jointly underpin the development of object recognition in children.

Methods

Figure1provides an overview of the dataset, our computational model of toddlers visual learning, and the evaluation procedure used in this study. In the following sections, we briefly describe each component. More details are supplied in AppendixA.

Overview of the experimental framework. (A) Eye‐tracking data from 38 toddlers and their caregivers were used. The resulting eye movement sequences formed a temporal dataset of visual observations. From each video frame, multiple crops are extracted using three methods to simulate the actual versus alternative central visual field experiences of participants: gaze‐guided crop (purple), centroid crop (orange), and random crop (green). AppendixA.1provides a more detailed description of these three cropping methods. (B) We trained a computational model of biological visual learning to make the representations of temporally adjacent frames more similar. (C) To assess the quality of the learned visual representations, we train a linear classifier on top of the frozen trained neural network to perform object recognition.

Overview of the experimental framework. (A) Eye‐tracking data from 38 toddlers and their caregivers were used. The resulting eye movement sequences formed a temporal dataset of visual observations. From each video frame, multiple crops are extracted using three methods to simulate the actual versus alternative central visual field experiences of participants: gaze‐guided crop (purple), centroid crop (orange), and random crop (green). AppendixA.1provides a more detailed description of these three cropping methods. (B) We trained a computational model of biological visual learning to make the representations of temporally adjacent frames more similar. (C) To assess the quality of the learned visual representations, we train a linear classifier on top of the frozen trained neural network to perform object recognition.

Datasets

We build on a dataset containing head‐camera videos and eye‐tracking data recorded from 38 dyads of toddlers and caregivers. Each video shows a dyad that plays with the same 24 toys for 15 min on average. The toddler participants had a mean age of 18.32 months (SD = 3.06 months, range = 12.3–24.3 months). To simulate the central visual field experience of a participant based on their eye gaze, we extract image patches (corresponding to 14°×14° of visual angle) from video frames of the scene camera mounted on the participant's head, which are centered on the recorded gaze position. FigureA1Ashows an example of a sequence of image patches extracted around subsequent gaze locations (Bambach et al.2018). We compare learning based on the visual input streams produced by these measured gaze behaviors (“Toddlers’/Adults’ eye movements”) of dyads against learning based on two simulated alternative gaze strategies. The first assumes that the camera‐wearer samples gaze locations uniformly at random within the field of view of the scene camera (“Random eye movements”). The second ignores eye movements (as in previous works) and instead assumes that the gaze location always remains in the center of the participant's field of view (“No eye movements”). In addition, we consider two idealized “oracles” that simulate a (biologically implausible) learner who fixates only on the objects and does not learn during transitions between objects. That is, this hypothetical learner already has perfect knowledge of when it is looking at an object of interest. We consider this strategy with natural backgrounds (“Object fixation”) and with blank backgrounds (“Blank background”). We show examples of visual sequences resulting from these strategies in FigureA1.

Computational Model

To model the learning process of humans, we train deep neural networks with two bio‐inspired self‐supervised learning models, namely SimCLR‐TT and BYOL‐TT (Schneider et al.2021). Both models learn visual representations that associate close‐in‐time visual inputs. They also include a hyper‐parameter ∆T(measured in seconds) that quantifies how slowly representations should change. We use a ResNet18 as our default neural net‐ work architecture and provide additional results with a ResNet50 in AppendixB.4. Networks are trained “from scratch,” that is, they are not already pretrained for object recognition or any other task. FigureA1Billustrates the learning process. For each gaze dataset, a separate model was trained while keeping the neural network architecture, optimization procedure, and hyperparameter settings identical across conditions. Thus, differences in performance reflect differences in the visual input streams rather than differences in model architecture or training settings.

Evaluation

We train the models with video frames recorded from 30 randomly chosen dyads and keep the other 8 for testing. Even though recording durations varied substantially across toddlers, during training we pooled data across all toddlers’ videos rather than applying explicit normalization, subsampling, or weighting procedures. This choice was motivated by the goal of leveraging the full amount of available visual experience for model training.

We also consider training on the recording of single participants. We assess the quality of the learned representations by training a linear classifier on top of the learned representation (right after the average pooling layer) in a supervised fashion (Chen et al.2020). Since our model of human visual representation learning does not use labeled images, we always train the linear classifier on the train split of the Objects Fixation dataset, which was manually labeled, and evaluate the object recognition accuracy on the test split of the Objects Fixation dataset. Statistical analyses were performed to assess the results. Unless otherwise specified, all significance tests were performed using two‐tailed independent two‐sample t‐tests. Correlation analyses were conducted using Pearson's correlation coefficient.

Results

Toddlers’ Central Visual Field Experience Supports the Learning of Invariant Object Representations via Time‐Based Self‐Supervised Learning

To test whether toddlers’ gaze behavior supports the learning of strong object representations, we compare the representations learned by the two unsupervised learning models based on slowness (see AppendixA.2for details of SimCLR‐TT and BYOL‐ TT) when trained on the different datasets introduced in Section2.1. Figure2A and Bshow that models trained with the Toddlers’ Eye movements dataset outperform those trained with the Random Eye movements dataset or the No Eye movements dataset (t‐tests,p <0.05 in all cases). This suggests that toddlers’ gaze behavior supports the learning of view‐invariant object representations. A similar result is observed for adults (AppendixB.3).

Toddlers’ gaze behavior supports the learning of invariant object representations. (A) and (B) show object recognition results of the linear classifier trained on top of representations learned with SimCLR‐TT and BYOL‐TT, respectively. Error bars represent the standard deviation over three random seeds. When training on data from just one toddler (“Single toddlers’ eye mov.”), we show the average accuracy and the standard deviation over the 38 toddlers. Results of statistical tests com‐ paring toddlers’ eye movements to other conditions are indicated as follows: “”:p‐value<0.001, “”:p‐value<0.01, and “”:p‐value<0.05. (C–G)t‐SNE visualizations of learned object representations for SimCLR‐TT trained on different datasets. Each point represents an image sample, color‐coded by object identity (24 in total). The number in the bottom‐right corner of each panel displays the corresponding inter/intra‐cluster distance ratio computed in the full latent space at the end of training (higher is better, see AppendixB.2).

Toddlers’ gaze behavior supports the learning of invariant object representations. (A) and (B) show object recognition results of the linear classifier trained on top of representations learned with SimCLR‐TT and BYOL‐TT, respectively. Error bars represent the standard deviation over three random seeds. When training on data from just one toddler (“Single toddlers’ eye mov.”), we show the average accuracy and the standard deviation over the 38 toddlers. Results of statistical tests com‐ paring toddlers’ eye movements to other conditions are indicated as follows: “”:p‐value<0.001, “”:p‐value<0.01, and “”:p‐value<0.05. (C–G)t‐SNE visualizations of learned object representations for SimCLR‐TT trained on different datasets. Each point represents an image sample, color‐coded by object identity (24 in total). The number in the bottom‐right corner of each panel displays the corresponding inter/intra‐cluster distance ratio computed in the full latent space at the end of training (higher is better, see AppendixB.2).

Next, we study the impact of toddlers’ gaze behaviors on the learnt visual representations at a qualitative level. For each model trained using different eye movement strategies, we extract their representations of images in the Objects Fixation dataset and project these representations into a 2‐dimensional embedding space using t‐ SNE (Maaten and Hinton2008). This allows us to visualize the (dis)similarity between representations of views of different objects (Figure2C–G). The model trained on the Toddlers’ Eye movements dataset shows well‐separated clusters, indicating strong and discriminative object representations. In contrast, models trained with the Random or No Eye movements datasets produce less distinct clusters, reflecting weaker object learning. In contrast, the two biologically implausible “oracle” methods that train with the Objects Fixation dataset (F) or the Blank Background dataset (G) exhibit improved clustering.

Furthermore, we quantitatively assess the representation quality in the full latent space across conditions using inter/intra‐cluster distance ratios (higher is better; see bottom‐right of Figure2C–G). AppendixB.2explains how the inter/intra‐cluster distance ratio is calculated and how it evolves throughout training. Notably, the toddlers’ eye movements condition achieves the highest ratios among all eye movement strategies. This indicates that toddlers’ visual experience supports compact, well‐ separated representations of the different objects, whereas random eye movements yield lower ratios and a less clustered representational structure. The highest ratios are observed under the idealized but unnatural blank background and object fixation conditions.

We wondered whether the visual experience of onlya singletoddler during a play session suffices to build good visual representations. To investigate this question, we train SimCLR‐TT on the individual recordings of each toddler and compute the average of linear accuracies. We train the neural network using all fixation data from a single toddler, followed by training and testing the linear classifier with the Objects fixation data from the same and different toddlers. We control the training set to comprise 75% of the total data, ensuring that the test set does not overlap with the training set. Figure2A and Bshow that the central visual experience of one toddler leads to representations almost as good as those resulting from the central visual experience of all toddlers.

Constraining Input to the Central Visual Field Improves Learning

Previous computational studies of learning from infants’ first person visual experience have neglected the importance of the constrained size of the central visual field for learning visual representations (Orhan et al.2020; Sheybani et al.2024). Here, we assess whether our simulated central visual field experience leads to better/worse object representations than learning from a wide field of view. We vary the size of the image portion extracted to simulate central vision. In Table1, we observe for both toddlers and adults that an image size of 128 × 128 (corresponding to 14°× 14 of visual angle) produces the best recognition accuracy for all gaze strategies. Importantly, results for toddlers’ and adults’ eye movements at an image size of 128 × 128 pixels present an accuracy boost of 8% compared to the no eye movements condition for an image size of 480×480 pixels, which simulates head‐camera recordings without eye‐tracking information. We conclude that accounting for the constrained size of the central visual field is crucial for learning powerful object representations. We speculate that this boost results from the property of a 128 × 128 gaze‐centered crop to frequently capture the complete structure of an object while minimizing irrelevant background information, at least in the present dataset. However, the optimal crop size likely depends on additional factors such as object size, viewing distance, scene depth, and the amount of surrounding background entering the crop. These factors were not explicitly controlled in the present study and may influence how much object structure versus contextual information is preserved within the gaze‐centered image patches.

Table: Linear object recognition accuracy for different cropping sizes.

Toddlers’ Gaze Behavior Favors Stronger Emphasis on Slowness

Semantic aspects of the visual experience have been shown to vary more slowly for toddlers than for adults (Sheybani et al.2023). Our learning model includes a hyper‐parameter ∆T(measured in seconds) that quantifies how slowly representations should change. Concretely, it specifies the (maximum) time interval between two positive input pairs (whose representations will be made more similar) in the learning algorithms (see AppendixA.2for the definition and implementation of ∆T). Interestingly, previous work has shown that increasing ∆Tcan improve the quality of object representations if visual inputs are sufficiently stable over time (Schneider et al.2021; Aubret et al.2022a). Thus, we wondered how changing ∆Tmay affect results and whether it may amplify any differences in the quality of representations learned from toddlers’ versus adults’ first person visual experience To test this, we varied ∆Tfrom 1/30 to 3.0 s (Figure3A). The models trained with the Toddlers' Eye movements dataset (red) achieve the highest recognition accuracy for an inter‐ mediate value of ∆T= 1.5s. In contrast, Figure3Bshows that, for models trained with the Adults’ Eye movements dataset, increasing ∆Tonly decreases the quality of the learned object representations. The results are consistent for both human eye movements (red) and the no eye movements datasets (blue). These results may reflect differences in gaze between toddlers and adults. Adults’ long gaze shifts and short fixation durations are more likely to bring their gaze to a different object within a few seconds. But making the representations of these views of different objects more similar is likely to make the resulting representation less good at object recognition, whereas toddlers’ longer object‐centered inspection periods may provide more tempo‐ rally stable object views that better support slowness‐based learning across larger ∆T(Franchak et al.2016). We conclude that toddlers’ gaze behavior favors a stronger emphasis on slowness (greater ∆T) than that of adults.

The impact of different ∆Ton recognition accuracy for (A) toddlers and (B) adults. Error bars represent the standard deviation over three random seeds.

The impact of different ∆Ton recognition accuracy for (A) toddlers and (B) adults. Error bars represent the standard deviation over three random seeds.

Toddlers’ Long Object Inspections Relative to Adults Facilitate Learning

So far, we have shown that the gaze behaviors of humans support object learning via an unsupervised slowness objective and that toddlers gaze behavior favors a stronger emphasis on slowness (greater ∆T). We wondered what differences between toddlers’ and adults’ eye movements may contribute to this discrepancy. We analyzed four metrics that characterize the temporal sequence of images: the average fixation duration, the average duration of bouts of looking at the same object when not holding the object, the average duration of looking at an object when holding it, and the cumulative duration of object looking in an entire recording session. For these analyses, we leverage manually labeled timestamps (by Bambach et al.2018) about when toddlers and adults look at/hold an object. See AppendixA.3for details on saccade detection and calculation of fixation durations.

We successfully extracted the data from 28 out of 38 toddlers and conducted all subsequent experiments using these 28 toddlers. The remaining participants are excluded from this analysis due to the lack of data on fixation durations. TableD3in AppendixDpresents the details of the 28 included toddlers.

In Figure4, we observe that object recognition accuracy is highly correlated with fixation durations, durations of object looking and duration of object looking while holding the object, but only weakly correlated with the cumulative duration of object looking. This indicates that long fixation bouts on the same object are important in explaining the relative quality of visual representations trained on Toddlers’ and Adults’ Eye movements datasets.

Correlation analysis between the recognition accuracy and (A) the average fixation duration; (B) the average duration of object looking while not holding the object; (C) the average duration of object looking while holding the object and (D) the cumulative duration of object looking. Models were all trained on the individual Toddlers’ and Adults’ Eye movements dataset. In each figure, the crosshairs represent the mean and standard deviation of the data values over the two axes. The legends show the Pearson correlation coefficients and their p‐values.

Correlation analysis between the recognition accuracy and (A) the average fixation duration; (B) the average duration of object looking while not holding the object; (C) the average duration of object looking while holding the object and (D) the cumulative duration of object looking. Models were all trained on the individual Toddlers’ and Adults’ Eye movements dataset. In each figure, the crosshairs represent the mean and standard deviation of the data values over the two axes. The legends show the Pearson correlation coefficients and their p‐values.

Figure4also shows that on average toddlers’ visual experience permits learning better representations than that of adults (t‐tests,p= 0.0053) with the given subset of dyads. To investigate which metric plays a crucial role in these differences. Figure5compares the distributions of average fixation durations for toddlers and adults. The t‐test statistics and p‐values are given in the titles. We observe that toddlers look longer at the object that they are holding (t‐test,p= 0.003). Other metrics do not exhibit statistically significant differences between adults and toddlers. We conclude that, compared to adults, toddlers’ longer periods of object inspection while manipulating objects allow for learning better viewpoint‐invariant object representations. More broadly, these findings raise questions about how embodied interaction structures temporal continuity in visual experience and how such structure may interact with slowness‐based learning objectives (see AppendixCfor additional discussion).

Comparison of (A) average fixation duration, (B) average duration of object looking while not holding the object, and (C) average duration of object looking while holding the object for toddlers and adults. Each panel includes the frequency distribution for the given metric, along with a density curve.

Comparison of (A) average fixation duration, (B) average duration of object looking while not holding the object, and (C) average duration of object looking while holding the object for toddlers and adults. Each panel includes the frequency distribution for the given metric, along with a density curve.

Discussion

The mechanisms permitting infants and toddlers to acquire sophisticated object recognition abilities from very limited first person visual experience are still poorly understood. Here, we investigated whether neuro‐biologically motivated models of visual learning can take advantage of toddlers’ gaze behavior to develop robust object representations. We extracted toddlers’ gaze locations from egocentric video recordings with head‐mounted eye‐tracking during play sessions. By cropping image regions around these gaze locations we reconstructed the moment‐to‐moment central visual field experience of the toddlers. We used this data to train self‐supervised deep learning models with a biologically inspired slowness objective.

Our findings indicate that toddlers’ gaze strategies permit the learning of representations that support view‐invariant object instance recognition within a single play session of 12 min. Interestingly, models trained on adults’ visual experience performed significantly worse. Our analysis showed that focusing learning on inputs from the central visual field is beneficial for learning viewpoint‐invariant object representations and that toddlers’ gaze behavior favors a stronger emphasis on slowness compared to that of adults. This is consistent with toddlers looking longer at objects while holding them. During these relatively long holding periods, toddlers often turn and move the object, giving them access to sequences of different views of an object over a short period of time.

By leveraging head‐mounted eye tracking and a biologically motivated learning mechanism, our model captures the development of children's invariant object recognition abilities more precisely than previous works. For example, a study by Orhan & Lake used video data from a head‐mounted camera from a single child, but the data used for learning were sparse (only one video frame per minute) and without any eye tracking information (Orhan et al.2020). An earlier study by Bambach et al. used head‐mounted eye tracking, but the model relied on a biologically implausible supervised learning approach (Bambach et al.2018). Our work is unique and goes beyond such earlier approaches in that it attempts to accurately model infants’ learning from the moment to moment first per‐ son central visual field experience to the biologically inspired self‐supervised learning objective.

From a developmental perspective, our work provides strong evidence that the development of viewpoint‐invariant representations can originate from a slowness learning objective, a mechanism supported by neuroscientific studies (Li and DiCarlo2008; Miyashita1988). Our results also suggest that toddlers’ gaze behavior during naturalistic interaction generates temporally structured visual experience that supports learning of visual representations.

Furthermore, engaging with these computational principles allows us to address established developmental accounts of early object recognition. Developmental literature suggests that younger toddlers, often under 18 months, rely heavily on part‐ and feature‐based representations (Rakison2003; Pereira and Smith2009), whereas older toddlers, between 18 and 24 months, increasingly exhibit holistic, viewpoint‐invariant object representations (Smith2003). We hypothesize that this representational transition may be functionally driven by systematic changes in infants’ own embodied behaviors over development. In early infancy, visual experience is relatively passive and fragmented. However, as toddlers mature, they develop more sophisticated object manipulation skills, such as holding, turning, and dynamically rotating objects closer to their eyes. This developmental milestone in active, coordinated hand‐eye behavior permits learning how, for example, controlled rotations of an object change its visual appearance. Under this view, the transition from feature‐based to holistic recognition observed in developmental psychology may not just be a matter of intrinsic neural maturation, but a direct consequence of the changing statistics of the sensory inputs curated by toddlers’ developing motor repertoires. In this work, we computationally test this hypothesis by demonstrating how such self‐generated, temporally continuous streams enable a self‐ supervised learning system with a slowness objective to bridge local, isolated features and bootstrap robust, holistic representations of object identity.

From a machine learning perspective, we show that combining head‐mounted eye‐ tracking video data with time‐based self‐supervised learning supports the emergence of viewpoint‐invariant object recognition. Our work therefore marks a step toward learning strong visual representations without handcrafted image datasets (e.g., Aubret et al.2022a).

Our work also has several limitations. Because the adult recordings were obtained from caregivers engaged in dyadic interaction with toddlers, adults’ gaze behavior in our dataset may differ from adult visual exploration in more independently structured or non‐social object‐interaction settings. Thus, the present adult‐toddler comparisons should be interpreted within the specific context of caregiver‐child interaction.

Moreover, refining our approach to utilize both central and peripheral vision for learning visual representations could provide a more accurate simulation of human visual physiology and of the development of object versus scene representations in different brain areas. Our current cropping procedure represents a simplified approximation of foveated vision, since it completely removes peripheral information rather than preserving it at reduced spatial resolution. Peripheral vision nevertheless carries coarse but potentially informative visual signals (Quaia and Krauzlis2024; Yu et al.2015; Zhaoping2024). Future work could therefore explore more biologically realistic approximations of foveated vision, such as graded peripheral blurring or continuous foveated representations, similar to approaches explored in previous computational studies of biologically inspired vision and gaze behavior (Wang et al.2021).

We analyzed gaze behavior of toddlers with a minimum age of 12.3 months, meaning they had substantial visual learning experience before the experiment. In contrast, our computational models learned from scratch. Expanding to younger toddlers and a more diverse visual diet and distinct visual exploration patterns, could offer further insights into early visual representation development. Studying how babies under one year engage with objects may reveal new aspects of gaze behavior that contribute to their visual learning (Sheybani et al.2024; Maurer2017). Such attempts should also be guided by knowledge of infants’ developing contrast sensitivity. Regarding the learning mechanism, while the slowness objective used in our model is biologically plausible, the detailed implementation via backpropagation of errors in a deep neural network is not. It is an interesting challenge to replace this mechanism with a more biologically realistic alternative. Finally, a more complete model of visual learning in infants and toddlers needs to also capture the computational mechanisms driving their gaze shifts, whose relevance for visual representation learning we have demonstrated here. Understanding these mechanisms will be an important next step in unraveling the mechanisms underlying the early development of human visual perception.

Finally, We Note that our focus on slowness learning was motivated by the present data and research question: the learning needs to extract relatively stable object representations from continuously changing viewpoints during naturalistic interaction. Toddlers’ visual experience during object manipulation often contains temporally stable object‐centered views over multi‐second intervals, making slowness‐based learning particularly well matched to the temporal statistics of the dataset.

At the same time, slowness learning is related to alternative temporal learning frameworks such as predictive coding and prediction‐based learning (Földiák1991; Rao and Ballard1999). However, the two approaches emphasize somewhat different objectives. Predictive coding and other prediction‐based learning approaches aim to generate accurate predictions of future sensory input, whereas slowness learning emphasizes suppressing rapidly changing fluctuations in appearance in order to extract stable latent structure (Földiák1991; Rao and Ballard1999; Lotter et al.2016). These mechanisms may therefore play complementary roles in biological vision systems and may contribute differently across ventral and dorsal processing streams.

The ventral stream, associated with object identity and recognition, may benefit more from a slowness objective that abstracts away transient appearance changes. The dorsal stream, associated with spatial processing and motion analysis, may instead rely more heavily on precise prediction of visual dynamics (Goodale and Milner1992). Future work incorporating both objectives within a unified model could provide a more complete account of the development of visual representation across these distinct cortical pathways.

Author Contributions

Arthur Aubret: conceptualization, investigation, writing – original draft, methodology, visualization, writing – review and editing, software, supervision, formal analysis.Zhengyang Yu: conceptualization, investigation, writing – original draft, writing – review and editing, visualization, methodology, software, formal analysis, validation.Marcel C. Raabe: conceptualization, investigation, writing – original draft, methodology, visualization, writing – review and editing, software, formal analysis, validation.Jochen Triesch: funding acquisition, writing – original draft, supervision, conceptualization, writing – review and editing, project administration.Chen Yu: conceptualization, funding acquisition, writing – original draft, writing – review and editing, project administration, supervision.Jane Yang: data curation, methodology, writing – review and editing.

Ethics Statement

All procedures involving the collection of data from toddlers and their caregivers were conducted at the Developing Intelligence Lab in the Department of Psychology at the University of Texas at Austin, and the collection was reviewed and approved by the Institutional Review Board at their institution.

Conflicts of Interest

The authors declare no conflicts of interest.

Permission to Reproduce Material From Other Sources

No copyrighted material from other sources was used in this manuscript.

References

  1. Abbas, A. , andS. Deny. 2023. “Progress and Limitations of Deep Networks to Recognize Objects in Unusual Poses. ” InProceedings of the AAAI Conference on Artificial Intelligence, vol. 37: 160–168.
  2. Aubret, A. , M. Ernst, C. Teulière, andJ. Triesch. 2022a. “Time to Augment Self‐Supervised Visual Representation Learning. ” InThe Eleventh International Conference on Learning Representations (ICLR). .
  3. Aubret, A. , C. Teulièr, andJ. Triesch. 2022b. “Toddler‐Inspired Embodied Vision for Learning Object Representations. ” In2022 IEEE International Conference on Development and Learning (ICDL), 81–87. IEEE.
  4. Aubret, A. , C. Teulière, andJ. Triesch. 2024. “Self‐Supervised Visual Learning From Interactions With Objects. ” InEuropean Conference on Computer Vision (ECCV), 54–71.
  5. Ayzenberg, V. , andM. Behrmann. 2024. “Development of Visual Object Recognition. ”Nature Reviews Psychology3, no. 2: 73–90.
  6. Bambach, S. , D. Crandall, L. Smith, andC. Yu. 2018. “Toddler‐Inspired Visual Object Learning. ” InProceedings of the 32nd International Conference on Neural Information Processing Systems (NeurlPS), 1209–1218.
  7. Chen, T. , S. Kornblith, M. Norouzi, andG. Hinton. 2020. “A Simple Framework for Contrastive Learning of Visual Representations. ” InInternational Conference on Machine Learning, 1597–1607. PMLR.
  8. Dong, Y. , S. Ruan, H. Su, C. Kang, X. Wei, andJ. Zhu. 2022. “Viewfool: Evaluating the Robustness of Visual Recognition to Adversarial Viewpoints. ”Advances in Neural Information Processing Systems35: 36789–36803.
  9. Földiák, P. 1991. “Learning Invariance From Transformation Sequences. ”Neural Computation3, no. 2: 194–200. doi.org/10.1162/neco.1991.3.2.194
  10. Franchak, J. M. , D. J. Heeger, U. Hasson, andK. E. Adolph. 2016. “Free Viewing Gaze Behavior in Infants and Adults. ”Infancy21, no. 3: 262–287. doi.org/10.1111/infa.12119
  11. Franzius, M. , N. Wilbert, andL. Wiskott. 2011. “Invariant Object Recognition and Pose Estimation with Slow Feature Analysis. ”Neural Computation23, no. 9: 2289–2323. doi.org/10.1162/NECO_a_00171
  12. Goodale, M. A. , andA. D. Milner. 1992. “Separate Visual Pathways for Perception and Action. ”Trends in Neurosciences15, no. 1: 20–25. doi.org/10.1016/0166-2236(92)90344-8
  13. Jonas, J. B. , U. Schneider, andG. O. Naumann. 1992. “Count and Density of Human Retinal Photoreceptors. ”Graefe's Archive for Clinical and Experimental Ophthalmology230, no. 6: 505–510. doi.org/10.1007/BF00181769
  14. Kraebel, K. S. , andP. C. Gerhardstein. 2006. “Three‐Month‐Old Infants' Object Recognition Across Changes in Viewpoint Using an Operant Learning Procedure. ”Infant Behavior and Development29, no. 1: 11–23. doi.org/10.1016/j.infbeh.2005.10.002
  15. Li, N. , andJ. J. DiCarlo. 2008. “Unsupervised Natural Experience Rapidly Alters Invariant Object Representation in Visual Cortex. ”Science321, no. 5895: 1502–1507. doi.org/10.1126/science.1160028
  16. Lotter, W. , G. Kreiman, andD. Cox. 2016. “Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning. ” InThe Fifth International Conference on Learning Representations (ICLR). .
  17. Maaten, L. , andG. Hinton. 2008. “Visualizing Data Using T‐SNE. ”Journal of Machine Learning Research9, no. 11: 2579–2605.
  18. Maurer, D. 2017. “Critical Periods Re‐Examined: Evidence From Children Treated for Dense Cataracts. ”Cognitive Development42: 27–36.
  19. Miyashita, Y. 1988. “Neuronal Correlate of Visual Associative Long‐Term Memory in the Primate Temporal Cortex. ”Nature335, no. 6193: 817–820. doi.org/10.1038/335817a0
  20. Orhan, A. E. , andB. M. Lake. 2024. “Learning High‐Level Visual Representations From a Child's Perspective Without Strong Inductive Biases. ”Nature Machine Intelligence6, no. 3: 271–283.
  21. Orhan, A. E. , W. Wang, A. N. Wang, M. Ren, andB. M. Lake. 2024. “Self‐Supervised Learning of Video Representations From a Child's Perspective. ” InProceedings of the Annual Meeting of the Cognitive Science Society, 46 (CogSci), 363–369.
  22. Orhan, E. , V. Gupta, andB. M. Lake. 2020. “Self‐Supervised Learning Through the Eyes of a Child. ”Advances in Neural Information Processing Systems33: 9960–9971.
  23. Pereira, A. F. , andL. B. Smith. 2009. “Developmental Changes in Visual Object Recognition Between 18 and 24 Months of Age. ”Developmental Science12, no. 1: 67–80. doi.org/10.1111/j.1467-7687.2008.00747.x
  24. Provis, J. M. , D. Van Driel, F. A. Billson, andP. Russell. 1985. “Development of the Human Retina: Patterns of Cell Distribution and Redistribution in the Ganglion Cell Layer. ”Journal of Comparative Neurology233, no. 4: 429–451. doi.org/10.1002/cne.902330403
  25. Quaia, C. , andR. J. Krauzlis. 2024. “Object Recognition in Primates: What Can Early Visual Areas Contribute?. ”Frontiers in Behavioral Neuroscience18: 1425496. doi.org/10.3389/fnbeh.2024.1425496
  26. Raabe, M. C. , F. M. López, Z. Yu, et al. 2023. “Saccade Amplitude Statistics are Explained by Cortical Magnification. ” In2023 IEEE International Conference on Development and Learning (ICDL), 300–305. IEEE.
  27. Rakison, D. H. 2003. “Seven Parts, Motion, and the Development of the Animate‐Inanimate Distinction in Infancy. ” InEarly Category and Concept Development: Making Sense of the Blooming, Buzzing Confusion, edited byD. H. Rakison, andL. M. Oakes, 159–192, Oxford University Press.
  28. Rao, R. P. , andD. H. Ballard. 1999. “Predictive Coding in the Visual Cortex: A Functional Interpretation of Some Extra‐Classical Receptive‐Field Effects. ”Nature Neuroscience2, no. 1: 79–87. doi.org/10.1038/4580
  29. Ruan, S. , Y. Dong, H. Su, J. Peng, N. Chen, andX. Wei. 2023. “Towards Viewpoint‐ Invariant Visual Recognition via Adversarial Training. ” InProceedings of the IEEE/CVF International Conference on Computer Vision, 4709–4719.
  30. Schneider, F. , X. Xu, M. R. Ernst, Z. Yu, andJ. Triesch. 2021. “Contrastive Learning Through Time. ” InSVRHM 2021 Workshop NeurIPS.
  31. Sheybani, S. , H. Hansaria, J. Wood, L. Smith, andZ. Tiganj. 2024. “Curriculum Learning With Infant Egocentric Videos. ”Advances in Neural Information Processing Systems36: 54199–54212.
  32. Sheybani, S. , Z. Tiganj, J. N. Wood, andL. B. Smith. 2023. “Slow Change: An Analysis of Infant Egocentric Visual Experience. ”Journal of Vision23, no. 9: 4685–4685.
  33. Smith, L. B. : 2003. “Learning to Recognize Objects. ”Psychological Science14, no. 3: 244–250. doi.org/10.1111/1467-9280.03439
  34. Tsutsui, S. , D. Crandall, andC. Yu. 2021. “Reverse‐Engineer the Distributional Structure of Infant Egocentric Views for Training Generalizable Image Classifiers. ” InEPIC 2021 Workshop CVPR.
  35. Wang, B. , D. Mayo, A. Deza, A. Barbu, andC. Conwell. 2021. “On the Use of Cortical Magnification and Saccades as Biological Proxies for Data Augmentation. ” InSVRHM 2021 Workshop NeurIPS.
  36. Wiskott, L. , andT. J. Sejnowski. 2002. “Slow Feature Analysis: Unsupervised Learning of Invariances. ”Neural Computation14, no. 4: 715–770. doi.org/10.1162/089976602317318938
  37. Xu, X. , andJ. Triesch. 2023. “Ciper: Combining Invariant and Equivariant Representations Using Contrastive and Predictive Learning. ” InInternational Conference on Artificial Neural Networks, 320–331. Springer.
  38. Yu, C. , andL. B. Smith. 2012. “Embodied Attention and Word Learning by Toddlers. ”Cognition125, no. 2: 244–262. doi.org/10.1016/j.cognition.2012.06.016
  39. Yu, H. ‐H. , T. Chaplin, andM. Rosa. 2015. “Representation of Central and Peripheral Vision in the Primate Cerebral Cortex: Insights From Studies of the Marmoset Brain. ”Neuroscience Research93: 47–61. doi.org/10.1016/j.neures.2014.09.004
  40. Zhaoping, L. 2024. “Peripheral Vision is Mainly for Looking Rather Than Seeing. ”Neuroscience Research201: 18–26. doi.org/10.1016/j.neures.2023.11.006

Republished from the open web under CC-BY. Authors: Yu Z, Aubret A, Raabe MC, Yang J, Yu C, Triesch J. Read the original.

0 comments

Sign in to join the discussion