
Multimodal AI: Top-Down or Bottom-Up?
DIC/ISC/CRIA Seminar in Cognitive Informatics — Université du Québec à Montréal (UQÀM), Autumn 2026 & Spring 2027
Click any title below to jump to that speaker’s abstract. For the Zoom link, please email Stevan Harnad — it isn’t posted publicly on this page.
Autumn 2026
Spring 2027
Autumn Talks — Details
Le pouvoir recombinant du langage
Tout dictionnaire complet peut être représenté comme un graphe définitionnel, où chaque mot-contenu (nom, verbe, adjectif, adverbe) renvoie, par sa définition, à d’autres mots-contenu. Nos travaux antérieurs ont montré que chaque graphe dictionnairique possède un « MinSet » : le plus petit ensemble de mots-contenu dont le référent doit être appris directement — par les sens et l’action — à partir duquel tous les autres mots-contenu peuvent être définis de façon purement récursive. Ce MinSet s’est révélé étonnamment petit : environ 1 % des mots-contenu. Cet exposé présente une généralisation de ce résultat : au lieu de ne considérer que le premier sens de chaque mot, nous désambiguïsons l’ensemble des sens de chaque mot-contenu, dans plusieurs dictionnaires anglais ainsi que dans de nouveaux dictionnaires en plusieurs langues. Bien que le nombre total de mots-contenu augmente considérablement une fois les sens désambiguïsés, la taille du MinSet ne croît que modérément, et ce profil se retrouve à travers les dictionnaires et les langues. Ce résultat renforce l’idée qu’un très petit noyau de mots ancrés sensorimoteurs suffit à engendrer, par recombinaison seule, la totalité du lexique — avec des implications pour l’apprentissage du langage chez l’enfant, la pédagogie, et les limites des grands modèles de langage, qui n’ont accès qu’à la recombinaison, jamais à l’ancrage direct.
Bio: Nicolas Goulet est doctorant en sciences des données à HEC et Mila. Il a étudié les bases neurales du comportement avant d’entamer une maîtrise en informatique et IA à l’UQAM dirigé par É. Harnad et A. Blondin Massé. Dirigées par T. Maharaj au Errata Lab, ses recherches doctorales lient l’apprentissage du langage, les fondements mathématiques de la perception catégorielle dans les LLMs et l’ancrage des symboles. Il s’intéresse aussi à la neuroscience, la sentience, et divers aspects de l’informatique-mathématique.
References:
- Navigli, R. (2026). Is word sense disambiguation dead in the LLM era? Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, No. 46, pp. 39753–39762).
- Meconi, D., Stirpe, S., Martelli, F., Lavalle, L., & Navigli, R. (2025). Do large language models understand word senses? In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 33885–33904).
- Goulet, N., Massé, A. B., & Abdendi, M. (2025). Approaching the Source of Symbol Grounding with Confluent Reductions of Abstract Meaning Representation Directed Graphs. arXiv preprint arXiv:2508.11068.
- Vincent-Lamarre, P., Massé, A. B., Lopes, M., Lord, M., Marcotte, O., & Harnad, S. (2016). The latent structure of dictionaries. Topics in Cognitive Science, 8(3), 625–659.
Mechanistic Emergence of Grounding
What does it mean for a language model to actually ground a word in the world? Much of the current discussion treats grounding as an observed correspondence: a word like “horse” aligns with the right image region, so the model appears grounded. But correlation alone leaves a deeper question unanswered: how does this connection arise during learning, and what inside the model actually implements it? In this talk, I will approach grounding as a process rather than a property of a finished model. Starting from a minimal setting inspired by child language learning, we trace how models learn to connect linguistic symbols with corresponding information from the environment. Interestingly, models initially rely heavily on simple co-occurrence statistics, but later develop mechanisms that go beyond these surface correlations. By following information flow across training and intervening on individual attention heads, we find that grounding becomes concentrated in specialized aggregation mechanisms in the model’s middle layers. I will then show how this picture extends from controlled experiments to vision-language models, and discuss what it suggests about language learning, multimodal model design, hallucination, and the broader symbol grounding debate.
Bio: Ziqiao Martin Ma is a Member of Technical Staff at Thinking Machines Lab. He obtained his Ph.D. at the University of Michigan. His research stands at the intersection of language, interaction, and embodiment from a scalable and cognitive perspective, with the goal of grounding and aligning language agents to non-linguistic modalities and rich interactive contexts. He received an Outstanding Paper Award at ACL 2023, and an Amazon Alexa Prize Award.
References:
- Wu, S., Ma, Z., Luo, X., Huang, Y., Torres-Fonseca, J., Shi, F., & Chai, J. (2025). The Mechanistic Emergence of Symbol Grounding in Language Models. ICML. https://arxiv.org/abs/2510.13796
- Bick, A., Xing, E., & Gu, A. (2025). Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism. Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, 4324–4344. https://arxiv.org/abs/2504.18574
- Wang, L., Li, L., Dai, D., Chen, D., Zhou, H., Meng, F., Zhou, J., & Sun, X. (2023). Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9840–9855. https://aclanthology.org/2023.emnlp-main.609/
- Bisk, Y., Holtzman, A., Thomason, J., Andreas, J., Bengio, Y., Chai, J., Lapata, M., Lazaridou, A., May, J., Nisnevich, A., Pinto, N., & Turian, J. (2020). Experience Grounds Language. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 8718–8735. https://aclanthology.org/2020.emnlp-main.703/
- Bousselham, W., Petersen, F., Ferrari, V., & Kuehne, H. (2024). Grounding Everything: Emerging Localization Properties in Vision-Language Transformers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3828–3837. https://arxiv.org/abs/2312.00878
- Szot, A., Mazoure, B., Attia, O., Timofeev, A., Agrawal, H., Hjelm, D., Gan, Z., Kira, Z., & Toshev, A. (2025). From multimodal LLMs to generalist embodied agents: Methods and lessons. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10644–10655. https://arxiv.org/abs/2412.08442
Grounding Reasoning
This talk will explore an unconventional way of thinking about grounding reasoning: that reasoning is a kind of action, properly predicated of an entire animal, and integrating cognition, affect, and conation in the case of humans. I will outline some consequences of this for what it means to reason well, and for when we should and should not offload reasoning to agents.
Bio: Reto Gubelmann is a researcher working at the intersection of philosophy and natural language processing (NLP). His primary research area is the philosophical theory and computational implementation of argumentation and logical inference, broadly conceived as logical reasoning. His research in NLP involves large language models (LLMs) with an emphasis on Neuro-Symbolic approaches.
References:
- Gubelmann, R., & Hongler, P. (2026, July). Too Fast, Too Shallow–LLMs, Including Reasoning LLMs, Are Unreliable Constitutional Reasoners. In Findings of the Association for Computational Linguistics: ACL 2026 (pp. 40554–40572).
- Gubelmann, R. (2026). Putting reasons back into reasoning: how genuine reasoning is inference-based and why neuro-symbolic NLI could achieve it. Frontiers in Artificial Intelligence, 9, 1801094.
Data efficiency in children and language models
Large language models show intriguing emergent behaviors, yet they receive at least three to four — and sometimes as much as six — orders of magnitude more language data than human children. What accounts for this vast difference in sample efficiency? I will describe steps towards a paradigm in which we can address this question. In particular, I’ll discuss the use of egocentric video (“baby headcam”) data for model training, and the use of developmental data for model evaluation. This paradigm provides a model-based framework for exploring the nature of children’s early development.
Bio: Michael C. Frank is Benjamin Scott Crocker Professor of Human Biology in the Department of Psychology at Stanford University and Director of the Symbolic Systems Program. He studies children’s language learning and development, with a focus on the use of large-scale datasets to understand the variability and consistency of learning across cultures. He is a founder of the ManyBabies Consortium, and has led open-data projects including Wordbank and the ongoing LEVANTE project.
References:
- Frank, M. C. (2026). Children, but not language models, show accelerating returns in word learning. https://arxiv.org/abs/2608.17120
- Frank, M. C., & Goodman, N. D. (2025). Cognitive modeling using artificial intelligence. Annual Review of Psychology. doi:10.1146/annurev-psych-030625-040748.
- Frank, M. C. (2023). Bridging the data gap between children and large language models. Trends in Cognitive Sciences. https://psyarxiv.com/qzbgx/
Forecasting Spoken Language Development
Can bottom-up neural signals — brain data collected from infancy — forecast how spoken language develops, better than top-down clinical categories like diagnosis or demographics? Using MRI and EEG, we build predictive models of language outcomes that outperform standard predictors across typical, hearing-impaired, and autism-spectrum populations. Beyond prediction, these models probe how early cortical and subcortical processing grounds native and non-native speech perception, and how restored sensorimotor input, via cochlear implantation, recruits brain regions to support language. The work asks whether grounding language forecasts in neural data, rather than symbolic or demographic proxies, better captures how language actually emerges.
Bio: Patrick C. M. Wong studies the cultural and biological factors underlying variation in language and cognition across individuals. His work is interdisciplinary, spanning infant and adult brain imaging, perceptual psychophysics, grammar learning, gene sequencing, and predictive modeling of developmental trajectories. Wong joined The Chinese University of Hong Kong (CUHK) in 2013 after nearly a decade on the faculty at Northwestern University. He is Founding Director of CUHK’s Brain and Mind Institute and Professor of Linguistics, Paediatrics, and Psychology.
References:
- Wong, P. C. M., Pan, S., Lai, C. M., Chan, P. H. Y., Feng, G., Lam, H. S., Leung, T. Y., Novitskiy, N., & Leung, T. F. (2026). Speech auditory brainstem response to predict language delay. Pediatrics, 157(4), e2025073409.
- Wang, Y., Yuan, D., Dettman, S., Choo, D., Xu, E. S., Thomas, D., Ryan, M. E., Wong, P. C. M., & Young, N. M. (2026). Forecasting spoken language development in children with cochlear implants using preimplant magnetic resonance imaging. JAMA Otolaryngology–Head & Neck Surgery, 152(3), 232–241.
- Szot, A., Mazoure, B., Attia, O., Timofeev, A., Agrawal, H., Hjelm, D., Gan, Z., Kira, Z., & Toshev, A. (2025). From multimodal LLMs to generalist embodied agents: Methods and lessons. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10644–10655.
- Smith, L. B., Jayaraman, S., Clerkin, E., & Yu, C. (2018). The developing infant creates a curriculum for statistical learning. Trends in Cognitive Sciences, 22(4), 325–336. https://pmc.ncbi.nlm.nih.gov/articles/PMC5866780/
- Best, C. A., Yim, H., & Sloutsky, V. M. (2013). The cost of selective attention in category learning: Developmental differences between adults and infants. Journal of Experimental Child Psychology, 116(2), 105–119.
Situated Multimodal Word Learning
Abstract to come.
Biography to come.
Sensorimotor Structure in Lexical Meaning
Abstract to come.
Biography to come.
How Cognitive Immaturity May Support Cognitive Development and Learning
Childhood is a period of broad and protracted immaturity resulting in massive limitations and imposing large costs on a caretaker. These limitations include poor control, seemingly chaotic behavior, and very little planning, all resulting in no self-reliance and requiring extensive care. However, there are also important benefits of immaturity, those that make development and adaptation possible. These include the importance of starting small (Elman, Newport), perceptual narrowing (Werker, Kuhl), or tolerance to failure (Bjorklund). I propose another important immaturity-based benefit — broad exploration and information sampling early in development. Typically, this broad exploration and information sampling have been attributed to early curiosity, or non-instrumental value of information. I propose an alternative view, suggesting that broad exploration and information sampling is a consequence of immature Working Memory-Attention system. I then present experimental and computational evidence of how such system may subserve learning and cognitive development. Overall, this research demonstrates how (paradoxically) the very limitations of the early cognition may support or even drive cognitive development.
Bio: Vladimir Sloutsky is Professor of Psychology and Cognitive Science at the Ohio State University. A developmental psychologist and cognitive scientist, he studies how categories emerge in the course of development, how they become lexicalized, and how immature attention and memory subserve this process. A Fellow of Cognitive Science Society, he has authored over 200 scientific papers and chapters and received over $20M in research funding.
References:
- Sloutsky, V. M., Wan, Q., & Turner, B. M. (2026). Working memory, exploration, and cognitive development. Trends in Cognitive Sciences.
- Wan, Q., & Sloutsky, V. M. (2025). Working memory shapes information sampling and attention allocation across development. Journal of Experimental Psychology: General, 155, 479–498.
- Wan, Q., & Sloutsky, V. M. (2024). Exploration, distributed attention, and development of category learning. Psychological Science, 35, 1164–1177.
- Sloutsky, V. M., Ralston, R., Turner, B. M., & Ghetti, S. (2025). A little imprecision goes a long way in launching memory development. Child Development Perspectives, 19, 139–145.
Developmental Sensorimotor Grounding
Abstract to come.
Biography to come.
Perception Tokens for Visual Reasoning
Abstract to come.
Biography to come.
Embodied contributions to word meaning
Abstract to come.
Biography to come.
World Models: Evaluation and Representation
Abstract to come.
Biography to come.
The coherence of multimodal communication: inquiry and inference
Coherence theory postulates that each unit of a discourse stands in specific pragmatic relations to other parts of the discourse, with each relation involving its own information goals and inferential connections. I will sketch how it can be used to characterize a wide range of communicative contributions to multimodal interaction, including co-verbal gesture, practical demonstration, and situated inquiry. As an illustration, I take text–image coherence as a case study: text accompanying an image may characterize what’s visible in it, explain how it was obtained, or offer the author’s reaction — an insight leading to new methods for image–text inference, caption generation, and caption evaluation.
Bio: Matthew Stone is Professor of Computer Science and Cognitive Science at Rutgers—New Brunswick and Dean for Mathematical and Physical Sciences in the School of Arts and Sciences. He has served as program chair for NAACL and general chair of SIGDIAL, and chaired the Computer Science department from 2019 to 2023.
References:
- Alikhani, M., Khalid, B., & Stone, M. (2023). Image–text coherence and its implications for multimodal AI. Frontiers in Artificial Intelligence, 6:1048874.
- Lascarides, A., & Stone, M. (2009). Discourse coherence and gesture interpretation. Gesture, 9(2), 147–180.
- Stojnić, U., & Stone, M. (2025). Inquiry and Logical Form. Philosophical Perspectives. (Online first.)
- Stone, M., & Stojnić, U. (2015). Meaning and demonstration. Review of Philosophy and Psychology, 6(1), 68–97.
Spring Talks — Details
Learning Dexterity from Demonstration
Abstract to come.
Biography to come.
Are generative video models the way to solve visual intelligence?
The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. Curiously, the same primitives — large, generative models trained on web-scale data — apply to today’s generative video models. We demonstrate that Veo 3 can solve a broad variety of tasks it wasn’t explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, and more. These abilities enable early forms of visual reasoning like maze and symmetry solving, indicating that video models are on a path to becoming unified, generalist vision foundation models.
Bio: Robert Geirhos is a Staff Research Scientist at Google DeepMind in Zurich. He obtained his PhD on comparing human and machine vision from the University of Tübingen and the International Max Planck Research School for Intelligent Systems. His research currently focuses on video models as visual foundation models.
References:
- Video models are zero-shot learners and reasoners
Boas, Shannon, and the origin of semantic categories
Abstract to come.
Biography to come.
Resource-limited language models for cognitive science and AI
Abstract to come.
Biography to come.
Adaptive Speech Perception
Abstract to come.
Biography to come.
Meaning in brains vs. LLMs
Abstract to come.
Biography to come.
Language-guided learning in embodied agents
Abstract to come.
Biography to come.
Grounding Constrains Language Learning
Abstract to come.
Biography to come.
Multisensory Robotic Grounding
Abstract to come.
Biography to come.
Infant Grounding in Action
Abstract to come.
Biography to come.
Toward Active Visual Intelligence
Visual problem solving is active. Effective reasoners choose what evidence to inspect, which strategies to pursue, and how to adapt when an attempt fails. Yet vision-language models are typically optimized for correct answers, with much less attention to the reasoning strategies behind them. I will show that multi-task reinforcement learning induces distinct strategies across visual task categories, that giving models tools to interact with visual evidence improves generalization, and that learning from repeated attempts after failure can produce effective exploration strategies where conventional reinforcement learning stalls. Finally, I will show how models can learn from human eye-tracking data to improve how they seek out and reason about visual evidence.
Bio: Gabriel Sarch is a Postdoctoral Research Fellow in Princeton Language and Intelligence at Princeton University. His research focuses on active perception and visual reasoning, combining reinforcement learning with insights from human cognition. He received his joint Ph.D. in Neural Computation and Machine Learning from Carnegie Mellon University in 2025.
References:
- Sarch, G. H., Cai, L., Wang, Q., Wu, H., Chen, D., & Liu, Z. (2026). Vero: An Open RL Recipe for General Visual Reasoning. European Conference on Computer Vision (ECCV).
- Sarch, G. H., Saha, S., Khandelwal, N., Jain, A., Tarr, M. J., Kumar, A., & Fragkiadaki, K. (2025). Grounded Reinforcement Learning for Visual Reasoning. Advances in Neural Information Processing Systems (NeurIPS).
Visual Contributions to Lexical Meaning (topic being finalized)
Abstract to come.
Biography to come.
Grounding Speech in Multimodal Perception
Current speech processing models can perform tasks such as speech recognition and translation with a high degree of accuracy. They rely on large-scale corpora of transcribed speech for training; their application is therefore limited to the small subset of languages that can support large-scale data curation. In contrast, humans acquire spoken language from a comparatively small amount of speech data before they learn to read and write. I will describe our ongoing work to develop models of spoken language that do not learn from conventional text annotations, instead using perceptual grounding (to the visual modality) as a learning signal, and how the learned representations implicitly capture linguistic structure from the raw speech waveform.
Bio: David Harwath is an associate professor in the computer science department, University of Texas at Austin, where he leads the Speech, Audio, and Language Technologies (SALT) Lab. His group’s research focuses on developing novel machine learning methods applied to speech, audio, and multimodal data.
References:
- Harwath, D., Recasens, A., Surís, D., Chuang, G., Torralba, A., & Glass, J. (2018). Jointly discovering visual objects and spoken words from raw sensory input. Proceedings of the European Conference on Computer Vision (ECCV).
- Berry, L., Shih, Y.-J., Wang, H.-F., Chang, H.-J., Lee, H., & Harwath, D. (2023). M-SpeechCLIP: Leveraging large-scale, pre-trained models for multilingual speech to image retrieval. Proceedings of ICASSP.
- Peng, P., Li, S.-W., Räsänen, O., Mohamed, A., & Harwath, D. (2023). Syllable discovery and cross-lingual generalization in a visually grounded, self-supervised speech model. Proceedings of INTERSPEECH.
- Peng, P., & Harwath, D. (2022). Word discovery in visually grounded, self-supervised speech models. Proceedings of INTERSPEECH.
title to be confirmed
Abstract to come.
Biography to come.