Multimodal AI: Top-Down or Bottom-Up?

DIC/ISC/CRIA Seminar in Cognitive Informatics — Université du Québec à Montréal (UQÀM), Autumn 2026 & Spring 2027

Click any title below to jump to that speaker’s abstract. For the Zoom link, please email harnad.stevan@uqam.ca — it isn’t posted publicly on this page.

Autumn 2026

DateSpeakerTitle
September 10, 2026 (Pavillon PK-5115)Nicolas GouletThe Recombinant Power of Language
September 17, 2026Ziqiao MaMechanistic Emergence of Grounding
September 24, 2026Reto GubelmannGrounding Reasoning
October 1, 2026Michael FrankData efficiency in children and language models
October 8, 2026Patrick WongForecasting Spoken Language Development
October 15, 2026Gabriella ViglioccoSituated Multimodal Word Learning
October 22, 2026Louise ConnellSensorimotor Structure in Lexical Meaning
November 5, 2026Vladimir SloutskyHow Cognitive Immaturity May Support Cognitive Development and Learning
November 12, 2026Linda SmithSecrets in the Training Data: How Infant Behavior Solves Hard Learning Problems
November 19, 2026Ranjay KrishnaPerception Tokens for Visual Reasoning
November 26, 2026Agata DymarskaEmbodied contributions to word meaning
December 3, 2026Yoav ArtziWorld Models: Evaluation and Representation
December 10, 2026Matthew StoneThe coherence of multimodal communication: inquiry and inference
October 29, 2026No seminar — reading week

Spring 2027

Spring is under construction: some abstracts are still awaited and some dates await re-confirmation.

DateSpeakerTitle
January 14, 2027Desmond Elliott & Stella FrankVisual Contributions to Lexical Meaning (topic being finalized)
January 28, 2027Freda ShiGrounding Constrains Language Learning
February 4, 2027Aude BillardDexterity in Humans and Robots
February 11, 2027Chen YuInfant Grounding in Action
February 18, 2027Robert GeirhosAre generative video models the way to solve visual intelligence?
February 25, 2027Tal LinzenResource-limited language models for cognitive science and AI
March 4, 2027Ann R. BradlowAdaptive Speech Perception
March 11, 2027Gemma BoledaLLMs and symbolic and distributed approaches to language
March 18, 2027Jesse ThomasonMultisensory Robot Reasoning
March 25, 2027Terry RegierBoas, Shannon, and the origin of semantic categories
April 1, 2027Anna IvanovaMeaning in brains vs. LLMs
April 8, 2027Gabriel SarchToward Active Visual Intelligence
April 15, 2027Joyce ChaiLanguage-guided learning in embodied agents
April 22, 2027David HarwathGrounding Speech in Multimodal Perception

Autumn Talks — Details

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

The Recombinant Power of Language

Nicolas Goulet — UQÀM / HEC Montréal & Mila


September 10, 2026 (Pavillon PK-5115), 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Any complete dictionary can be represented as a definitional graph, in which every content-word (noun, verb, adjective, adverb) points (i.e., refers), via its definition, to other content words. Our earlier work showed that every dictionary graph has a “MinSet”: the smallest set of content-words whose referents must be learned directly — through the senses and action — from which all other content words can be defined indirectly in a purely recursive way, through recombinations of other content-words. The MinSet size turned out to be surprisingly small: about 1% of all the content-words. This talk presents a generalization of that result: rather than considering only the first sense of each word, we disambiguate the full set of senses of each content word, across several English dictionaries as well as new dictionaries in multiple languages. Although the total number of content words grows considerably once senses are disambiguated, the size of the MinSet grows only modestly, and this pattern holds across dictionaries and languages. This result reinforces the idea that a very small set of words already “grounded” directly through nonverbal, sensorimotor learning is enough to generate the entire lexicon indirectly, through verbal recombination alone — with implications for children’s language learning, pedagogy, and the limits of large language models, which have access only to recombination, never to direct grounding.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Nicolas Goulet

Nicolas Goulet is a doctoral student in data science at HEC and Mila. He studied the neural basis of behavior before beginning a master’s degree in computer science and AI at UQAM, supervised by É. Harnad and A. Blondin Massé. His doctoral research, supervised by T. Maharaj at the Errata Lab, connects language learning, the mathematical foundations of categorical perception in LLMs, and symbol grounding. He is also interested in neuroscience, sentience, and various aspects of mathematical computer science.

References:

Navigli, R. (2026). Is word sense disambiguation dead in the LLM era? Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, No. 46, pp. 39753–39762).

Meconi, D., Stirpe, S., Martelli, F., Lavalle, L., & Navigli, R. (2025). Do large language models understand word senses? In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 33885–33904).

Goulet, N., Massé, A. B., & Abdendi, M. (2025). Approaching the Source of Symbol Grounding with Confluent Reductions of Abstract Meaning Representation Directed Graphs. arXiv preprint arXiv:2508.11068.

Vincent-Lamarre, P., Massé, A. B., Lopes, M., Lord, M., Marcotte, O., & Harnad, S. (2016). The latent structure of dictionaries. Topics in Cognitive Science, 8(3), 625–659.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Mechanistic Emergence of Grounding

Ziqiao Ma — University of Michigan


September 17, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: What does it mean for a language model to actually ground a word in the world? Much of the current discussion treats grounding as an observed correspondence: a word like “horse” aligns with the right image region, so the model appears grounded. But correlation alone leaves a deeper question unanswered: how does this connection arise during learning, and what inside the model actually implements it? In this talk, I will approach grounding as a process rather than a property of a finished model. Starting from a minimal setting inspired by child language learning, we trace how models learn to connect linguistic symbols with corresponding information from the environment. Interestingly, models initially rely heavily on simple co-occurrence statistics, but later develop mechanisms that go beyond these surface correlations. By following information flow across training and intervening on individual attention heads, we find that grounding becomes concentrated in specialized aggregation mechanisms in the model’s middle layers. I will then show how this picture extends from controlled experiments to vision-language models, and discuss what it suggests about language learning, multimodal model design, hallucination, and the broader symbol grounding debate.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Ziqiao Ma

Ziqiao Martin Ma is a Member of Technical Staff at Thinking Machines Lab. He obtained his Ph.D. at the University of Michigan. His research stands at the intersection of language, interaction, and embodiment from a scalable and cognitive perspective, with the goal of grounding and aligning language agents to non-linguistic modalities and rich interactive contexts. He received an Outstanding Paper Award at ACL 2023, and an Amazon Alexa Prize Award.

References:

Wu, S., Ma, Z., Luo, X., Huang, Y., Torres-Fonseca, J., Shi, F., & Chai, J. (2025). The Mechanistic Emergence of Symbol Grounding in Language Models. ICML.

Bick, A., Xing, E., & Gu, A. (2025). Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism. Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, 4324–4344.

Wang, L., Li, L., Dai, D., Chen, D., Zhou, H., Meng, F., Zhou, J., & Sun, X. (2023). Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9840–9855.

Bisk, Y., Holtzman, A., Thomason, J., Andreas, J., Bengio, Y., Chai, J., Lapata, M., Lazaridou, A., May, J., Nisnevich, A., Pinto, N., & Turian, J. (2020). Experience Grounds Language. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 8718–8735.

Bousselham, W., Petersen, F., Ferrari, V., & Kuehne, H. (2024). Grounding Everything: Emerging Localization Properties in Vision-Language Transformers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3828–3837.

Szot, A., Mazoure, B., Attia, O., Timofeev, A., Agrawal, H., Hjelm, D., Gan, Z., Kira, Z., & Toshev, A. (2025). From multimodal LLMs to generalist embodied agents: Methods and lessons. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10644–10655.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Grounding Reasoning

Reto Gubelmann — University of Zurich


September 24, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: This talk will explore an unconventional way of thinking about grounding reasoning: that reasoning is a kind of action, properly predicated of an entire animal, and integrating cognition, affect, and conation in the case of humans. I will outline some consequences of this for what it means to reason well, and for when we should and should not offload reasoning to agents.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Reto Gubelmann

Reto Gubelmann is a researcher working at the intersection of philosophy and natural language processing (NLP). His primary research area is the philosophical theory and computational implementation of argumentation and logical inference, broadly conceived as logical reasoning. His research in NLP involves large language models (LLMs) with an emphasis on Neuro-Symbolic approaches.

References:

Gubelmann, R., & Hongler, P. (2026, July). Too Fast, Too Shallow – LLMs, Including Reasoning LLMs, Are Unreliable Constitutional Reasoners. In Findings of the Association for Computational Linguistics: ACL 2026 (pp. 40554–40572).

Gubelmann, R. (2026). Putting reasons back into reasoning: how genuine reasoning is inference-based and why neuro-symbolic NLI could achieve it. Frontiers in Artificial Intelligence, 9, 1801094.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Data efficiency in children and language models

Michael Frank — Stanford University


October 1, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Large language models show intriguing emergent behaviors, yet they receive at least three to four — and sometimes as much as six — orders of magnitude more language data than human children. What accounts for this vast difference in sample efficiency? I will describe steps towards a paradigm in which we can address this question. In particular, I’ll discuss the use of egocentric video (“baby headcam”) data for model training, and the use of developmental data for model evaluation. This paradigm provides a model-based framework for exploring the nature of children’s early development.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Michael Frank

Michael C. Frank is Benjamin Scott Crocker Professor of Human Biology in the Department of Psychology at Stanford University and Director of the Symbolic Systems Program. He received his PhD from MIT in Brain and Cognitive Sciences in 2010. He studies children’s language learning and development, with a focus on the use of large-scale datasets to understand the variability and consistency of learning across cultures. He is a founder of the ManyBabies Consortium, and has led open-data projects including Wordbank and the ongoing LEVANTE project. He has received awards including the Troland Award from the National Academy of Sciences and the FABBS Early Career Impact award. He served as President of the Cognitive Science Society, has edited for journals including Cognition and Child Development, and is current co-Editor in Chief of the Open Encyclopedia of Cognitive Science.

References:

Frank, M. C. (2026). Children, but not language models, show accelerating returns in word learning.

Frank, M. C., & Goodman, N. D. (2025). Cognitive modeling using artificial intelligence. Annual Review of Psychology. doi:10.1146/annurev-psych-030625-040748.

Frank, M. C. (2023). Bridging the data gap between children and large language models. Trends in Cognitive Sciences.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Forecasting Spoken Language Development

Patrick Wong — The Chinese University of Hong Kong


October 8, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Can bottom-up neural signals — brain data collected from infancy — forecast how spoken language develops, better than top-down clinical categories like diagnosis or demographics? Using MRI and EEG, we build predictive models of language outcomes that outperform standard predictors across typical, hearing-impaired, and autism-spectrum populations. Beyond prediction, these models probe how early cortical and subcortical processing grounds native and non-native speech perception, and how restored sensorimotor input, via cochlear implantation, recruits brain regions to support language. The work asks whether grounding language forecasts in neural data, rather than symbolic or demographic proxies, better captures how language actually emerges.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Patrick Wong

Patrick C. M. Wong studies the cultural and biological factors underlying variation in language and cognition across individuals. His work is interdisciplinary, spanning infant and adult brain imaging, perceptual psychophysics, grammar learning, gene sequencing, and predictive modeling of developmental trajectories. Wong joined The Chinese University of Hong Kong (CUHK) in 2013 after nearly a decade on the faculty at Northwestern University. He is Founding Director of CUHK’s Brain and Mind Institute and Professor of Linguistics, Paediatrics, and Psychology.

References:

Wong, P. C. M., Pan, S., Lai, C. M., Chan, P. H. Y., Feng, G., Lam, H. S., Leung, T. Y., Novitskiy, N., & Leung, T. F. (2026). Speech auditory brainstem response to predict language delay. Pediatrics, 157(4), e2025073409.

Wang, Y., Yuan, D., Dettman, S., Choo, D., Xu, E. S., Thomas, D., Ryan, M. E., Wong, P. C. M., & Young, N. M. (2026). Forecasting spoken language development in children with cochlear implants using preimplant magnetic resonance imaging. JAMA Otolaryngology–Head & Neck Surgery, 152(3), 232–241.

Szot, A., Mazoure, B., Attia, O., Timofeev, A., Agrawal, H., Hjelm, D., Gan, Z., Kira, Z., & Toshev, A. (2025). From multimodal LLMs to generalist embodied agents: Methods and lessons. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10644–10655.

Smith, L. B., Jayaraman, S., Clerkin, E., & Yu, C. (2018). The developing infant creates a curriculum for statistical learning. Trends in Cognitive Sciences, 22(4), 325–336.

Best, C. A., Yim, H., & Sloutsky, V. M. (2013). The cost of selective attention in category learning: Developmental differences between adults and infants. Journal of Experimental Child Psychology, 116(2), 105–119.

Novitskiy, N., Maggu, A. R., Lai, C. M., Chan, P. H. Y., Wong, K. H. Y., Lam, H. S., Leung, T. Y., Leung, T. F., & Wong, P. C. M. (2022). Early development of neural speech encoding depends on age but not native language status: Evidence from lexical tone. Neurobiology of Language, 3(1), 67–86.

Feng, G., Ingvalson, E. M., Grieco-Calub, T. M., Roberts, M. Y., Ryan, M. E., Birmingham, P., Burrowes, D., Young, N., & Wong, P. C. M. (2018). Neural preservation underlies speech improvement from auditory deprivation in young cochlear implant recipients. Proceedings of the National Academy of Sciences, 115(5), E1022–E1031.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Situated Multimodal Word Learning

Gabriella Vigliocco — University College London


October 15, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: I will discuss similarities and differences between the way people learn about new things in everyday life from caregivers or teachers and the way language models learn. In everyday face-to-face contexts, in addition to the information in the linguistic input, human learners can take advantage of situated information, including the visual communicative signals produced by teachers, the physical availability of what is being talked about (especially if concrete) and, crucially, also the contingency between what the learner and the teacher say and do. Studies show that multimodal communicative signals like gaze and gesture, active engagement in the learning, and the contingency between teacher and learner behaviours all influence learning in a dynamic way across ages (from childhood to adulthood). I will discuss possible cognitive and neural mechanisms mediating these effects and then open the discussion to the challenges and opportunities these findings pose for learning in artificial systems, focusing on both multimodality and social contingency.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Gabriella Vigliocco

Gabriella Vigliocco is Professor of Psychology and Language Sciences in the Department of Experimental Psychology at University College London, where she directs the Language and Cognition Lab. Her research examines how language and cognition are grounded in embodied, multimodal experience — including gesture, iconicity, and sensorimotor and emotional information — and how this shapes word learning, sentence processing, and language development across the lifespan.

References:

Motamedi, Y., Murgiano, M., Grzyb, B., Gu, Y., Kewenig, V., Brieke, R., Donnellan, E., Marshall, C., Wonnacott, E., Perniss, P., & Vigliocco, G. (2024). Language development beyond the here-and-now: Iconicity and displacement in child-directed communication. Child Development, 95(5), 1539–1557.

Murgiano, M., Motamedi, Y., & Vigliocco, G. (2021). Situating Language in the Real-World: The Role of Multimodal Iconicity and Indexicality. Journal of Cognition, 4(1).

Szegedi, D., Rozic, G., Fonagy, P., Hamilton, A., & Vigliocco, G. (2026, under review). Learning about “Inflation”: Mapping Caregivers’ Mentalizing and Pedagogical Strategies that Support 8–9-Year-Old Children’s Learning of Abstract Concepts.

De Felice, S., Hamilton, A. F. de C., Ponari, M., & Vigliocco, G. (2023). Learning from others is good, with others is better: The role of social interaction in human acquisition of new knowledge. Philosophical Transactions of the Royal Society B, 378(1870), 20210357.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Sensorimotor Structure in Lexical Meaning

Louise Connell — Maynooth University


October 22, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: The cognitive sciences view concepts as the heart of cognition, yet it remains a matter of debate how we mentally represent lexical meanings as diverse as cat, affection, and calculus. Grounded theories of representation hold that the same neural systems engaged during perception and action experience are also engaged when meaning is processed during conceptual tasks, thus providing a grounding mechanism for semantic memory. However, the contribution of sensorimotor information beyond the senses of vision and hearing (and to a lesser extent touch and smell) is not well understood, nor is the role of sensorimotor information in grounding abstract concepts. By incorporating a multidimensional variety of perceptual experience from a range of distinct modalities and action experience from a range of distinct bodily effectors, and doing so at the full scale of adult semantic memory (approximately 40k lexical concepts), we have tested grounded theories across a broad variety of cognitive tasks. From semantic similarity to visual word recognition to categorical structure, we find that the sensorimotor experience underlying meaning representations plays an essential role in both concrete and abstract domains. Overall, these findings suggest that sensorimotor information helps provide humans with a robust, flexible conceptual system, where the distinction between abstract and concrete words is not as clear-cut as ontological assumptions might suggest.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Louise Connell

Louise Connell is a Professor in the Department of Psychology at Maynooth University, part of the National University of Ireland. An interdisciplinary researcher in psychology and cognitive science, she specializes in the sensorimotor basis of concepts in semantic memory and the role of language in cognition.

References:

Banks, B., & Connell, L. (2022). Multidimensional sensorimotor grounding of concrete and abstract categories. Philosophical Transactions of the Royal Society B: Biological Sciences, 378(1870), Article 20210366.

Connell, L., & Lynott, D. (2014). Principles of representation: Why you can’t represent the same concept twice. Topics in Cognitive Science, 6, 390-406.

Connell, L., & Lynott, D. (2024). What can language models tell us about human cognition? Current Directions in Psychological Science, 33(3), 181-189.

Lynott, D., Connell, L., Brysbaert, M., Brand, J., & Carney, J. (2020). The Lancaster Sensorimotor Norms: Multidimensional measures of perceptual and action strength for 40,000 English words. Behavior Research Methods, 52, 1271-1291.

Wingfield, C., & Connell, L. (2023). Sensorimotor distance: A fully grounded measure of semantic similarity for 800 million concept pairs. Behavior Research Methods, 55, 3416–3432.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

How Cognitive Immaturity May Support Cognitive Development and Learning

Vladimir Sloutsky — The Ohio State University


November 5, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Childhood is a period of broad and protracted immaturity resulting in massive limitations and imposing large costs on a caretaker. These limitations include poor control, seemingly chaotic behavior, and very little planning, all resulting in no self-reliance and requiring extensive care. However, there are also important benefits of immaturity, those that make development and adaptation possible. These include the importance of starting small (Elman, Newport), perceptual narrowing (Werker, Kuhl), or tolerance to failure (Bjorklund). I propose another important immaturity-based benefit — broad exploration and information sampling early in development. Typically, this broad exploration and information sampling have been attributed to early curiosity, or non-instrumental value of information. I propose an alternative view, suggesting that broad exploration and information sampling is a consequence of immature Working Memory-Attention system. I then present experimental and computational evidence of how such system may subserve learning and cognitive development. Overall, this research demonstrates how (paradoxically) the very limitations of the early cognition may support or even drive cognitive development.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Vladimir Sloutsky

Vladimir Sloutsky is Professor of Psychology and Cognitive Science at the Ohio State University. A developmental psychologist and cognitive scientist, he studies how categories emerge in the course of development, how they become lexicalized, and how immature attention and memory subserve this process. A Fellow of Cognitive Science Society, he has authored over 200 scientific papers and chapters and received over $20M in research funding.

References:

Sloutsky, V. M., Wan, Q., & Turner, B. M. (2026). Working memory, exploration, and cognitive development. Trends in Cognitive Sciences.

Wan, Q., & Sloutsky, V. M. (2025). Working memory shapes information sampling and attention allocation across development. Journal of Experimental Psychology: General, 155, 479–498.

Wan, Q., & Sloutsky, V. M. (2024). Exploration, distributed attention, and development of category learning. Psychological Science, 35, 1164–1177.

Sloutsky, V. M., Ralston, R., Turner, B. M., & Ghetti, S. (2025). A little imprecision goes a long way in launching memory development. Child Development Perspectives, 19, 139–145.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Secrets in the Training Data: How Infant Behavior Solves Hard Learning Problems

Linda B. Smith — Indiana University, Bloomington


November 12, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Much of the information in the world is latent: it is not revealed without action by the perceiver. This creates direct, real-time relations between sensory input and behavior, between the current and next states of the learning system, and between inputs and the outcomes of actions on the world. I will present findings from our analyses of the visual statistics of infants’ egocentric images, collected at the scale of daily life in the home. These findings suggest that the quality and spatiotemporal structure of infants’ self-generated input contribute to their efficient learning about visual objects. When human learning is more efficient than current learning mechanisms can explain, theorists often posit intrinsic “inductive biases” that constrain learning outcomes, enabling faster and more reliable learning from complex, variable, and noisy training data. The visual statistics generated by infants and toddlers interacting with their everyday world reveal another source of constraint: the learner’s own behavior structures and biases the training data itself, rather than merely constraining the inferences drawn from it. These findings may suggest principles for designing training data that support efficient learning even in machines whose learning mechanisms differ from those of humans.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Linda B. Smith

Linda B. Smith is Distinguished Professor of Psychological and Brain Sciences at Indiana University – Bloomington. Taking a complex systems perspective, she seeks to understand the interdependencies among perceptual, motor, and cognitive developments during the first three years of post-natal life. She uses wearable sensors to study how the young learner’s own behavior creates the statistical structure of the learning environment at the scale of everyday life.

References:

Ramirez, F. M., Clerkin, E. M., Crandall, D. J., & Smith, L. B. (2026). A solution to generalized learning from small training sets found in infants’ repeated visual experiences of individual objects. Proceedings of the National Academy of Sciences, 123(33), e2536143123.

Petroff, Z. J., Jayaraman, S., Smith, L. B., Candy, T. R., & Bonnen, K. (2025). The world through infant eyes: Evidence for the early emergence of the cardinal orientation bias. PNAS, 122(16), Article e2421277122.

Sheybani, S., Hansaria, H., Wood, J. N., Smith, L. B., & Tiganj, Z. (2023). Curriculum learning with infant egocentric videos. Advances in Neural Information Processing Systems (NeurIPS 2023).

Karmazyn-Raz, H., & Smith, L. B. (2023). Sampling statistics are like story creation: A network analysis of parent–toddler exploratory play. Philosophical Transactions of the Royal Society B, 378(1870).

Smith, L. B., Jayaraman, S., Clerkin, E., & Yu, C. (2018). The developing infant creates a curriculum for statistical learning. Trends in Cognitive Sciences, 22(4), 325–336.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Perception Tokens for Visual Reasoning

Ranjay Krishna — University of Washington


November 19, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Today’s vision-language models can describe images fluently, yet they still struggle with the spatial and perceptual reasoning required for robotics. I will argue that visual reasoning—reasoning through points, sketches, trajectories, depth maps, and other grounded representations—will become more important than language reasoning for physical agents. It traces progress from benchmarks exposing fundamental perceptual failures to methods that teach models to reason directly in visual space. These ideas culminate in MolmoAct and MolmoAct2, open vision-language-action models that convert visual plans into robot actions. The resulting systems are interpretable, steerable, deployable zero-shot on affordable hardware, and competitive with proprietary models trained on substantially more data—advancing the goal of accessible, general-purpose robots for real-world environments.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Ranjay Krishna

Ranjay Krishna is Assistant Professor in the Paul G. Allen School of Computer Science & Engineering at the University of Washington, where he co-directs the RAIVN lab. His research bridges computer vision, natural language processing, robotics, and human-computer interaction, focusing on teaching machines to perceive the visual world and interact with people by drawing on behavioral and social-science frameworks. He previously directed the multimodal and embodied AI team at the Allen Institute for AI.

References:

Bigverdi, M., Luo, Z., Hsieh, C.-Y., Shen, E., Chen, D., Shapiro, L. G., & Krishna, R. (2025). Perception Tokens Enhance Visual Reasoning in Multimodal Language Models. Conference on Computer Vision and Pattern Recognition (CVPR).

Liu, B., Dong, Y., Wang, Y., Ma, Z., Tang, Y., Tang, L., Rao, Y., Ma, W.-C., & Krishna, R. (2025). Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model. Conference on Computer Vision and Pattern Recognition (CVPR), 3783–3792.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Embodied contributions to word meaning

Agata Dymarska — Adam Mickiewicz University, Poznań


November 26, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Our perceptual and motor experience with the world shapes the way we process and understand language. Depending on concept type, different aspects of sensorimotor experience are at play, as evidenced by facilitation of lexical decision and word recognition performance. However, at times, a rich sensorimotor representation can be detrimental to correct word recognition. This talk will focus on research findings from lexical decision and word recognition memory tasks. These studies show which forms of sensorimotor experience are activated when viewing different words. They reveal when that activation helps identify and maintain accurate conceptual representations, and when it actually harms performance. Evidence from non-native English speakers who draw on sensorimotor information differently shows how the pattern of sensorimotor grounding can change with language experience.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Agata Dymarska

Agata Dymarska is a researcher at Adam Mickiewicz University in Poznań, Poland. She completed her PhD at Lancaster University, UK, in the Embodied Cognition Lab, examining the interplay of language and sensorimotor information in word recognition and memory. Her research interests include sensorimotor activation in second language users and the contribution of sensorimotor information to emotional word processing.

References:

Dymarska, A. (2025). Frequency over semantic richness: Word recognition in non-native English speakers. Bilingualism: Language and Cognition.

Dymarska, A., & Connell, L. (2025). Sensorimotor effects in surprise word memory – A registered report. Cortex.

Dymarska, A., Connell, L., & Banks, B. (2023). More is not necessarily better: How different aspects of sensorimotor experience affect recognition memory for words. Journal of Experimental Psychology: Learning, Memory, and Cognition, 49(10), 1572–1587.

Lynott, D., Connell, L., Brysbaert, M., Brand, J., & Carney, J. (2020). The Lancaster Sensorimotor Norms: Multidimensional measures of perceptual and action strength for 40,000 English words. Behavior Research Methods, 52, 1271–1291.

Yap, M. J., Tan, S. E., Pexman, P. M., & Hargreaves, I. S. (2011). Is more always better? Effects of semantic richness on lexical decision, speeded pronunciation, and semantic classification. Psychonomic Bulletin & Review, 18(4), 742–750.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Multi-modal Evaluation / State Computation

Yoav Artzi — Cornell University / Google DeepMind


December 3, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: This talk briefly describes three distinct projects. The first two focus on evaluating multi-modal capabilities in models through the evaluation of dual perspective reasoning and fine-grained manipulation. For the former, we focus on spatial reasoning across egocentric and allocentric observations, following setups designed for the evaluation of cognitive development in children. The latter utilizes knot manipulation, a scenario well modeled by mathematical theory, which provides a clear avenue to gauge complexity and generalization. The last part of the talk is focused on a fundamental capability critical for cross-modal reasoning: the ability to separate between how one represents state and one’s prediction of outputs. We hypothesize that mechanically separating the two in LLMs is beneficial, and show this in the context of LLM pre-training.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Yoav Artzi

Yoav Artzi is Associate Professor in the Department of Computer Science and Cornell Tech at Cornell University, and visiting faculty research at Google DeepMind. His research focuses on language modeling and learning in interactive and situated scenarios. His work was acknowledged by awards and honorable mentions at ACL, EMNLP, NAACL, and IROS, as well as a TACL test-of-time award. He was previously arXiv’s associate faculty director, and co-founded COLM.

References:

Chen, Z., & Artzi, Y. (2025). Knot so simple: A minimalistic environment for spatial reasoning. Advances in Neural Information Processing Systems: Datasets and Benchmarks Track.

Monea, G., Godey, N., Brantley, K., & Artzi, Y. (2026). The state-prediction separation hypothesis. arXiv preprint arXiv:2607.01218.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

The coherence of multimodal communication: inquiry and inference

Matthew Stone — Rutgers University


December 10, 2026, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115

Abstract: Coherence theory postulates that each unit of a discourse stands in specific pragmatic relations to other parts of the discourse, with each relation involving its own information goals and inferential connections. I will sketch how it can be used to characterize a wide range of communicative contributions to multimodal interaction, including co-verbal gesture, practical demonstration, and situated inquiry. As an illustration, I take text–image coherence as a case study: text accompanying an image may characterize what’s visible in it, explain how it was obtained, or offer the author’s reaction — an insight leading to new methods for image–text inference, caption generation, and caption evaluation.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Matthew Stone

Matthew Stone is Professor of Computer Science and Cognitive Science at Rutgers—New Brunswick and Dean for Mathematical and Physical Sciences in the School of Arts and Sciences. He has served as program chair for NAACL and general chair of SIGDIAL, and chaired the Computer Science department from 2019 to 2023.

References:

Alikhani, M., Khalid, B., & Stone, M. (2023). Image–text coherence and its implications for multimodal AI. Frontiers in Artificial Intelligence, 6:1048874.

Lascarides, A., & Stone, M. (2009). Discourse coherence and gesture interpretation. Gesture, 9(2), 147–180.

Stojnić, U., & Stone, M. (2025). Inquiry and Logical Form. Philosophical Perspectives. (Online first.)

Stone, M., & Stojnić, U. (2015). Meaning and demonstration. Review of Philosophy and Psychology, 6(1), 68–97.

↑ back to calendar

Spring Talks — Details

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Visual Contributions to Lexical Meaning (topic being finalized)

Desmond Elliott & Stella Frank — University of Copenhagen

January 14, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: awaiting date re-confirmation

Abstract: please provide 50-100-word abstract






The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Desmond Elliott

Desmond Elliott is Associate Professor in the Artificial Intelligence Section, Department of Computer Science, University of Copenhagen. He works on building and understanding multimodal and multilingual models, with a focus on vision-and-language learning, grounded machine translation, and tokenization-free approaches to language processing.

References (Elliott):

Elliott, D., Frank, S., Sima’an, K., & Specia, L. (2016). Multi30K: Multilingual English-German image descriptions. Proceedings of the 5th Workshop on Vision and Language (VL’16).

Elliott, D., & Kádár, Á. (2017). Imagination improves multimodal translation. Proceedings of the Eighth International Joint Conference on Natural Language Processing (Vol. 1).

Photo of Stella Frank

Stella Frank is a Postdoctoral Researcher at the Pioneer Centre for AI, Department of Computer Science, University of Copenhagen. Her research designs computational models of how humans and machines connect language to the world, focusing on grounded and multimodal language learning, cross-modal representation, and computational psycholinguistics.

References (Frank):

Frank, S., Bugliarello, E., & Elliott, D. (2021). Vision-and-language or vision-for-language? On cross-modal influence in multimodal transformers. Proceedings of EMNLP 2021.

Oneata, D., Elliott, D., & Frank, S. (2025). Seeing what tastes good: Revisiting multimodal distributional semantics in the billion parameter era. Findings of ACL 2025.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Grounding Constrains Language Learning

Freda Shi — University of Waterloo

January 28, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: awaiting date re-confirmation

Abstract: please provide 50-100-word abstract






The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Freda Shi

Freda Shi is an Assistant Professor in the Cheriton School of Computer Science at the University of Waterloo and a Faculty Member at the Vector Institute, where she leads the CompLING Lab. Her research focuses on grounded and multimodal language learning, unsupervised language acquisition, and connections between language and other modalities such as vision.

References:

Shi, H., Mao, J., Gimpel, K., & Livescu, K. (2019). Visually grounded neural syntax acquisition. Proceedings of ACL 2019.

Mao, J., Shi, F., Wu, J., Levy, R., & Tenenbaum, J. (2021). Grammar-based grounded lexicon learning. Advances in Neural Information Processing Systems 34 (NeurIPS 2021).

Zhang, Z., Hu, F., Lee, J., Shi, F., Kordjamshidi, P., Chai, J., & Ma, Z. (2025). Do vision-language models represent space and how? Evaluating spatial frame of reference under ambiguities. ICLR 2025 (Oral).

Ogezi, M., & Shi, F. (2025). SpaRE: Enhancing spatial reasoning in vision-language models with synthetic data. Proceedings of ACL 2025 (Vol. 1).

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Dexterity in Humans and Robots

Aude Billard — EPFL

February 4, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

The human hand has long been a valuable source of inspiration for the design of robotic hands. It is often regarded as the ultimate embodiment of dexterity. But this raises the question: what is dexterity? This talk will walk you through a quest to better understand how humans acquire such dexterity. We will start with a longitudinal study of apprenticeship in watchmaking and see how the central nervous system learns simultaneously a forward model of the task and new hand postures capable of generating the required highly precise movements. We will then turn to a longitudinal study of the acquisition of microsurgery and look at the CNS ability to model the subtle interactions between the hand and tissue. If these studies advance our understanding of human motor control, they provide only general guidance in developing algorithms for robots to acquire such dexterity. We will then turn to the question of what robotic hands are capable of, what are the essential algorithmic components necessary to manipulate objects dexterously. The talk will conclude by questioning whether human dexterity is really the holy grail when it comes to robotic dexterity, showing an example of a non anthropomorphic hand that can in some aspects exceed human capability.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Aude Billard

Aude Billard is full professor and the head of the LASA laboratory, and director of the Robotics Center at the Swiss Institute of Technology Lausanne (EPFL). She was the President of the IEEE Robotics and Automation Society. She currently serves as director of the EPFL Robotics Center and Vice-President of the Swiss Robotics Association. Her research spans machine learning, robotics, computational neuroscience, and motor control.

References:

Hu, L., Gholami, S., Dindelegan, G., Meling, T. R., & Billard, A. (2026). Quantitative outcome-oriented assessment of microsurgical anastomosis. arXiv.

Billard, A., & Kragic, D. (2019). Trends and challenges in robot manipulation. Science, 364(6446), 8414.

Gao, X., Yao, K., Junge, K., Hughes, J., & Billard, A. (2026). A detachable crawling robotic hand. Nature Communications, 17(1), 428.

Yao, K., Sternad, D., & Billard, A. (2021). Hand pose selection in a bimanual fine-manipulation task. Journal of Neurophysiology, 126(1), 195–212.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Infant Grounding in Action

Chen Yu — University of Texas at Austin

February 11, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

Abstract: please provide 50-100-word abstract






The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Chen Yu

Chen Yu, Charles and Sarah Seay Regents Professor of Psychology at the University of Texas at Austin, directs the Developing Intelligence Lab. His research uses head-mounted eye-tracking and multimodal sensing to study how infants’ own actions — gaze, reaching, manipulation — generate the everyday visual and statistical experience underlying early word and category learning.

References:

Yu, C., & Smith, L. B. (2012). Embodied attention and word learning by toddlers. Cognition, 125(2), 244–262.

Smith, L. B., Jayaraman, S., Clerkin, E., & Yu, C. (2018). The developing infant creates a curriculum for statistical learning. Trends in Cognitive Sciences, 22(4), 325–336.

Clerkin, E. M., Hart, E., Rehg, J. M., Yu, C., & Smith, L. B. (2017). Real-world visual statistics and infants’ first-learned object names. Philosophical Transactions of the Royal Society B, 372(1711), 20160055.

Bambach, S., Crandall, D., Smith, L., & Yu, C. (2018). Toddler-inspired visual object learning. Advances in Neural Information Processing Systems 31 (NeurIPS 2018).

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Are generative video models the way to solve visual intelligence?

Robert Geirhos — Google DeepMind

February 18, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. Curiously, the same primitives — large, generative models trained on web-scale data — apply to today’s generative video models. We demonstrate that Veo 3 can solve a broad variety of tasks it wasn’t explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, and more. These abilities enable early forms of visual reasoning like maze and symmetry solving, indicating that video models are on a path to becoming unified, generalist vision foundation models.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Robert Geirhos

Robert Geirhos is a Staff Research Scientist at Google DeepMind in Zurich. He obtained his PhD on comparing human and machine vision from the University of Tübingen and the International Max Planck Research School for Intelligent Systems. His research currently focuses on video models as visual foundation models.

References:

Wiedemer, T., Li, Y., Vicol, P., Gu, S. S., Matarese, N., Swersky, K., Kim, B., Jaini, P., & Geirhos, R. (2025). Video models are zero-shot learners and reasoners. arXiv preprint.

Kataoka, H., et al., & Geirhos, R. (2026). Visual general intelligence: A white paper. arXiv preprint.

Ramirez, F., Clerkin, E. M., Crandall, D. J., & Smith, L. B. (2026). A solution to generalized learning from small training sets found in infants’ repeated visual experiences of individual objects. Proceedings of the National Academy of Sciences, 123(33), e2536143123.

Video demonstrations: Video models are zero-shot learners and reasoners.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Resource-limited language models for cognitive science and AI

Tal Linzen — New York University

February 25, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

Abstract: AI models surpass human capabilities in many areas: as a simple example, models can recall long lists of digits without errors, but humans can’t. In this talk, I’ll discuss ongoing work that highlights two consequences of the growing gap between models and humans, and proposes ways to address them. The first consequence of models’ superhumanness is diminished effectiveness as cognitive models: I will show this through the example of human next-word prediction in reading. The second is their limited usefulness as human simulators; I will argue that if models incorporated human resource limitations, they could be an effective way to evaluate and improve AI models’ ability to interact with humans.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Tal Linzen

Tal Linzen is Associate Professor of Linguistics and Data Science at New York University, where he directs the Computation and Psycholinguistics Lab. His research uses computational and behavioral methods to study how people and neural language models learn and understand syntax and meaning, focusing on cognitively-inspired evaluation, syntactic generalization, and sample-efficient language model training.

References:

Linzen, T., Dupoux, E., & Goldberg, Y. (2016). Assessing the ability of LSTMs to learn syntax-sensitive dependencies. Transactions of the Association for Computational Linguistics, 4, 521–535.

Linzen, T. (2020). How can we accelerate progress towards human-like linguistic generalization? Proceedings of the 58th Annual Meeting of the ACL (pp. 5210–5217).

Warstadt, A., Mueller, A., Choshen, L., et al., & Cotterell, R. (2023). Findings of the BabyLM Challenge: Sample-efficient pretraining on developmentally plausible corpora. Proceedings of the BabyLM Challenge at CoNLL 2023.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Adaptive Speech Perception

Ann R. Bradlow — Northwestern University

March 4, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

Abstract: Speech perception is necessarily adaptive because speech production is variable. Top-down perspectives on the mechanisms that drive adaptive speech perception seek to link sources of systematic speech variation to dimensions of generalization of perceptual adaptation. For example, exposure to multiple utterances spoken by a single foreign-accented talker should generalize to novel utterances by that talker, while exposure to utterances spoken by multiple talkers of a given accent should generalize to novel talkers of that accent. While this pattern of generalization following high-variability exposure has received substantial empirical support, low-variability (e.g., single-talker) training can also lead to generalized perceptual adaptation. In this talk, I will present a series of studies that examines how top-down, linguistically-guided adaptation may combine with bottom-up, signal-driven similarity to yield generalized perceptual adaptation to speech variation. Key empirical developments that underlie this research are availability of large, curated databases of first-language and second-language speech, objective intelligibility ratings from large groups of human listeners, and recent applications of utterance embeddings in a pre-trained semi-supervised machine learning model.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Ann R. Bradlow

Ann R. Bradlow is the Abraham Harris Professor of Linguistics at Northwestern University. Her research examines speech perception and production, focusing on second-language/accented speech, cross-talker intelligibility, and perceptual learning and adaptation in spoken language processing. She is a Fellow of the Acoustical Society of America and of AAAS.

References:

Kim, S.-E., Goldrick, M., Keshet, J., & Bradlow, A. R. (2026). Generalization of perceptual adaptation to second-language speech in a single-talker training paradigm: Predictors remain elusive. Attention, Perception, & Psychophysics, 88, 192.

Kim, S.-E., Chernyak, B. R., Keshet, J., Goldrick, M., & Bradlow, A. R. (2025). Predicting relative intelligibility from inter-talker distances in a perceptual similarity space for speech. Psychonomic Bulletin & Review, 32(4), 1664–1675.

Bradlow, A. R., Bassard, A. M., & Paller, K. A. (2023). Generalized perceptual adaptation to second-language speech: Variability, similarity, and intelligibility. Journal of the Acoustical Society of America, 154(3), 1601–1613.

Bradlow, A. R., & Bent, T. (2008). Perceptual adaptation to non-native speech. Cognition, 106(2), 707–729.

Additional references:

Bradlow, A. R., & Alexander, J. A. (2007). Semantic and phonetic enhancements for English sentence-in-noise recognition by native and non-native listeners. Journal of the Acoustical Society of America, 121(4), 2339–2349.

Bent, T., & Bradlow, A. R. (2003). The interlanguage speech intelligibility benefit. Journal of the Acoustical Society of America, 114(3), 1600–1610.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

LLMs and symbolic and distributed approaches to language

Gemma Boleda — Universitat Pompeu Fabra

March 11, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

Abstract: The language we produce on a given occasion is influenced by two main factors: the linguistic system in our brains, on the one hand, and the context in which we use it—where, when, with whom, about what we speak. In turn, how we use language ends up influencing the system itself; for instance, over time, it can change the meaning of words. My talk will center around top-down (from system to use) and bottom-up (from use to system) phenomena and their dynamics. The primary focus will be the lexicon, the secondary focus the representation of meaning vs grammar in LLMs.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Gemma Boleda

Gemma Boleda is an ICREA Research Professor in the Department of Translation and Language Sciences at Universitat Pompeu Fabra, Barcelona, where she co-directs the Computational Linguistics and Linguistic Theory group. Her research combines computational and distributional semantics with cognitive and linguistic theory to study how word meaning is grounded, represented, and used.

References:

Boleda, G. (2020). Distributional semantics and linguistic theory. Annual Review of Linguistics, 6, 213–234.

Gualdoni, E., & Boleda, G. (2024). Why do objects have many names? A study on word informativeness in language use and lexical systems. Proceedings of EMNLP 2024.

Boleda, G. (2025). LLMs as a synthesis between symbolic and distributed approaches to language. Findings of EMNLP 2025.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Multisensory Robot Reasoning

Jesse Thomason — Georgia Institute of Technology

March 18, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: awaiting date re-confirmation

Abstract: please provide 50-100-word abstract






The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Jesse Thomason

Jesse Thomason is an Associate Professor at the Georgia Institute of Technology’s School of Interactive Computing, where he leads the GLAMOR Lab (Grounding Language in Actions, Multimodal Observations, and Robots). His research integrates language, perception, and action, focusing on robot learning, natural-language instruction following, and multisensory grounding for embodied AI systems.

References:

Zhang, X., Amiri, S., Sinapov, J., Thomason, J., Stone, P., & Zhang, S. (2023). Multimodal embodied attribute learning by robots for object-centric action policies. Autonomous Robots.

Singh, I., Blukis, V., Mousavian, A., et al., & Garg, A. (2023). ProgPrompt: Generating situated robot task plans using large language models. IEEE ICRA.

Mitra, C., Anwar, A., Corona, R., Klein, D., Darrell, T., & Thomason, J. (2024). Which one? Leveraging context between objects and multiple views for language grounding. Proceedings of NAACL 2024.

Zhang, J., Luo, Y., Anwar, A., et al. (2025). ReWiND: Language-guided rewards teach robot policies without new demonstrations. Conference on Robot Learning (CoRL 2025).

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Boas, Shannon, and the origin of semantic categories

Terry Regier — UC Berkeley

March 25, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

Abstract: please provide 50-100-word abstract






The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Terry Regier

Terry Regier is Professor in the Department of Linguistics and the Cognitive Science Program at UC Berkeley. His research combines computational modeling and cross-linguistic analysis to explain how semantic categories — including color terms and kinship terms — arise from pressures toward efficient communication. He holds an honorary doctorate from the University of Gothenburg.

References:

Kemp, C., Xu, Y., & Regier, T. (2018). Semantic typology and efficient communication. Annual Review of Linguistics, 4, 109–128.

Zaslavsky, N., Kemp, C., Regier, T., & Tishby, N. (2018). Efficient compression in color naming and its evolution. Proceedings of the National Academy of Sciences, 115(31), 7937–7942.

Regier, T., Kay, P., & Khetarpal, N. (2007). Color naming reflects optimal partitions of color space. Proceedings of the National Academy of Sciences, 104(4), 1436–1441.

Zaslavsky, N., Regier, T., Tishby, N., & Kemp, C. (2019). Semantic categories of artifacts and animals reflect efficient coding. Proceedings of the 41st Annual Meeting of the Cognitive Science Society.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Meaning in brains vs. LLMs

Anna Ivanova — Georgia Institute of Technology

April 1, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

Abstract: please provide 50-100-word abstract






The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Anna Ivanova

Anna Ivanova is Assistant Professor of Psychology at the Georgia Institute of Technology, where she directs the Language, Intelligence, and Thought (LIT) Lab. Her research compares how meaning is represented in the human brain and in large language models, combining fMRI neuroimaging with computational/NLP methods to probe language, world knowledge, and cognition in both humans and machines.

References:

Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2024). Dissociating language and thought in large language models. Trends in Cognitive Sciences, 28(6), 517–540.

Ivanova, A. A., Mineroff, Z., Zimmerer, V., Kanwisher, N., Varley, R., & Fedorenko, E. (2021). The language network is recruited but not required for nonverbal event semantics. Neurobiology of Language, 2(2), 176–201.

Ivanova, A. A., Schrimpf, M., Anzellotti, S., Zaslavsky, N., Fedorenko, E., & Isik, L. (2022). Beyond linear regression: Mapping models in cognitive neuroscience should align with research goals. Neurons, Behavior, Data Analysis, and Theory.

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Toward Active Visual Intelligence

Gabriel Sarch — Princeton University

April 8, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: awaiting date re-confirmation

Visual problem solving is active. Effective reasoners choose what evidence to inspect, which strategies to pursue, and how to adapt when an attempt fails. Yet vision-language models are typically optimized for correct answers, with much less attention to the reasoning strategies behind them. I will show that multi-task reinforcement learning induces distinct strategies across visual task categories, that giving models tools to interact with visual evidence improves generalization, and that learning from repeated attempts after failure can produce effective exploration strategies where conventional reinforcement learning stalls. Finally, I will show how models can learn from human eye-tracking data to improve how they seek out and reason about visual evidence.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Gabriel Sarch

Gabriel Sarch is a Postdoctoral Research Fellow in Princeton Language and Intelligence at Princeton University. His research focuses on active perception and visual reasoning, combining reinforcement learning with insights from human cognition. He received his joint Ph.D. in Neural Computation and Machine Learning from Carnegie Mellon University in 2025.

References:

Sarch, G. H., Cai, L., Wang, Q., Wu, H., Chen, D., & Liu, Z. (2026). Vero: An open RL recipe for general visual reasoning. European Conference on Computer Vision (ECCV).

Sarch, G. H., Saha, S., Khandelwal, N., Jain, A., Tarr, M. J., Kumar, A., & Fragkiadaki, K. (2025). Grounded reinforcement learning for visual reasoning. Advances in Neural Information Processing Systems (NeurIPS).

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Language-guided learning in embodied agents

Joyce Chai — University of Michigan

April 15, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: date confirmed

Abstract: please provide 50-100-word abstract






The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of Joyce Chai

Joyce Chai is Professor of Computer Science and Engineering at the University of Michigan, where she directs the SLED (Situated Language and Embodied Dialogue) Lab. An ACL Fellow, her research focuses on grounded language processing, human-robot communication, and embodied agents that learn language through situated interaction.

References:

She, L., & Chai, J. (2017). Interactive learning of grounded verb semantics towards human-robot communication. Proceedings of ACL 2017 (Vol. 1).

Chai, J. Y., Gao, Q., She, L., Yang, S., Saba-Sadiya, S., & Xu, G. (2018). Language to action: Towards interactive task learning with physical agents. Proceedings of IJCAI-18 (pp. 2–9).

Zhang, Y., Yang, J., Pan, J., et al., & Chai, J. (2022). DANLI: Deliberative agent for following natural language instructions. Proceedings of EMNLP 2022 (pp. 1280–1298).

Xi, J., He, Y., Yang, J., Dai, Y., & Chai, J. (2024). Teaching embodied reinforcement learning agents: Informativeness and diversity of language use. Proceedings of EMNLP 2024 (pp. 4097–4114).

↑ back to calendar

Séminaire en Informatique Cognitive DIC/ISC/CRIA Seminar in Cognitive Informatics

2026-2027 Seminar theme: Polymodal AI: Top/Down or Bottom/UP ?
Full year programme: https://skywritingspress.ca/?page_id=1384

Grounding Speech in Multimodal Perception

David Harwath — UT Austin

April 22, 2027, 10:30am – noon ET (speakers join at 10:15am)
ZOOM: please email HERE for the zoom link
Université du Québec à Montréal (UQÀM), PK-5115
Status: awaiting date re-confirmation

Current speech processing models can perform tasks such as speech recognition and translation with a high degree of accuracy. They rely on large-scale corpora of transcribed speech for training; their application is therefore limited to the small subset of languages that can support large-scale data curation. In contrast, humans acquire spoken language from a comparatively small amount of speech data before they learn to read and write. I will describe our ongoing work to develop models of spoken language that do not learn from conventional text annotations, instead using perceptual grounding (to the visual modality) as a learning signal, and how the learned representations implicitly capture linguistic structure from the raw speech waveform.

The video of the talk will be available shortly after the talk at https://dic.uqam.ca/seminaire-dic/

Photo of David Harwath

David Harwath is an associate professor in the computer science department, University of Texas at Austin, where he leads the Speech, Audio, and Language Technologies (SALT) Lab. His group’s research focuses on developing novel machine learning methods applied to speech, audio, and multimodal data.

References:

Harwath, D., Recasens, A., Surís, D., Chuang, G., Torralba, A., & Glass, J. (2018). Jointly discovering visual objects and spoken words from raw sensory input. Proceedings of ECCV.

Berry, L., Shih, Y.-J., Wang, H.-F., Chang, H.-J., Lee, H., & Harwath, D. (2023). M-SpeechCLIP: Leveraging large-scale, pre-trained models for multilingual speech to image retrieval. Proceedings of ICASSP.

Peng, P., Li, S.-W., Räsänen, O., Mohamed, A., & Harwath, D. (2023). Syllable discovery and cross-lingual generalization in a visually grounded, self-supervised speech model. Proceedings of INTERSPEECH.

Peng, P., & Harwath, D. (2022). Word discovery in visually grounded, self-supervised speech models. Proceedings of INTERSPEECH.

↑ back to calendar