Prof. Gemma Boleda

Research Professor
Universitat Pompeu Fabra / ICREA
website

Title: LLMs as a synthesis between symbolic and distributed approaches to language

Abstract: Since the middle of the 20th century, a fierce battle is being fought between symbolic and distributed approaches to language and cognition. In this talk, I reflect on deep learning models of language in the context of this broader discussion, and consider the possibility that LLMs represent a synthesis between these two antagonistic traditions. This possibility is supported by the fact that deep learning architectures allow for both distributed/continuous/fuzzy and symbolic/discrete/categorical-like representations and processing; what is left to see is how modern language models make use of this flexibility. I will discuss recent research in interpretability that suggests that grammar is processed in a near-symbolic fashion in modern LLMs, while lexical semantics receives a much more distributed treatment; and I will also discuss work in progress in our group that is starting to probe this view. I will end the talk by providing some results from a recent systematic review of interpretability work on syntactic knowledge in LLMs that we have carried out.

Short bio: Gemma Boleda is an ICREA Research Professor in the Department of Translation and Language Sciences of the Universitat Pompeu Fabra, where she co-leads the Computational Linguistics and Linguistic Theory (COLT) research group. She previously held post-doctoral positions at the University of Trento (Italy), The University of Texas at Austin (USA), and Universitat Politècnica de Catalunya (Spain). In her research, Prof. Boleda uses quantitative and computational methods to investigate how languages work; in particular, how they are shaped by human cognitive constraints, on the one hand, and communicative needs, on the other.


Prof. Ryan Cotterell

Assistant Professor
ETH Zürich
website

Title: Surprisal Theory is Tautological (without Rational Grounding)

Abstract: Surprisal theory posits that the processing difficulty of a word in context is an affine function of its surprisal—the negative log-probability assigned by a language model. The theory has been enormously popular, and empirical studies continue to show that surprisal predicts reading times across languages, paradigms, and model architectures. But I argue that this success is less informative than it appears. When the language model is drawn from an unrestricted family—as is now standard practice—any pattern of processing difficulty is consistent with some model’s surprisal, rendering the theory unfalsifiable. I develop this argument rigorously, attending to mathematical subtleties hidden in the formal statement. I then draw an analogy to Darwin’s theory of evolution, where Popper famously argued that “survival of the fittest” was also unfalsifiable until fitness was defined independently of reproductive success. Applying the same logic to psycholinguistics, I propose constraining the model family using predictive information—an information-theoretic measure of how much a language model remembers about the past for future prediction. Bounding predictive information at a cognitively plausible level, calibrated from independent psycholinguistic experiments on working memory rather than from reading times, restores falsifiability: the theory can now fail, and a good fit tells us something about cognition. I also discuss the BabyLM Challenge, a community effort I helped spawn to train language models on developmentally plausible amounts of data, as a complementary approach to grounding the theory. Zooming out, I believe this talk is an exercise in formal functionalism—using the mathematical tools of formal language theory and information theory to make functional claims about human language processing precise, testable, and scientifically meaningful.

Short bio: Ryan Cotterell is an assistant professor of computer science at ETH Zürich in the Institute for Machine Learning, where he has been since 2020. Previously, he was a lecturer at the University of Cambridge. He received his PhD in computer science from Johns Hopkins University, advised by Jason Eisner, and holds a BS in cognitive science from the same institution with a focus on linguistics. With respect to machine learning, he focuses on the mathematical foundations of large language models, interpretable methods in machine learning, and approximate sampling procedures. With respect to linguistics, he works on mathematical linguistics (e.g., formal properties of grammar formalisms), information-theoretic and functionalist theories of diachronic linguistic data. He has been a co-organizer of the BabyLM Challenge, SIGTYP, and SIGMORPHON workshops, among others. He regularly serves as a reviewer, area chair, and senior area chair at major NLP and machine learning venues. He publishes at both NLP venues (ACL, EMNLP, NAACL) and machine learning venues (NeurIPS, ICML, ICLR), and has received several best paper awards, including the best paper award at ACL 2017 and ACL 2026 as well as one of the two Outstanding papers at ICLR 2026.