Our Breakstone Speaker series visitor (and last week’s colloquium speaker), Roni Katzir will be giving a mini-course this week as well:
Since the mid-1980s, artificial neural networks (ANNs) have been trained almost exclusively using a particular approach that has proven to be very useful for improving how the ANN fits its training data and, in turn, has been instrumental in the impressive engineering successes of ANNs on linguistic tasks over the past decades. ANNs trained using this method are typically extremely large, they require huge training corpora, and their inner workings are opaque. They also generalize in ways that seem inconsistent with common assumptions about rational inference. In these classes we will look at what happens when we replace the standard training approach for neural networks with Minimum Description Length (MDL), a simplicity principle that helps explain what makes some generalizations better than others. MDL has a long history in cognitive science, and among other things it provides a possible answer to how humans learn abstract grammars from unanalyzed surface data.
MDL also provides a way for machines to do the same: with MDL as the training method, we obtain small, transparent networks that learn complex recursive patterns perfectly and from very little data. These MDL networks help illustrate just how far standard ANNs (even the most successful of them) are from what we would expect from an intelligent system that attempts to extract regularities from the input data: given hundreds of billions of parameters and huge training corpora the performance of standard ANNs is sufficiently good to fool us on many common examples, but even then what the networks offer is just a superficial approximation of the regularities that reveals a complete lack of understanding of what what these regularities actually are. The MDL networks show us that it is possible for ANNs to learn intelligently and acquire systematic regularities perfectly from small training corpora but that this requires a very different learning approach from what current networks are based on.
Fox, D. and Katzir, R. (2024). Large language models and theoretical linguistics. Theoretical Linguistics, 50(1–2):71–76. https://doi.org/10.1515/tl-2024-2005
Lan, N., Geyer, M., Chemla, E., and Katzir, R. (2022). Minimum description length recurrent neural networks. Transactions of the Association for Computational Linguistics, 10:785–799. https://doi.org/10.1162/tacl_a_00489