Time: 4:30-5:30 p.m.
Date: Wednesday, September 3, 2026
Location: 350 Computing and Information Science Building
Speaker: Yuqi Gu, Assistant Professor, Department of Statistics, member of the Data Science Institute, Columbia University
Title: Statistical Foundations of Identifiable Latent Representations: From Deep Generative Models to Causal Structure
Abstract: Identifiability is a fundamental requirement for interpretable representation learning: if very different latent explanations can generate the same observed data, the meaning of a learned representation, and any scientific conclusions built on it, can be unstable. Modern representation learning therefore raises a basic statistical question: when is hidden structure uniquely determined by complex, high-dimensional observations?
In this talk, I will develop a statistical approach to this question using discrete latent representations. I will first introduce Deep Discrete Encoders, a class of deep generative models whose layered latent structure can be identified under interpretable conditions, while retaining flexible modeling and scalable estimation. Applications to images, text data, and multimodal educational assessment data illustrate how the learned representations can capture meaningful hierarchical structure. I will then ask a harder question: once latent factors are identifiable, can we also learn how they causally influence one another? Discrete Causal Representation Learning shows that, under suitable structure, noisy and heterogeneous observations can reveal both how latent factors generate the data and the causal structure among those factors that can be recovered from observational data alone. Together, these results provide a path from flexible representation learning to statistically grounded interpretation, causal discovery, and ultimately intervention.
Bio: Yuqi Gu is an Assistant Professor in the Department of Statistics and a member of the Data Science Institute at Columbia University. Her research develops statistical theory and methodology for learning interpretable latent structure from complex, high-dimensional data, with a particular emphasis on identifiability, high-dimensional inference, and representation learning. Her recent work studies identifiable deep generative models, causal representation learning, and statistical measurement and evaluation of modern AI systems. She received her Ph.D. in Statistics from the University of Michigan in 2020 and was a postdoctoral researcher at Duke University before joining Columbia in 2021.