A day of scholarly exchange and community building

The inaugural Rising Stars in Statistics and Data Science workshop will spotlight emerging leaders, share new ideas, and strengthen connections across the broader research community. 

The workshop will take place Sept. 18, 2026, with arrival on Sept. 17, and will feature 11 invited talks, a curated poster session, and optional hiking activities on Saturday, Sept. 19. 
 

Tentative Agenda

Location: Computing and Information Science Building

Time
Activity
Thursday, Sept. 17Participant arrival
Friday, Sept. 18All day workshop
8:30 - 9:00 a.m.Breakfast and registration
9:00 - 9:05 a.m.Opening remarks
9:05 - 9:35 a.m.Ian Waudby-Smith: Advances in large-sample sequential inference
9:35 - 10:05 a.m.

Junu Lee: Compound p-values for competition-based testing

10:05 - 10:35 a.m.Jiadong Liang: Low-Dimensional Adaptation of Diffusion Generative Models
10:35 - 10:50 a.m. Morning Break
10:50 - 11:20 a.m.Yang Cao: MoDaH achieves embedding-level rate optimal batch correction
11:20 - 11:50 a.m.Eric Xia: Offline dynamic pricing: Estimation and Optimization
11:50 a.m. - 12:20 p.m.Yuli Slavutsky: From Fixed Covariates to Distribution Shift: Adaptive Bayesian Uncertainty
12:20 - 1:30 p.m.Lunch
1:30 - 2:00 p.m.Renyuan (Jack) Ma: The Anisotropic Local Law For Sample Covariance Matrices Under Quadratic-Form Concentration
2:00 - 2:30 p.m.Kangjie Zhou: When Does Model Collapse Occur in Structured Interactive Learning?
2:30 - 3:00 p.m.Margalit Glasgow: When do Infinite-Width Neural Networks explain Finite-Width Networks
3:00 - 3:15 p.m.Afternoon Break
3:15 - 3:45 p.m.Adam Jaffe: Wasserstein-Cramér-Rao Theory of Unbiased Estimation
3:45 - 4:15 p.m.Aram-Alexandre Pooladian: Trajectory inference via Acceleration Matching
4:15 - 4:20 p.m. Closing remarks
4:30 - 6:00 p.m.Poster Session & Reception
6:00 p.m. Dinner

Invited speakers:

Advances in large-sample sequential inference

Abstract: In this talk, I will discuss a theory of large-sample sequential statistical inference. Historically, anytime-valid tools like confidence sequences, sequential hypothesis tests, and anytime p-values have been justified non-asymptotically (i.e. in finite samples). While such guarantees are strong, they necessarily rely on strong moment assumptions and tend to be conservative.  By contrast, large-sample procedures such as those based on the central limit theorem occupy an important part of the statistical toolbox for their simplicity, universality, and sharpness, as well as for the weak assumptions that they impose. I will discuss some recent works that aim to bring those desiderata to sequential inference. In particular, I will discuss a Berry–Esseen-type inequality for anytime-valid inference that resembles the usual inequality for the central limit theorem up to a root-logarithmic factor.

Bio: Ian Waudby-Smith is a Miller postdoctoral fellow at the University of California, Berkeley where he is hosted in the Department of Statistics by Michael I. Jordan. Before joining UC Berkeley, he obtained his PhD in Statistics from Carnegie Mellon University where he was advised by Aaditya Ramdas and was awarded the Umesh K. Gavaskar Best Dissertation Award. He obtained his Bachelor's degree in mathematics from the University of Waterloo in Canada. His recent research interests include anytime-valid sequential inference, e-values, causal inference, concentration inequalities, and strong limit theorems. 

MoDaH achieves embedding-level rate optimal batch correction

Abstract: Batch effects pose a significant challenge in the analysis of single-cell omics data, introducing technical artifacts that confound biological signals. While various computational methods have achieved empirical success in correcting these effects, they lack the formal theoretical guarantees required to assess their reliability and generalization. To bridge this gap, we introduce Mixture-Model-based Data Harmonization (MoDaH), a principled batch correction algorithm grounded in a rigorous statistical framework. Under a new Gaussian-mixture-model with explicit parametrization of batch effects, we establish the minimax optimal error rates for batch correction at the embedding level and prove that MoDaH achieves this rate by leveraging the recent theoretical advances in clustering data from anisotropic Gaussian mixtures. This constitutes, to the best of our knowledge, the first theoretical guarantee for batch correction. Extensive experiments on diverse single-cell RNA-seq and spatial proteomics datasets demonstrate that MoDaH not only attains theoretical optimality but also achieves empirical performance comparable to or even surpassing those of state-of-the-art heuristics (e.g., Harmony, Seurat-V5, and LIGER), effectively balancing the removal of technical noise with the conservation of biological signal.

Bio: Yang Cao is a Postdoctoral Associate in the Department of Statistics and Data Science at Yale University. His research focuses on statistical inference and applications in single-cell and spatial multi-omics. He earned his PhD in Mathematics from the Hong Kong University of Science and Technology and his bachelor’s degree in Statistics from Peking University.

The Anisotropic Local Law For Sample Covariance Matrices Under Quadratic-Form Concentration

Abstract: We study sample covariance matrices $K = \frac{1}{N} \sum_{i=1}^N \x_i \x_i^* \in \R^{n \times n}$ in the proportional high-dimensional regime $n \asymp N$. The columns $\x_1, \ldots, \x_N \in \R^n$ are independent and centered, with common covariance $\E \x_i \x_i^* = \Sigma$, but may otherwise have strongly and nonlinearly dependent coordinates. Assuming only that the quadratic forms of the columns concentrate uniformly at the optimal rate $| \x_i^* A \x_i - \Tr \Sigma A | \prec \| A \|_F$, together with polynomial norm moments and a standard nondegeneracy condition, we prove the optimal anisotropic local law. More precisely, on regular spectral domains and uniformly down to spectral scales $\eta:= \Im z \geq N^{-1 + \eps}$,

    \[

    \big| \< \u , \big( (K-z)^{-1} - (-zI_n-z\widetilde m_0(z)\Sigma \big)^{-1} \bv \> \big| \prec \sqrt{\frac{\Im \widetilde m_0 (z)}{N\eta}} + \frac{1}{N\eta}

    \] 

    for any fixed unit vectors $\u,\bv \in \C^n$, where $\widetilde m_0(z)$ is the deterministic Stieltjes transform of the deformed Marchenko-Pastur law. This removes the higher-cumulant tensor assumption of Fan, Ma, Paquette, and Wang (2026), thereby answering the question raised in their work. The result applies, among other examples, to every centered log-concave column distributions with bounded, nondegenerate covariance, nonlinear tilts of Gaussian vectors, deep random features, and a high-temperature spherical 4-spin model for which the cumulant assumption is known to fail.

Bio: I’m Renyuan (Jack) Ma, a 6th year PhD student at Yale SDS supervised by Prof. Zhou Fan and Prof. Hemant Tagare. My research interests lies in Random matrix theory, free probability, statistical physics, and their applications to solve problems in high dimensional statistics and machine learning. 

Trajectory inference via Acceleration Matching

Abstract:Trajectory inference is a fundamental problem in many scientific domains: given a collection of unpaired snapshots of observations at discrete time points, the goal is to generate smooth trajectories that best resemble and interpolate the data. Existing algorithms exhibit computational challenges: they either rely on preprocessing subroutines to enforce smoothness or on simulation-based training objectives, both of which can be expensive. In order to overcome these limitations, we propose a new algorithm called Acceleration Matching (\texttt{AM}). Our approach consists of lifting the original interpolation problem to phase space and then regressing onto an explicit conditional acceleration field that induces random, smooth trajectories that agree with the prescribed marginals. Importantly, our resulting training algorithm only requires positional data, avoids trajectory simulation during training, and is devoid of expensive preprocessing. We provide ample numerical evidence suggesting that \texttt{AM} is competitive with or superior to existing algorithms on several benchmark problems from the existing literature.

Bio: Aram-Alexandre Pooladian is a Foundations of Data Science Postdoctoral Associate at Yale University. His research surrounds mathematical statistics and algorithms for probabilistic inference, the latter through the lens of optimal transport. He holds a PhD from New York University Center for Data Science, and BA and MSc in Applied Mathematics from McGill University.

Offline dynamic pricing: Estimation and Optimization

Abstract:Dynamic pricing revolves around understanding the (conditional) distributional properties of some latent utility variable, based on observational data involving covariates, an offered price, and an indicator for whether that utility exceeds the offered price. Two commonly studied problems involves estimating the conditional median of this utility variable, and learning a price-offering policy that maximizes the expected revenue. Much of this literature has focused on settings where the distribution of the utility is smooth; however this contradicts much of the work from the business community which reports that utilities are quite discrete in nature. We establish, amongst other results, that under certain forms of discreteness on the utilities, the maximum score estimator can achieve a super-consistent estimation rate of $O_\PP(\frac{1}{\numobs})$. These estimation guarantees are extended to solve the policy learning problem in batch settings. We then address the computational challenges associated with the maximum score estimator through the use of convex surrogates. 

Bio: Eric Xia is a postdoctoral research associate in the ORFE department at Princeton University. His work focuses on statistical theory and methodology for analyzing nonstandard data regimes, such as classification imbalance, and surrogate-aided prediction. He received his PhD in Electrical Engineer and Computer Science from MIT in 2025. 

Wasserstein-Cramér-Rao Theory of Unbiased Estimation

Abstract:The quantity of interest in the classical Cramér-Rao theory of unbiased estimation (e.g., the Cramér-Rao lower bound, its exact attainment for exponential families, and asymptotic efficiency of maximum likelihood estimation) is the variance, which represents the instability of an estimator when its value is compared to the value for an independently-sampled data set from the same distribution. In this talk we are interested in a quantity which represents the instability of an estimator when its value is compared to the value for an infinitesimal additive perturbation of the original data set; we refer to this as the "sensitivity" of an estimator. The resulting theory of sensitivity is based on the Wasserstein geometry in the same way that the classical theory of variance is based on the Fisher-Rao (equivalently, Hellinger) geometry, and this insight allows us to determine a collection of results which are analogous to the classical case: a Wasserstein-Cramér-Rao lower bound for the sensitivity of any unbiased estimator, a characterization of models in which there exist unbiased estimators achieving the lower bound exactly, and general result on the asymptotic sensitivity-efficiency of Wasserstein projection estimators. We use these results to treat many statistical examples, sometimes revealing new optimality properties for existing estimators and other times revealing entirely new estimators. Based on joint work with Nicolás García Trillos and Bodhisattva Sen.

Bio: Adam is a postdoctoral researcher in the Department of Statistics at Columbia University, working under the supervision of Bodhisattva Sen. His research focuses on interactions between probability, statistics, and geometry, including clustering, dimensionality reduction, empirical Bayes, Fréchet means, information geometry, optimal transport, and Riemannian optimization. Previously, he completed his PhD at UC Berkeley and his undergraduate at Stanford.

When Does Model Collapse Occur in Structured Interactive Learning?

Abstract: The proliferation of generative artificial intelligence has given rise to an interactive learning environment, where model parameters are continuously updated using not only data generated by natural processes, but also synthetic outputs produced by other models. This paradigm introduces two major challenges: (1) training data are no longer drawn exclusively from the target population, undermining a core assumption of classical statistical learning, and (2) model training processes become inherently correlated, as models interact with one another through repeated exposure to each other's synthetic outputs in a potentially complex manner. Establishing reliable statistical inference in such structured interactive learning environments therefore remains an important open problem. In particular, there is growing concern about model collapse, a phenomenon in which the performance of generative models progressively degrades as they are trained on synthetic data produced by earlier model generations. In this talk, we develop a framework for investigating the performance of generative models in an interactive learning environment with general interaction patterns. In particular, we formalize model interactions using directed graphs and show that the occurrence of model collapse depends critically on the topology of the interaction graph. We further derive an explicit necessary and sufficient condition characterizing when model collapse occurs, and establish finite-sample results for linear regression and asymptotic guarantees for general empirical risk minimizers.

Bio: Kangjie Zhou is an assistant research professor at the Center for Data Science for Enterprise and Society, Cornell University. He was a Founder’s Postdoctoral Fellow at the Department of Statistics, Columbia University. In 2024, he received PhD in Statistics from Stanford University, advised by Andrea Montanari. His research lies at the intersection of high-dimensional statistics and probability, non-convex optimization, and deep learning, with an emphasis on the mathematical and algorithmic foundations of modern AI.

When do Infinite-Width Neural Networks explain Finite-Width Networks

Abstract:A longstanding question in deep learning theory asks why gradient descent (GD) finds good solutions in non-convex neural-network optimization landscapes. One compelling theory is that overparameterization makes the optimization landscape benign, leading GD to find global optima. These global convergence guarantees have even been proven rigorously in some settings for infinite-width neural networks. In this talk, I'll address the question of when such infinite-width convergence guarantees can be transferred to finite-width 2-layer networks. I'll show that whenever the convergence rate of the infinite-width network is faster than 1/t^2, comparable guarantees can be attained in finite-width networks. A key takeaway of our result is that whenever the convergence rate of the infinite-width, population-loss dynamics is faster than 1/t^2 , we can attain a loss of ϵ with only poly(d/ϵ) neurons, training samples, and GD steps.

Bio: Margalit Glasgow is a postoc at MIT advised by Sasha Rakhlin.  Before that, she completed her PhD in Computer Science at Stanford university, advised by Mary Wootters and Tengyu Ma.  She is broadly interested in theoretical machine learning and high dimensional probability, with a focus on deep learning theory.

From Fixed Covariates to Distribution Shift: Adaptive Bayesian Uncertainty

Abstract: Modern AI systems offer impressive predictive abilities, but often struggle to provide reliable uncertainty estimates. A particular challenge arises when new inputs differ from those seen during training, as can happen, for example, when broadly pretrained models are adapted to specialized downstream tasks. Classical Bayesian methods provide a principled framework for uncertainty quantification, but do not directly address this setting: a new covariate changes where the posterior predictive distribution is evaluated, but not the posterior over the predictive mechanism itself. At the same time, methods designed to account for distribution shifts respond to covariate shifts, even when the predictive ability remains unchanged.

In this talk, I will introduce covariate-dependent Bayesian models, in which the prior on model parameters depends on the training and queried covariates (but not on the outcomes). We will then discuss how defining this prior through the predictive model’s likelihood yields uncertainty estimates that respond to shifts affecting predictive performance, rather than to distributional change alone.

Bio: Yuli Slavutsky is a Founder’s Postdoctoral Research Scientist in the Department of Statistics at Columbia University. Her research develops Bayesian and variational methods for reliable prediction, with a focus on uncertainty quantification and robust representation learning. She received her PhD in Statistics and Data Science from the Hebrew University of Jerusalem, advised by Yuval Benjamini, and her master’s degree in Statistics also from the Hebrew University, advised by Or Zuk. Prior to her PhD, she worked in industry as a data scientist and research team leader.

Compound p-values for competition-based testing

Abstract: Standard competition-based multiple testing methods—such as the knockoff filter and the target-decoy procedure—are typically formulated as head-to-head competitions between a test unit and its assigned negative control unit. Motivated by real-world applications, which leverage any available negative controls, we study this competition-based framework in more intricate settings where competitions vary in size and structure. We propose new methods based on compound p-values—p-values that are superuniform only on aggregate—that enable powerful multiple testing with broader competition structures, via the Bonferroni and Benjamini–Hochberg procedures. The former provides exact familywise error rate control, whereas the latter may incur a false discovery rate inflation factor, for which we derive upper bounds across a range of settings. We showcase the performance of our general competition-based testing methodology through extensive simulations and real-data applications, including analyses of virtual spatial transcriptomics and CRISPR screen datasets that exhibit precisely the complex competition structures that motivate our methods

Bio: Junu Lee is a fifth-year Ph.D. student in the Department of Statistics and Data Science at the University of Pennsylvania, advised by Dr. Zhimei Ren and Dr. T. Tony Cai. Previously, he received an A.B. in Mathematics and Statistics and an S.M. in Computer Science from Harvard University. His research focuses on developing methods in multiple hypothesis testing and selective inference with model- and distribution-free guarantees, drawing on ideas such as e-values, knockoffs, and conformal inference to address the complexities of modern datasets. He also works on agentic data science systems, designing rigorous inference frameworks that leverage the capabilities of agentic AI for large-scale scientific discovery while preserving statistical validity.

Low-Dimensional Adaptation of Diffusion Generative Models

Abstract: Diffusion models have emerged as one of the most prominent approaches to generative modeling. They combine a ​``score-matching'' training procedure with a short sequence of iterative updates to transform pure noise into samples from a target distribution. Existing theoretical bounds for both iteration complexity and training sample complexity, however, depend heavily on the ambient dimension $d$ of the data space containing the target distribution, leaving a substantial gap between theory and the practical efficiency of diffusion models. In this talk, I will show how low-dimensional structure can close this gap. Under suitable low-dimensional assumptions on the target distribution, I will demonstrate, first for iteration complexity and then for training sample complexity, that the resulting bounds can depend only on the intrinsic dimension $k$, with no dependence on $d$. These results apply to a broad class of target distributions without requiring smoothness or log-concavity.

Bio: Jiadong Liang is a postdoctoral researcher in Statistics at the Wharton School, University of Pennsylvania, where he works with Professor Yuxin Chen. He received his PhD in Statistics from the School of Mathematical Sciences at Peking University, advised by Professor Zhihua Zhang. His research focuses on sequential statistical and machine learning problems, including online inference, diffusion generative models, and other sampling methods.

Poster presenters:

Awni Altabaa
Ph.D. Student, Department of Statistics and Data Science, Yale University
Research Areas: Foundations of language models; latent reasoning and representation learning; relational reasoning; transformer theory

Anna Brandenberger
Ph.D. Student, Department of Mathematics, MIT
Research Areas: Probability and random structures; random graphs and networks; high-dimensional and combinatorial probability

Yuchen Chen
Ph.D. Student, Department of Statistics & Data Science, Carnegie Mellon University
Research Areas: High-dimensional statistics; nonconvex learning dynamics; single-index models; statistical theory of machine learning

Daniil Dmitriev
Postdoctoral Researcher, Department of Statistics and Data Science, The Wharton School, University of Pennsylvania
Research Areas: High-dimensional learning theory; random features; robust mixture learning; discrete diffusion models; differential privacy

Avrajit Ghosh
Postdoctoral Fellow, Simons Institute for the Theory of Computing and BAIR, University of California, Berkeley
Research Areas: Optimization dynamics in deep learning; implicit bias and implicit regularization; generalization; inverse problems

Wenjie Guan
Ph.D. Student, Department of Statistics and Data Science, Cornell University
Research Areas: Statistical machine learning; theory of transformer reasoning and generalization; high-dimensional inference

Yu Gui
Postdoctoral Researcher, Department of Statistics and Data Science, The Wharton School, University of Pennsylvania
Research Areas: Selective and distribution-free inference; conformal inference; causal inference; multimodal representation learning

Rohan Hore
Postdoctoral Fellow, Department of Statistics, Stanford University
Research Areas: Distribution-free inference; conformal inference; multiple testing and variable selection; conditional-independence testing

Feiyang Yi
Ph.D. Student, Department of Statistics and Data Science, Cornell University
Research Area: Causal Inference

Seunghyun (Sky) Lee
Ph.D. Candidate, Department of Statistics, Columbia University
Research Areas: High-dimensional latent-variable models; empirical Bayes and variational inference; graphical models; statistical theory of generative models

Bernardo Marenco
Universidad de la República
Research Areas: Statistical network models; random graphs; graph representation learning

Kenta Takatsu
Columbia University
Research Areas: Causal inference; semiparametric and nonparametric inference; doubly robust machine learning; inference for stochastic optimization

Boyu Wang
Ph.D. Student, Department of Statistics and Data Science, Cornell University
Research Area: Statistics and data science

Yuepeng Yang
Postdoctoral Researcher, Department of Statistics and Data Science, The Wharton School, University of Pennsylvania
Research Areas: High-dimensional statistical learning; reinforcement learning; matrix completion and multi-matrix estimation; ranking