Denny Wu
|
New York University Contact:dwu [at] flatironinstitute (dot) org |
About Me
I am a faculty fellow at the Center for Data Science, New York University and the Flatiron Institute. My research focuses on the mathematical foundations of modern machine learning. I aim to develop a predictive theory of how neural networks learn, generalize, and scale, with particular interests in optimization, feature learning, and scaling laws.
I received my Ph.D. in computer science from the University of Toronto and the Vector Institute, advised by Jimmy Ba and Murat A. Erdogdu. Before that, I was an undergraduate at Carnegie Mellon University, where I was advised by Ruslan Salakhutdinov.
[CV] [Google Scholar]
Selected Publications
Sharp capacity scaling of spectral optimizers in learning associative memory.
Juno Kim, Eshaan Nichani, Denny Wu, Alberto Bietti, and Jason D. Lee.
Understanding the mechanisms of fast hyperparameter transfer.
Nikhil Ghosh, Denny Wu, and Alberto Bietti.
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws.
Gérard Ben Arous, Murat A. Erdogdu, N. Mert Vural, and Denny Wu.
Emergence and scaling laws in SGD learning of shallow neural networks.
Yunwei Ren, Eshaan Nichani, Denny Wu, and Jason D. Lee.
Learning compositional functions with transformers from easy-to-hard data.
Zixuan Wang, Eshaan Nichani, Alberto Bietti, Alex Damian, Daniel Hsu, Jason D. Lee, and Denny Wu.
Propagation of chaos in one-hidden-layer neural networks beyond logarithmic time.
Margalit Glasgow, Denny Wu, and Joan Bruna.
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit.
Jason D. Lee, Kazusato Oko, Taiji Suzuki, and Denny Wu.
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations.
Kazusato Oko, Yujin Song, Taiji Suzuki, and Denny Wu.
Nonlinear spiked covariance matrices and signal propagation in deep neural networks.
Zhichao Wang, Denny Wu, and Zhou Fan.
Learning in the presence of low-dimensional structure: a spiked random matrix perspective.
Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, and Denny Wu.
High-dimensional asymptotics of feature learning: how one gradient step improves the representation.
Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang.
Convex analysis of the mean-field Langevin dynamics.
Atsushi Nitanda, Denny Wu, and Taiji Suzuki.
When does preconditioning help or hurt generalization?
Shun-ichi Amari, Jimmy Ba, Roger Grosse, Xuechen Li, Atsushi Nitanda, Taiji Suzuki, Denny Wu, and Ji Xu.
On the optimal weighted ℓ₂ regularization in overparameterized linear regression.
Denny Wu and Ji Xu.
Stochastic Runge-Kutta accelerates Langevin Monte Carlo and beyond.
Xuechen Li, Denny Wu, Lester Mackey, and Murat A. Erdogdu.