Skip to content
Home
Home
Research
Teaching
CV

Denny Wu

Denny Wu

New York University
· Center for Data Science
· Courant Institute of Mathematical Sciences
Flatiron Institute
· Center for Computational Mathematics

Contact:

dwu [at] flatironinstitute (dot) org

About Me

I am a faculty fellow at the Center for Data Science, New York University and the Flatiron Institute. My research focuses on the mathematical foundations of modern machine learning. I aim to develop a predictive theory of how neural networks learn, generalize, and scale, with particular interests in optimization, feature learning, and scaling laws.

I received my Ph.D. in computer science from the University of Toronto and the Vector Institute, advised by Jimmy Ba and Murat A. Erdogdu. Before that, I was an undergraduate at Carnegie Mellon University, where I was advised by Ruslan Salakhutdinov.

[CV] [Google Scholar]

Selected Publications

  • Sharp capacity scaling of spectral optimizers in learning associative memory.
    Juno Kim, Eshaan Nichani, Denny Wu, Alberto Bietti, and Jason D. Lee.

  • Understanding the mechanisms of fast hyperparameter transfer.
    Nikhil Ghosh, Denny Wu, and Alberto Bietti.

  • Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws.
    Gérard Ben Arous, Murat A. Erdogdu, N. Mert Vural, and Denny Wu.

  • Emergence and scaling laws in SGD learning of shallow neural networks.
    Yunwei Ren, Eshaan Nichani, Denny Wu, and Jason D. Lee.

  • Learning compositional functions with transformers from easy-to-hard data.
    Zixuan Wang, Eshaan Nichani, Alberto Bietti, Alex Damian, Daniel Hsu, Jason D. Lee, and Denny Wu.

  • Propagation of chaos in one-hidden-layer neural networks beyond logarithmic time.
    Margalit Glasgow, Denny Wu, and Joan Bruna.

  • Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit.
    Jason D. Lee, Kazusato Oko, Taiji Suzuki, and Denny Wu.

  • Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations.
    Kazusato Oko, Yujin Song, Taiji Suzuki, and Denny Wu.

  • Nonlinear spiked covariance matrices and signal propagation in deep neural networks.
    Zhichao Wang, Denny Wu, and Zhou Fan.

  • Learning in the presence of low-dimensional structure: a spiked random matrix perspective.
    Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, and Denny Wu.

  • High-dimensional asymptotics of feature learning: how one gradient step improves the representation.
    Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang.

  • Convex analysis of the mean-field Langevin dynamics.
    Atsushi Nitanda, Denny Wu, and Taiji Suzuki.

  • When does preconditioning help or hurt generalization?
    Shun-ichi Amari, Jimmy Ba, Roger Grosse, Xuechen Li, Atsushi Nitanda, Taiji Suzuki, Denny Wu, and Ji Xu.

  • On the optimal weighted ℓ₂ regularization in overparameterized linear regression.
    Denny Wu and Ji Xu.

  • Stochastic Runge-Kutta accelerates Langevin Monte Carlo and beyond.
    Xuechen Li, Denny Wu, Lester Mackey, and Murat A. Erdogdu.


Misc.

I like Wes Anderson.

I also like Japanese food.

Page generated by jemdoc.