Abstract
We study shallow neural networks with monomial activations and output dimension one. The function space for these models can be identified with a set of symmetric tensors with bounded rank. We describe general features of these networks, focusing on the relationship between width and optimization. We then consider teacher-student problems, which can be viewed as problems of low-rank tensor approximation with respect to nonstandard inner products that are induced by the data distribution. In this setting, we introduce a teacher-metric data discriminant which encodes the qualitative behavior of the optimization as a function of the training data distribution. Finally, we focus on networks with quadratic activations, presenting an in-depth analysis of the optimization landscape. In particular, we present a variation of the Eckart-Young theorem characterizing all critical points and their Hessian signatures for teacher-student problems with quadratic networks and Gaussian training data.
| Original language | English |
|---|---|
| Pages (from-to) | 174-209 |
| Number of pages | 36 |
| Journal | SIAM Journal on Applied Algebra and Geometry |
| Volume | 10 |
| Issue number | 2 |
| DOIs | |
| State | Published - 20 Apr 2026 |
Bibliographical note
Publisher Copyright:© 2026 Society for Industrial and Applied Mathematics
Keywords
- Eckart-Young theorem
- Polynomial neural networks
- data discriminant
- optimization landscape
- symmetric tensor rank
- teacher-student problems
Fingerprint
Dive into the research topics of 'Geometry and Optimization of Shallow Polynomial Networks'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver