Data Geometry and Deep Learning
This is a postgraduate course aimed at Mathematics students interested in the theory and mathematics of Deep Learning.
Abstract. The idea is to give the basic ingredients to understand what Deep Learning theory is, and what the underlying mathematical/theoretical structures are. Hopefully at the end of the course, the mathematics students that follow it will be able to read Deep Learning research papers without being lost, and will have the basic tools to start working on some important maths-rich topics. Lectures 7–12 below each cover an important direction of research for which some first mathematical steps have been done, but which has many subtopics and extensions left open. Particular emphasis is given to the view that the geometry of datapoints and of learning algorithms has important practical consequences, some of which have started to emerge in the last 3–4 years.
- Dates
- November 15th 2022 – December 20th 2022
- Timetable
- Tuesdays 3–5 PM, Thursdays 3–5 PM
- Place
- Math Department "Guido Castelnuovo" of "La Sapienza" University, Room B (first floor)
Links to slides will be posted on this page. Please fill this short form (2 mins) if you are interested in being added to the mailing list of the course!
The plan is to spend one lecture on each topic (the plan below may change as the course progresses). Note that the slides are intended as study material: clickable links in the slides will send you to more detailed references.
Part I: Introduction to Deep Learning
- Introduction and a brief history of Neural Networks. Overview of the course. (Slides) Video recordings (in Italian): part 1, 6 min and part 2, 80 min
- Backpropagation, Stochastic Gradient Descent, Convergence. (Slides) (Video recording (in Italian), 100 min)
- Some very common Neural Network architectures and their motivations. (Slides) (Video (Ita), 120 min)
Neuromorphic Neural Networks(skipped)
Part II: Staples of classical Deep Learning theory
- Generalization: some mathematical interpretations. (Slides) (Video (Ita), 110 min)
- PAC learning, VC dimension, and expressivity tests for DNN. (Slides) (Video part 2 (Ita), 50 min)
- Introduction to Information Theory, and the Information Bottleneck Principle. (Slides) (Video part 1, 45 min, part 2, 35 min)
No class on Dec. 1st.
Part III: Selected topics of research
- Network pruning: the "Lottery ticket hypothesis", and sparsity. (Slides) (Video in Italian, 90 min)
No class on Dec. 8th.
- Hyperbolic Neural Networks. (Slides)
- Equivariant Neural Networks. (Slides)
- Persistence diagrams and Topological Data Analysis (guest lecture by Sara Scaramuccia). (Slides)
Final presentations by students
- Luca Falorsi: Renormalization Group and Restricted Boltzmann Machines
- Alessio Oliviero: Physics Informed Neural Networks for Optimal Control (Slides)
- Jacopo Ulivelli: Geometric interpretation of GANs and link with Optimal Transport (link to exposed paper)
- Andrea Pizzi: Graph Neural Networks generalized to simplicial complexes (Slides)
- Lorenzo D'Arca: Teoremi di approssimazione in spazi funzionali (Slides) (Slides in English, paper 1, paper 2)
Bibliography
Given before the course start, for the first 6 lectures (see precise references within each lecture's slides).
- History began with AlexNet: A comprehensive survey on Deep Learning, arXiv:1803.01164 (useful for lectures 1 and 3)
- Deep Learning Book, deeplearningbook.org (useful for lectures 1–5, but especially for lectures 2 and 3)
- Some famous Deep Learning architectures, towardsdatascience.com
- Foundations of Machine Learning, cs.nyu.edu/~mohri/mlbook (useful for lectures 5–6)
- The modern mathematics of deep learning, arXiv:2105.04026