ORF 526 · Princeton · Fall 2026

Probability for
Modern Machine Learning

The Gaussian The semicircle
MW 10:40–12:00 24 sessions Sep 2 – Dec 7 No measure theory

The course

A graduate probability course built out of tools and examples rather than a slow development of the foundations.

Twenty-four sessions, in two modules. Each one is organized around a technique worth owning — Lindeberg swapping, Stein’s method, Wick’s theorem, the resolvent, exponential tilting — and around an example that shows what the technique is for. A substantial part of that second half is modern machine learning: wide networks at initialization, kernel regression, the spectrum of a Gaussian random matrix, the maxima of Gaussian processes, stochastic gradient descent, and diffusion models.

No measure theory. Nothing is assumed beyond linear algebra, multivariable calculus and an undergraduate probability course. Where a standard fact from analysis is needed — Fourier inversion, the spectral theorem — it is quoted honestly and marked as such in the notes.

How it works

Homework, assigned every sessionnot graded
Weekly quiz, 15 min, in class20%
Midterm, in class30%
Final50%

The lowest quiz is dropped. Both examinations are closed book.

Grades

A90–100
B80–89
C70–79
D60–69
F0–59

On the homework and the quizzes. The exercises live inside the lecture notes, typed A, B and C: A checks that nothing was missed, B asks whether the argument was understood, and C withholds exactly one idea. They are not collected and not graded — use whatever help you like, including an AI assistant, and if you are short of time do the C problems. The weekly quiz is where that work is cashed in: closed book, by hand, drawn from the week’s assigned exercises. The homework is the exploration layer; the quiz is the verification layer, and it is the only thing that needs policing.

Module I · Gaussians 12 sessions · Sep 2 – Oct 14
1Sep 2
Gaussians: definition and linear structure
  • Definition by the Fourier transform, and why not the density
  • Linear images, existence, rotational invariance
  • Uncorrelated and jointly Gaussian implies independent
  • Maximum entropy at fixed covariance
notes
2Sep 9
Gaussians: conditioning as projection
  • Conditional expectation as orthogonal projection
  • The Schur complement, derived
  • Tower property and total variance, read off the picture
  • Bayesian posteriors and ridge regression; Tweedie's formula
notes
3Sep 14
CLT: Lindeberg's method
  • Maximum entropy, recalled; why a Gaussian at all
  • What convergence in distribution means, and why smooth test functions
  • Lindeberg's swapping argument, proved
  • Lindeberg's and Feller's conditions; the CLT for triangular arrays
  • Records: why random search improves only log n times, and why the Gaussian/Poisson dial is s_n
notes
4Sep 16
Stein's method
  • The Gaussian characterized by an identity, not a limit
  • Stein's equation; the CLT with a rate
  • Chen–Stein for Poisson; dependence
notes
5Sep 21
Wick's theorem, and a Gaussian random matrix
  • Every moment of a Gaussian vector, as a sum over pair partitions
  • Proved from the characteristic function — no density, no invertibility
  • The Gaussian Wigner matrix
  • ‖Mu‖² for a fixed direction: its mean and variance, by Wick
  • Cumulants, in the reading
notes
6Sep 23
The semicircle law by counting paths
  • The theorem, and the strategy: moments are traces
  • E tr M² and E tr M⁴ by hand
  • Matrix powers as sums over paths; from a path to a pair partition
  • Only tree paths survive; non-crossing pairings and the Catalan numbers
  • The fluctuations are small
notes
7Sep 28
The Stieltjes transform
  • The resolvent and mN(z) = (1/N) tr (M − zI)−1
  • Eigenvalues are the poles: counting them in an interval by a contour integral
  • What to compute: limε→0 limN→∞ (1/π) Im mN(x + iε), and why the limits come in that order
  • Near z = ∞ the transform is the generating function of the traces
notes
8Sep 30
The semicircle law by the Stieltjes transform
  • Schur complement: one diagonal entry of a resolvent through a smaller resolvent
  • The self-consistent equation m² + zm + 1 = 0, with explicit error bounds
  • The trace concentrates: a Poincaré inequality and a cancellation of N's
  • Stability of the equation, and the density read off the unit circle
notes
9Oct 5
Gaussian processes: definition, feature maps, and maxima
  • A Gaussian process is a Gaussian vector indexed by an arbitrary set; mean μt and kernel Σs,t
  • Feature maps: every kernel is an inner product; Mu on the sphere, Brownian motion, the wide network
  • The canonical metric
  • The union bound E max ≤ σ√(2 log n), with or without independence; the Gaussian width as homework
notes
10Oct 7
Gaussian processes: suprema and chaining
  • Independent Gaussians reach √(2 log n)
  • Covering numbers; Dudley's inequality stated, applied to the sphere, and proved by chaining
  • Sudakov's lower bound, stated
notes
11Oct 12
Gaussian processes: comparison
  • Sudakov–Fernique and Slepian by Gaussian interpolation
  • The operator norm of a Gaussian matrix: E‖M‖ ≤ √N + √d
  • Sudakov's lower bound, proved; what the midterm covers
notes
12Oct 14
Midterm examination
  • Sessions 1–11, closed book, in class
Module II · Dynamics, Deviations, Diffusion 12 sessions · Oct 26 – Dec 7
1Oct 26
Martingales: conditional expectation and definitions
  • Conditional expectation as projection
  • Filtrations and the tower property
  • The Doob decomposition
2Oct 28
Martingales: limit theorems
  • Doob's maximal inequality; convergence in L²
  • Optional stopping
  • The martingale central limit theorem
3Nov 2
Martingales: concentration
  • Azuma–Hoeffding and Freedman
  • Bounded differences
4Nov 4
SGD: stochastic approximation
  • Drift plus martingale difference
  • Robbins–Monro and the ODE method
5Nov 9
SGD: fluctuations and averaging
  • Asymptotic normality of the iterates
  • Polyak–Ruppert averaging
  • Concentration along the trajectory
6Nov 11
SGD: high-dimensional dynamics
  • A deterministic ODE plus a fluctuation term
  • The data covariance spectrum, through the resolvent
7Nov 16
Large deviations: energy versus entropy
  • Cramér's theorem by exponential tilting
  • The rate function as a competition
8Nov 18
Large deviations: Sanov's theorem
  • The method of types
  • Relative entropy as the rate function
  • The contraction principle
9Nov 23
Large deviations: Varadhan and examples
  • Varadhan's lemma and the Gibbs variational principle
  • Curie–Weiss and its phase transition
  • Hypothesis testing; escape from a basin
10Nov 30
Diffusion models: the forward process
  • Noising as a Gaussian channel
  • Tweedie's formula from integration by parts
  • Denoising as score estimation
11Dec 2
Diffusion models: score matching
  • The reverse chain
  • Denoising score matching
  • The DDPM objective
12Dec 7
Diffusion models: the PDE picture
  • Fokker–Planck and the heat equation
  • The Gaussian as fundamental solution
  • Time reversal; the variational bound