Sampling the parameter space of artificial neural networks according to a Boltzmann distribution provides insight into the geometry of low-loss solutions and offers an alternative to conventional loss minimization for training. However, exact sampling methods such as hybrid Monte Carlo (hMC), while formally correct, become computationally prohibitive for larger datasets because they require repeated evaluation of full–batch gradients. We introduce a pseudo–Langevin (pL) dynamics that enables efficient Boltzmann sampling of feed-forward neural networks trained with large datasets by using mini–batches in a controlled manner. The method exploits the statistical properties of mini–batch gradient noise and adjusts fictitious masses and friction coefficients to ensure that the induced stochastic process samples efficiently the desired equilibrium distribution. We validate numerically the approach by comparing its equilibrium statistics with those obtained from exact hMC sampling. Performance benchmarks demonstrate that, while hMC rapidly becomes inefficient as network size increases, the pL scheme maintains the best computational diffusion among exact and non–exact methods and scales favorably to networks with over one million parameters. Additionally, we show that sampling at intermediate temperatures yields optimal generalization performance, comparable to stochastic gradient descent, without requiring a validation set or early stopping procedure. Finally, we show that our sampling method can achieve optimal generalization on the MNIST dataset, independently of the architecture. These results establish controlled mini–batch Langevin dynamics as a practical and scalable tool for exploring and exploiting the solution space of large neural networks.
Controlled Langevin dynamics for sampling of feedforward neural networks trained with minibatches / A. Zambon, F.C.. - In: JOURNAL OF STATISTICAL MECHANICS: THEORY AND EXPERIMENT. - ISSN 1742-5468. - 2026:7(2026 Jul 09), pp. 073401.1-073401.26. [10.1088/1742-5468/ae8246]
Controlled Langevin dynamics for sampling of feedforward neural networks trained with minibatches
A. ZambonPrimo
;G. Tiana
Ultimo
2026
Abstract
Sampling the parameter space of artificial neural networks according to a Boltzmann distribution provides insight into the geometry of low-loss solutions and offers an alternative to conventional loss minimization for training. However, exact sampling methods such as hybrid Monte Carlo (hMC), while formally correct, become computationally prohibitive for larger datasets because they require repeated evaluation of full–batch gradients. We introduce a pseudo–Langevin (pL) dynamics that enables efficient Boltzmann sampling of feed-forward neural networks trained with large datasets by using mini–batches in a controlled manner. The method exploits the statistical properties of mini–batch gradient noise and adjusts fictitious masses and friction coefficients to ensure that the induced stochastic process samples efficiently the desired equilibrium distribution. We validate numerically the approach by comparing its equilibrium statistics with those obtained from exact hMC sampling. Performance benchmarks demonstrate that, while hMC rapidly becomes inefficient as network size increases, the pL scheme maintains the best computational diffusion among exact and non–exact methods and scales favorably to networks with over one million parameters. Additionally, we show that sampling at intermediate temperatures yields optimal generalization performance, comparable to stochastic gradient descent, without requiring a validation set or early stopping procedure. Finally, we show that our sampling method can achieve optimal generalization on the MNIST dataset, independently of the architecture. These results establish controlled mini–batch Langevin dynamics as a practical and scalable tool for exploring and exploiting the solution space of large neural networks.| File | Dimensione | Formato | |
|---|---|---|---|
|
Zambon_2026_J._Stat._Mech._2026_073401.pdf
accesso aperto
Tipologia:
Publisher's version/PDF
Licenza:
Creative commons
Dimensione
2.41 MB
Formato
Adobe PDF
|
2.41 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




