Sampling the parameter space of artificial neural networks according to a Boltzmann distribution provides insight into the geometry of low-loss solutions and offers an alternative to conventional loss minimization for training. However, exact sampling methods such as hybrid Monte Carlo (hMC), while formally correct, become computationally prohibitive for larger datasets because they require repeated evaluation of full–batch gradients. We introduce a pseudo–Langevin (pL) dynamics that enables efficient Boltzmann sampling of feed-forward neural networks trained with large datasets by using mini–batches in a controlled manner. The method exploits the statistical properties of mini–batch gradient noise and adjusts fictitious masses and friction coefficients to ensure that the induced stochastic process samples efficiently the desired equilibrium distribution. We validate numerically the approach by comparing its equilibrium statistics with those obtained from exact hMC sampling. Performance benchmarks demonstrate that, while hMC rapidly becomes inefficient as network size increases, the pL scheme maintains the best computational diffusion among exact and non–exact methods and scales favorably to networks with over one million parameters. Additionally, we show that sampling at intermediate temperatures yields optimal generalization performance, comparable to stochastic gradient descent, without requiring a validation set or early stopping procedure. Finally, we show that our sampling method can achieve optimal generalization on the MNIST dataset, independently of the architecture. These results establish controlled mini–batch Langevin dynamics as a practical and scalable tool for exploring and exploiting the solution space of large neural networks.

Controlled Langevin dynamics for sampling of feedforward neural networks trained with minibatches / A. Zambon, F.C.. - In: JOURNAL OF STATISTICAL MECHANICS: THEORY AND EXPERIMENT. - ISSN 1742-5468. - 2026:7(2026 Jul 09), pp. 073401.1-073401.26. [10.1088/1742-5468/ae8246]

Controlled Langevin dynamics for sampling of feedforward neural networks trained with minibatches

A. Zambon
Primo
;
G. Tiana
Ultimo
2026

Abstract

Sampling the parameter space of artificial neural networks according to a Boltzmann distribution provides insight into the geometry of low-loss solutions and offers an alternative to conventional loss minimization for training. However, exact sampling methods such as hybrid Monte Carlo (hMC), while formally correct, become computationally prohibitive for larger datasets because they require repeated evaluation of full–batch gradients. We introduce a pseudo–Langevin (pL) dynamics that enables efficient Boltzmann sampling of feed-forward neural networks trained with large datasets by using mini–batches in a controlled manner. The method exploits the statistical properties of mini–batch gradient noise and adjusts fictitious masses and friction coefficients to ensure that the induced stochastic process samples efficiently the desired equilibrium distribution. We validate numerically the approach by comparing its equilibrium statistics with those obtained from exact hMC sampling. Performance benchmarks demonstrate that, while hMC rapidly becomes inefficient as network size increases, the pL scheme maintains the best computational diffusion among exact and non–exact methods and scales favorably to networks with over one million parameters. Additionally, we show that sampling at intermediate temperatures yields optimal generalization performance, comparable to stochastic gradient descent, without requiring a validation set or early stopping procedure. Finally, we show that our sampling method can achieve optimal generalization on the MNIST dataset, independently of the architecture. These results establish controlled mini–batch Langevin dynamics as a practical and scalable tool for exploring and exploiting the solution space of large neural networks.
artificial neural networks; Boltzmann measure Contents;
Settore PHYS-04/A - Fisica teorica della materia, modelli, metodi matematici e applicazioni
9-lug-2026
https://iopscience.iop.org/article/10.1088/1742-5468/ae8246
Article (author)
File in questo prodotto:
File Dimensione Formato  
Zambon_2026_J._Stat._Mech._2026_073401.pdf

accesso aperto

Tipologia: Publisher's version/PDF
Licenza: Creative commons
Dimensione 2.41 MB
Formato Adobe PDF
2.41 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1260203
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact