Wasserstein gradient flows for energy-kernel discrepancies

Research announcement
Continuum and particle dynamics for energy-kernel MMDs, including a two-timescale transport-and-relaxation mechanism.
Author

Matthew Rosenzweig

Published

August 6, 2026

Modified

August 6, 2026

Dejan Slepčev, Lihan Wang, and I have posted our paper Wasserstein gradient flows of Maximum Mean Discrepancy with energy kernels. It studies the continuum and particle dynamics generated by a family of nonsmooth discrepancies that includes the classical energy distance. The paper is a companion to my recent work with Antonin Chodron de Courcel on Wasserstein gradient flows for Coulomb discrepancies, but the energy kernels present a rather different combination of analytical and dynamical challenges.

Energy kernels as transport energies

Fix a target probability distribution \(\mu\) on \(\mathbb R^d\). For \(0<q<2\), we consider the kernel

\[ K(z)=-|z|^q. \]

When the relevant moments are finite, the resulting squared Maximum Mean Discrepancy (MMD) is

\[ \mathsf{MMD}_q^2(\rho,\mu) =-\frac12\iint_{\mathbb R^d\times\mathbb R^d}|x-y|^q \,d(\rho-\mu)(x)\,d(\rho-\mu)(y). \]

For \(q=1\), this is the energy distance used in statistics [1, 2]. More generally, it is, up to a constant, the squared homogeneous Sobolev norm \(\|\rho-\mu\|_{\dot H^{-(d+q)/2}}^2\). We fix \(\mu\) and evolve \(\rho\) by the Wasserstein gradient flow of this discrepancy. Formally, \(\rho\) is transported by the nonlocal velocity field generated by \(\rho-\mu\).

Energy-kernel MMDs combine sensitivity to large-scale separation with dimension-free empirical approximation: the MMD between a distribution and the empirical measure of \(N\) independent samples is typically of order \(N^{-1/2}\), independently of the ambient dimension. This also connects the problem to deterministic quantization. Fornasier and Hütter established qualitative consistency of equal-weight power-kernel quantizers, while Colasanto, Focardi, Fornasier, and Mattesini recently obtained sharp empirical-quantization rates for power kernels [3, 4]. In joint work with Hess-Childs and Serfaty, we have also obtained sharp empirical-quantization rates for a broader class of homogeneous and screened Riesz MMDs and for their singular diagonal-excluded counterparts, known as modulated energies [5].

From continuum flow to particle dynamics

The nonsmoothness that makes these kernels useful also puts the dynamics outside standard theory. The interaction force is not Lipschitz at the particle diagonal, and in dimensions \(d\ge2\) the energy is not displacement semiconvex in Wasserstein space. The continuum equation and the particle system therefore have to be constructed separately.

In the range \(d+q-2>0\), we prove global existence and uniqueness for probability densities in subcritical \(L^p\) spaces, with finite moments; we also include the endpoint \(d=q=1\), whose special structure makes it explicitly solvable. For the diagonal-free particle system, we prove global noncollision and convergence, at fixed \(N\), to the collision-free critical set. We construct explicit collision-free saddle equilibria, implying the particles do not in general evolve to an optimal empirical quantizer of the target.

We then connect the two descriptions through a modulated-energy estimate, which is a type of weak-strong stability estimate widely used in the equilibrium and nonequilibrium analysis of Coulomb/Riesz gases (see, e.g., [6]). It gives a quantitative mean-field limit on every fixed time interval, and for well-prepared random initial data it propagates the dimension-free \(N^{-1/2}\) MMD scale from the initial empirical approximation to the evolving particle system.

The long-time picture is subtler. In most of the parameter range, an absolutely continuous Lagrangian critical point must equal the target. The exceptions at the natural moment level are confined to \(0<q<1\) in dimensions one and three: in dimension three we recover rigidity under several additional hypotheses, while in dimension one, outside our present continuum well-posedness range, we construct compactly supported non-minimizing critical points. Energy dissipation gives convergence toward the critical set without additional long-time bounds for \(1\le q<2\), and with uniform moment bounds and subcritical \(L^p\) bounds for \(0<q<1\). In the rigid regimes, those bounds also give convergence to the target.

There is overlap here with recent work of Fornasier, Huang, and Sun on a broader attraction–repulsion equation [7]. When their attractive and repulsive exponents agree, their model contains, as a special case, the energy-kernel MMD flow for \(1\le q<2\); their main emphasis is the unequal-exponent regime, and their results provide a complementary Lagrangian well-posedness and stationary-state theory.

Transport first, relaxation second

One of the main messages of the paper is that convergence has two distinct scales. If a source is initially placed a distance \(R\) from a compactly supported target, then the far-field velocity is of order \(R^{q-1}\). The source therefore needs a time of order \(R^{2-q}\) merely to reach the target region. On the whole space, this rules out both a global Polyak–Łojasiewicz inequality and any multiplicative MMD decay rate that is uniform over all initial data. It does not rule out rapid relaxation after a waiting time determined by the initial geometry.

The distinction becomes completely explicit for the one-dimensional energy distance. The ordered particle equations decouple through the target cumulative distribution function. A particle outside the target support first moves toward it at constant velocity. Once it enters, it converges exponentially to its equilibrium quantile provided the target density is bounded below on its support. Thus a finite entrance time separates a transport phase from a relaxation phase. The lower bound on the target density is essential. For example, if a target density vanishes quadratically at its midpoint, then a displaced middle particle relaxes only like \(t^{-1/2}\), and its excess energy decays like \(t^{-2}\). The continuum flow displays the same mechanism quantile by quantile, although for a source with unbounded support there need not be a finite time by which every quantile has entered the target region.

Earlier work of Di Francesco, Fornasier, Hütter, and Matthes analyzed the same balanced one-dimensional power flow through pseudo-inverse distribution functions and proved well-posedness and long-time convergence in several regimes [8]. The distance-kernel flow was also recently studied through quantiles by Duong, Stein, Beinert, Hertrich, and Steidl [9]. Our example isolates the transport and post-entry relaxation scales and shows explicitly how degeneracy of the target changes the second scale. We expect this two-timescale mechanism to be much more general than the example in which we can presently prove it.

What remains open

Three questions come to mind immediately.

First, can the moment bounds and subcritical \(L^p\) bounds needed to identify long-time limits be propagated uniformly in time? In the rigid regimes, a positive answer would turn our conditional convergence theory into unconditional convergence to the target.

Second, does rigidity at the natural moment level also hold in the remaining regime \(d=3\), \(0<q<1\), or can there be an absolutely continuous non-minimizing critical point there?

Third, can the two-timescale picture be established for the full family of energy-kernel flows in arbitrary dimension? This would require a theory that separates the datum-dependent time needed to transport mass into the relevant target region from the subsequent local relaxation and that explains how the latter depends on target positivity and degeneracy. On the torus, Chizat, Colombo, Colombo, and Fernández-Real prove polynomial convergence near sufficiently regular positive targets in the corresponding non-Coulomb Sobolev regime [10]. Their local-in-data theory complements our global obstruction: the failure of a global PL inequality does not preclude local decay, but it explains why such a local argument does not yield a global exponential theory. Understanding the analogous local theory on the whole space and its interaction with particle approximation and the long-time mean-field limit is part of the same problem.

The paper is available on arXiv; the PDF can be downloaded here.

I thank Massimo Fornasier for bringing his related work on this topic to our attention.

References

  1. G. J. Székely and M. L. Rizzo, “A new test for multivariate normality,” Journal of Multivariate Analysis 93 (2005), no. 1, 58–80. doi:10.1016/j.jmva.2003.12.002.

  2. G. J. Székely and M. L. Rizzo, “Energy statistics: A class of statistics based on distances,” Journal of Statistical Planning and Inference 143 (2013), no. 8, 1249–1272. doi:10.1016/j.jspi.2013.03.018.

  3. M. Fornasier and J.-C. Hütter, “Consistency of probability measure quantization by means of power repulsion–attraction potentials,” Journal of Fourier Analysis and Applications 22 (2016), no. 3, 694–749. doi:10.1007/s00041-015-9432-z.

  4. F. Colasanto, M. Focardi, M. Fornasier, and F. Mattesini, “Sharp rates of MMD empirical estimation with power kernels,” arXiv:2605.18497 (2026). arXiv.

  5. E. Hess-Childs, M. Rosenzweig, and S. Serfaty, “Optimal quantization for Riesz MMDs,” manuscript in preparation (2026).

  6. S. Serfaty, Lectures on Coulomb and Riesz Gases, Colloquium Publications 70, American Mathematical Society, Providence, RI, 2026. doi:10.1090/coll/070.

  7. M. Fornasier, H. Huang, and L. Sun, “The nonlocal attraction-repulsion transport equation with power kernels,” arXiv:2607.04424 (2026). arXiv.

  8. M. Di Francesco, M. Fornasier, J.-C. Hütter, and D. Matthes, “Asymptotic behavior of gradient flows driven by nonlocal power repulsion and attraction potentials in one dimension,” SIAM Journal on Mathematical Analysis 46 (2014), no. 6, 3814–3837. doi:10.1137/140951497.

  9. R. Duong, V. Stein, R. Beinert, J. Hertrich, and G. Steidl, “Wasserstein gradient flows of MMD functionals with distance kernel and Cauchy problems on quantile functions,” ESAIM: Control, Optimisation and Calculus of Variations 32 (2026), Paper No. 10. doi:10.1051/cocv/2025097.

  10. L. Chizat, M. Colombo, R. Colombo, and X. Fernández-Real, “Quantitative convergence of Wasserstein gradient flows of kernel mean discrepancies,” arXiv:2603.01977 (2026). arXiv.

Back to top

Citation

BibTeX citation:
@online{rosenzweig2026,
  author = {Rosenzweig, Matthew},
  title = {Wasserstein Gradient Flows for Energy-Kernel Discrepancies},
  date = {2026-08-06},
  url = {https://matthewrosenzweigwork-max.github.io/posts/wasserstein-gradient-flows-energy-kernel-discrepancies/},
  langid = {en}
}
For attribution, please cite this work as:
Rosenzweig, Matthew. 2026. “Wasserstein Gradient Flows for Energy-Kernel Discrepancies.” August 6. https://matthewrosenzweigwork-max.github.io/posts/wasserstein-gradient-flows-energy-kernel-discrepancies/.