skip to content

Department of Applied Mathematics and Theoretical Physics

Local SGD, also known as Federated Averaging, is a simple approach to communication-efficient distributed optimization in which each client performs several stochastic-gradient steps between communication rounds. Despite its empirical success [1], a basic theoretical question remains: when do sequential local updates outperform the use of the same stochastic gradients in a large mini-batch? In this talk, I will summarize a line of work that develops a hierarchy of assumptions on data heterogeneity and clarifies when local updates are worthwhile [1, 2, 3, 4, 5, 6, 7].

We begin with the homogeneous setting, reviewing the optimal rates and thereby establishing the benchmark that any heterogeneous guarantee could hope to match [2, 3]. We then study a natural first-order condition capturing approximate simultaneous realizability and show that it is too weak to guarantee an advantage for Local SGD [4]. At the opposite extreme, uniform gradient-similarity assumptions can recover the improvements seen in the homogeneous setting [5], but at the cost of excluding meaningful heterogeneity across clients [4]. These observations, together with analyses of related local-update methods in nonconvex optimization [6], motivate bounded second-order heterogeneity, which controls differences in local objective geometry rather than requiring local gradients to remain uniformly close. I will review the resulting theory in the strongly convex setting [7], including convergence rates and the roles of heterogeneity, stochastic noise, and higher-order smoothness, before presenting recent guarantees for general convex objectives and nearly matching lower bounds. A central idea in these analyses is a trajectory-dependent treatment of client drift: consensus error is controlled using only the gradient disagreement encountered along the optimization path, and the resulting bounds are closed through self-bounding recursions. The theory identifies regimes in which Local SGD provably outperforms mini-batch SGD and suggests a simple principle: effective local updates require aligned local geometry rather than globally similar gradients. I will conclude by discussing how these insights and techniques inform our understanding of data heterogeneity in other distributed optimization problems.

Further information

Time:

13Oct
Oct 13th 2026
14:00 to 15:00

Venue:

MR19 (Potter Room, Pavilion B), CMS

Series:

Machine learning theory