[2502.11034] AdaGC: Improving Training Stability for Large Language Model Pretraining

[2502.11034] AdaGC: Improving Training Stability for Large Language Model Pretraining

arXiv - Machine Learning 4 min read Article

Summary

The paper presents AdaGC, a novel adaptive gradient clipping method aimed at enhancing training stability in large language model pretraining by addressing loss spikes caused by various factors.

Why It Matters

Training stability is crucial for the performance of large language models. Loss spikes can hinder model accuracy and efficiency. AdaGC provides a solution that is optimizer-agnostic and reduces communication costs, making it a significant advancement for researchers and practitioners in machine learning.

Key Takeaways

  • AdaGC mitigates loss spikes during large language model training.
  • The method is optimizer-agnostic and integrates seamlessly with existing optimizers.
  • Empirical results show significant improvements in accuracy and stability over previous methods.
  • AdaGC reduces communication costs in distributed training environments.
  • The approach addresses the confluence of multiple factors causing training instability.

Computer Science > Machine Learning arXiv:2502.11034 (cs) [Submitted on 16 Feb 2025 (v1), last revised 22 Feb 2026 (this version, v2)] Title:AdaGC: Improving Training Stability for Large Language Model Pretraining Authors:Guoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng, Jiabin Yang, Dianhai Yu, Yanjun Ma, Li Shen View a PDF of the paper titled AdaGC: Improving Training Stability for Large Language Model Pretraining, by Guoxia Wang and 7 other authors View PDF HTML (experimental) Abstract:Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause of loss spikes by investigating individual factors, we observe that, in practice, such spikes are typically triggered by the confluence of heterogeneous factors. Empirically, loss spikes may arise from a combination of data outliers, hardware or transient computational faults, numerical precision issues, and hyperparameter settings. Regardless of the underlying cause, these spikes manifest as unstable optimizer updates, as abnormal gradients contaminate both first- and second-moment states. In this paper, we propose a principled gradient-centric remedy: AdaGC, an adaptive per-tensor gradient clipping scheme that mitigates such contamination by bounding gradient norms relative to a tensor-wise exponential moving average of their historical clipped values. AdaGC is optimizer-agnostic, introduces negligible memory overhead, and reduces communic...

Related Articles

Llms

I think we’re about to have a new kind of “SEO”… and nobody is talking about it.

More people are asking ChatGPT things like: “what’s the best CRM?” “is this tool worth it?” “alternatives to X” And they just… trust the ...

Reddit - Artificial Intelligence · 1 min ·
Llms

Why would Claude give me the same response over and over and give others different replies?

I asked Claude to "generate me a random word" so I could do some word play. Then I asked it again in a new prompt window on desktop after...

Reddit - Artificial Intelligence · 1 min ·
Anthropic blocks OpenClaw from Claude subscriptions
Llms

Anthropic blocks OpenClaw from Claude subscriptions

Anthropic forces pay-as-you-go pricing for OpenClaw users after creator joins OpenAI

AI Tools & Products · 6 min ·
Llms

wtf bro did what? arc 3 2026

The Physarum Explorer is a high-speed, bio-inspired neural model designed specifically for ARC geometry. Here is the snapshot of its curr...

Reddit - Artificial Intelligence · 1 min ·
More in Llms: This Week Guide Trending

No comments

No comments yet. Be the first to comment!

Stay updated with AI News

Get the latest news, tools, and insights delivered to your inbox.

Daily or weekly digest • Unsubscribe anytime