# POV: Chinese AI Lab Teaching Everyone How To Save Millions of Dollars

Source: https://www.youtube.com/watch?v=cJeqGq0Bx1M
Recap page: https://rapidrecap.app/video/cJeqGq0Bx1M
Generated: 2025-07-19T19:02:56.336+00:00

---
## Quick Overview

ByteDance Seed, the AI lab behind TikTok, is rapidly emerging as a top Chinese AI research entity, outperforming competitors like DeepSeek and Google's Veo 3 in video generation. They recently published a groundbreaking paper on 'Model Merging in Pre-training of Large Language Models,' introducing Pre-trained Model Averaging (PMA), a novel technique that significantly reduces pre-training costs by predicting final model performance early, offering free accuracy gains of 3-7%, and stabilizing training, potentially saving millions in compute resources.

**Key Points:**
- ByteDance Seed's Seedance 1.0 video model ranks first on the Artificial Analysis Video Arena Leaderboard, outperforming Google's Veo 3.
- ByteDance plans to invest $12 billion in AI chips by 2025, demonstrating a massive commitment to AI development.
- Their new paper introduces Pre-trained Model Averaging (PMA), a novel technique for merging models during the pre-training phase.
- PMA allows predicting a model's final performance early in training, saving 3-6 days of compute and approximately 15% of the budget.
- The technique provides a 'free' accuracy gain of 3-7% and enhances training stability, aiding in crash recovery and noisy setups.
- The optimal interval for saving checkpoints in PMA scales with model size, aligning with tendencies for larger models to use larger batch sizes.
- Runpod, a sponsor, offers serverless GPU services and a Hub for easy deployment of AI models, including a revenue share program for creators.

![Screenshot at 0:23: Artificial Analysis Video Arena Leaderboard showing ByteDance Seed's Seedance 1.0 model ranked first.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-00-23.png)

**Context:** The video discusses the advancements of ByteDance Seed, the AI research arm of ByteDance (owner of TikTok), in the competitive landscape of artificial intelligence. It highlights their significant financial investment and research breakthroughs, particularly in large language model (LLM) pre-training. The core focus is on a novel technique called Pre-trained Model Averaging (PMA) and its implications for reducing computational costs and improving training efficiency, a critical challenge in large-scale AI development.

## Detailed Analysis

ByteDance Seed, the AI lab associated with TikTok and Douyin, is quickly establishing itself as a leading force in Chinese AI research, even surpassing DeepSeek in some metrics and Google's Veo 3 in video generation with their Seedance 1.0 model. Their substantial budget, including a planned $12 billion investment in AI chips for 2025, positions them to compete directly with global leaders like Google and OpenAI. A recent pivotal paper from ByteDance Seed introduces Pre-trained Model Averaging (PMA), a novel strategy for model-level weight merging during the pre-training phase of large language models (LLMs). This technique involves saving checkpoints during the constant learning rate phase of training and averaging them to create a merged model. This merged model surprisingly achieves performance comparable to a fully annealed model much earlier in the training cycle, effectively saving 3-6 days of training time and approximately 15% of compute resources. The paper demonstrates PMA's effectiveness across various dense models (from 411M to 70B parameters) and Mixture-of-Experts (MoE) architectures (from 0.7B/7B to 20B/200B parameters), an undertaking estimated to cost around $15 million in GPU time. PMA not only provides early performance estimates and free accuracy gains (3-7%) but also stabilizes training, making it resilient to issues like loss spikes and beneficial for noisy setups or distributed computing. The research highlights that simple moving average (SMA) is the most effective merging method, as it acts like a one-shot low-pass filter, effectively removing high-frequency noise from the model's weight trajectory without the need for iterative annealing. This open-source contribution from ByteDance Seed is poised to significantly impact how top AI labs approach and optimize expensive pre-training processes.

### ByteDance Seed's AI Prowess

- ByteDance Seed is positioned as a top Chinese AI lab, outperforming Baidu and DeepSeek in AI research
- Their Seedance 1.0 video model ranks first on the Artificial Analysis Video Arena Leaderboard, surpassing Google's Veo 3
- ByteDance plans to invest $12 billion in AI chips by 2025, indicating significant resources for AI development.

### The Challenge of Pre-training Model Merging

- Model merging in the post-training stage is common, but pre-training merging remains largely unexplored due to its immense cost and risk
- Training a 70B model can cost $2 million per run, making extensive experimentation prohibitively expensive for most labs
- Publishing successful pre-training merging techniques is rare, as it provides free competitive intelligence to rivals.

### Introducing Pre-trained Model Averaging (PMA)

- ByteDance Seed published a paper detailing PMA, a novel strategy for model-level weight merging during pre-training
- PMA involves saving multiple checkpoints during the constant learning rate phase and averaging their weights into a single model
- This merged model accurately predicts the final performance of a fully annealed model much earlier, saving 3-6 days of training and ~15% compute.

### PMA's Performance and Cost Savings

- PMA was demonstrated on various dense models (411M to 70B parameters) and Mixture-of-Experts (MoE) architectures (0.7B/7B to 20B/200B parameters)
- The extensive experiments conducted for this research are estimated to have cost up to $15 million in GPU time
- PMA provides a free accuracy gain of 3-7% and stabilizes training, making it useful for crash recovery and noisy setups.

### Mechanism and Optimal Implementation

- PMA works by acting as a low-pass filter, effectively averaging out high-frequency noise in the model's weight trajectory, similar to annealing but in a single step
- Simple Moving Average (SMA) proved to be the most effective merging method, outperforming more complex exponential and weighted moving averages
- Optimal checkpoint saving intervals scale with model size, for example, every 4 billion tokens for 0.7B/7B models and every 80 billion tokens for 10B/100B models.

### Runpod Serverless and Hub

- Runpod offers serverless GPU services and a Hub for one-click deployment of open-source AI repositories
- This platform simplifies model training and deployment by abstracting server management and automatically scaling resources
- Runpod is launching a revenue share program for Hub creators, allowing them to earn money when their repositories are deployed, supporting open-source AI development.

![Screenshot at 0:07: Bar chart comparing AI lab performance, showing ByteDance as the top Chinese AI lab.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-00-07.png)
![Screenshot at 0:15: Reuters headline: 'TikTok owner ByteDance plans to spend $12 billion on AI chips in 2025, FT reports'.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-00-15.png)
![Screenshot at 0:23: Artificial Analysis Video Arena Leaderboard showing ByteDance Seed's Seedance 1.0 model at the top for text-to-video generation.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-00-23.png)
![Screenshot at 0:47: Diagram illustrating the concept of model merging, where multiple models combine into a final model for various tasks.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-00-47.png)
![Screenshot at 1:04: Screenshot of multiple research papers on model merging, highlighting the scarcity of pre-training merging research.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-01-04.png)
![Screenshot at 1:30: Table showing GPU rental prices per hour, emphasizing the high cost of AI model training.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-01-30.png)
![Screenshot at 2:14: Title slide of the ByteDance Seed paper: 'Model Merging in Pre-training of Large Language Models'.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-02-14.png)
![Screenshot at 2:22: Charts demonstrating the performance impact of different model merging hyperparameters and the ability to predict annealing behavior.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-02-22.png)
![Screenshot at 2:42: Runpod serverless interface showing various GPU options and their hourly prices for AI model deployment.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-02-42.png)
![Screenshot at 4:01: Graph illustrating a typical pre-training cycle with warm-up, constant, and annealing phases for learning rate.](https://ss.rapidrecap.app/screens/cJeqGq0Bx1M/00-04-01.png)
