跳到正文
原文
NVIDIA AI· @NVIDIAAI · X·· 2 天前AI 评分50
AI 导读

CoreWeave 借助 NVIDIA Dynamo 的 ModelExpress 和 Router,在对 Nemotron 3.5 Lightning 做 RL 后训练时实现模型重载速度比其基线快 15 倍,且停机时间极短。RL 后训练需反复训练、生成、再训练,推理 worker 每次都要重新加载更新后的模型权重,模型越大 GPU 空等越严重。

正文

Congrats @CoreWeave on RL Rollouts!

RL post-training involves a lot of back and forth: train the model, generate responses, then train again. Inference workers need to load the updated model weights each time. As models get bigger, that can leave GPUs waiting.

CoreWeave’s new service uses ModelExpress and Router in NVIDIA Dynamo to speed up those reloads with minimal downtime.

Working with us and @youdotcom, CoreWeave achieved 15× faster model reloads compared with its baseline while post-training Nemotron 3.5 Lightning.

Check out their blog below for details

引用CoreWeave@CoreWeave
ICYMI: CoreWeave Forge is here 🎉 A production trace that never reaches the next training run is a signal you paid for and threw away. Most teams do it every day, because the tool that catches the trace and the tool that runs the training came from different vendors and were never built to talk. Forge closes that gap by unifying @wandb, post-training from @OpenPipeAI, and @marimo_io notebooks in one connected environment with CoreWeave Training, Inference, Sandboxes, and Registry. Run, observe, curate, improve, evaluate. The traces you flag in production become the datasets you train on and the evaluations you gate with. Open across any model, framework, or cloud. @MasterClass and @canva are already building on it. Free, Pro, and Enterprise available today: https://crwv.co/utcAw
在 X 查看被引用的帖子

来源:NVIDIA AI · x.com