Home/Events/Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

Confirmed
Confidence
90%
Impact: 80%
Updated 44m ago

Consensus Brief

The AsyncGRPOTrainer now supports training a LoRA adapter and synchronizing only that adapter to vLLM, significantly reducing data transfer requirements. This new approach allows for separate training and inference processes across different machines, utilizing Storage Buckets for efficient data handling.

Sourced from
Primary: Hugging Face

What Changed Since Last Update

44m ago

New official source added: Hugging Face published an update on Thu, 10 Se ("Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL").

Claim Ledger

4 claims tracked across sources

Confirmed Fact

AsyncGRPOTrainer can now train a LoRA adapter and sync only that adapter to vLLM.

Confirmed Fact

A rank-1 adapter for a 1.5B model is a few megabytes, while the full model is around 3 GB.

Confirmed Fact

The AsyncGRPO metrics show where the bottleneck sits, with five runs taking the same recipe from 3 h 27 min to 53 min for 500 steps.

Official Claim

LoRA training is particularly suited for RL, as shown in Thinking Machines's blog LoRA Without Regret.

Role-Based Impact Analysis

Source Timeline

30 sources corroborating