White Circle 发布开源分布式后训练框架 Halo,面向已超出 Huggingface TRL 但不需要完整 Megatron 技术栈的模型。在 8× B300 上,Halo 训练吞吐最高约为原生 Huggingface Transformers RL 的 2.8 倍,双方都启用 ZeRO-3 分片时为 2.7 倍且峰值内存少 25%。
White Circle just unveiled Halo, an open-source, distributed post-training framework for models that have outgrown Huggingface TRL but don't need the full Megatron stack.
On 8× B300, Halo delivers up to ~2.8× the training throughput of stock Huggingface Transformers Reinforcement Learning (2.7× at 25% less peak memory when both sides shard ZeRO-3), with larger margins over the other frameworks benchmarked.
you can run multi-turn RL directly on Hugging Face models, from one GPU to a cluster.
Rather than reimplementing model architectures, Halo wraps existing Transformers models with expert, context, tensor and expert-tensor parallelism.
来源:@rohanpaul_ai · x.com