ai·radar
Bản tin sáng Tra cứu

Looping Beyond Twice: Một công thức có thể mở rộng cho các chuyên gia

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

Đoạn trích bài viết

Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons.

Toàn văn bài viết

Đọc bài viết đầy đủ trên huggingface

Mở bài gốc để xem trọn vẹn chi tiết và dẫn chứng.

Mở bài viết gốc