跳到正文

#开源/仓库

今日 1 条
今天10月2日周五
10月1日周四
9月29日周二
  1. Thomas Wolf56

    modded-nanogpt 传入新的历史纪录 39.9 秒,较此前 67.6 秒快 27.7 秒,核心思路是在单个 flop 级别做稀疏优化而非只优化矩阵乘法。主要手段包括采样 softmax(约 8 秒)、稀疏 n-gram 嵌入更新与优化器状态、稀疏通信、最后 300 步 EMA(约 4 秒)、新优化器 Anvil2(约 1 秒)等,稀疏嵌入参数扩展到 65B,占本次提升的 25%。详见 https://github.com/KellerJordan/modded-nanogpt/pull/360 和 https://hyperstition.cc/training-nanogpt-in-39-9-seconds。

    引用Larry Dial@classiclarryd

    New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds

9月28日周一
9月27日周日
9月24日周四
9月23日周三
  1. ViggleAI46

    我们为开源社区推出了首个 Qwen-Image-2.1 turbo。 试试 Viggle-Turbo,一个经 DMD 蒸馏的 Qwen-Image-2.1,仅需 4 个采样步即可生成和编辑,无需 classifier-free guidance。 权重:https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

    引用Hugging Apps@HuggingApps

    Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo

9月14日周一
9月11日周五
8月27日周四
6月15日周一
12月23日周二