聊聊安全话题:GLM5.3 无审查版已上线 HuggingFace。好戏在后头。
here u go https://huggingface.co/Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw
聊聊安全话题:GLM5.3 无审查版已上线 HuggingFace。好戏在后头。
here u go https://huggingface.co/Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw
Ai2 发布 Olmo-core 3,为 Olmo 框架带来重新设计的开源 MoE 训练系统,可扩展到万亿参数规模并保持计算效率。新实现从 FSDP 切换到基于 DDP 的方案,47B 参数 MoE 在 8 张 NVIDIA B300 上达到每 GPU 每秒 52,000 tokens,约为旧实现的 2.7 倍;启用 MXFP8 后吞吐比 BF16 高约 21%。
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds
开源 RL 环境就是赢。当然是在 Hugging Face 上!
was looking for a quiet weekend but xiaomi dropped their rl envs repo last night to put in perspective, if you have to buy some tasks like this its usually hundred to thousand dollars per task so this repo is literally worth millions https://huggingface.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
蚂蚁集团联合多家金融机构与领域专家推出 Ling-3.0-flash-Fin,这是 Ling 系列首个金融增强模型,在 Ling-3.0-flash 基础上用高质量金融数据继续训练,总参数 124B、激活参数 5.1B、上下文窗口 256K。
Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo
inclusionAI 推出 SingProbe,一个基于 meta-llama/Llama-3.2-1B-Instruct 的流式安全探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,仅增加不到 0.5% 解码开销。
inclusionAI 发布 Llama-3.1-8B-Instruct-singprobe,一个基于 meta-llama/Llama-3.1-8B-Instruct 的流式安全探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI 发布 gpt-oss-20b-singprobe,这是构建在 openai/gpt-oss-20b 上的流式护栏探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
Ling-3.0-flash-VL 两项更新: - 现已在 OpenRouter 上线,提供两周免费访问 - FP4 和 INT4 量化版本现已开源 可通过 API 试用或自行部署——由你选择。


更多好消息:GLM-5.3 的权重将于明天发布。 https://huggingface.co/zai-org/GLM-5.3
Z Lab、SGLang 与 Modal 联合发布面向 Qwen 3.5 397B-A17B 的 DFlash 投机解码草拟模型,并随 SGLang 新的 Spec V2 引擎默认可用。
SpecForge 团队联合蚂蚁集团、美团、Nex-AGI 和 EigenAI 发布 SpecBundle(Phase 1),提供覆盖 8B 到 1T 参数主流 instruct 模型的生产级 EAGLE3 草稿模型权重,最高可达 4 倍端到端推理加速。