Andrew Ng 推出新课程 Transformers in Practice,与 AMD 合作并由 Sharon Zhou 主讲,免费版可看所有视频和基础代码。课程用交互可视化演示 token 逐个生成、注意力头与层的分工以及幻觉成因,并讲解量化、KV Cache、Flash Attention、投机解码等推理优化技巧。
做LLM生产落地的开发老哥们,可以看Andrew Ng刚出的这门课,免费版可以看所有视频和基础代码。
这个课程不是又一遍Attention is All You Need的数学推导,
也不是又一套调prompt的玄学技巧,
更不是又一个从零写Transformer的玩具项目,它直接把LLM的黑箱给你拆开了。
会让你亲手玩自回归循环,
看着模型一个token一个token生成,看着某一步概率采样走偏,
看着幻觉是怎么一步步从无到有长出来的。
甚至会让你拖动滑块调整temperature,实时看到输出多样性的变化,
看到不同的采样策略到底在改变什么。
以及让你点开每一层每一个注意力头,
看到哪个头在管语法,
哪个头在管事实,
哪个头在管逻辑推理。
最狠的是推理优化部分,
这是所有生产工程师每天都在踩的坑,慢推理,OOM,成本爆炸。
以前所有人都告诉你要换更大的GPU。要加更多的机器。
这门课告诉你,
70%以上的延迟根本不是参数量的问题,是内存带宽的问题,是注意力计算的问题。
量化,KV Cache,Flash Attention,投机解码,
每一个技巧都能让你的模型速度翻2到5倍,精度损失几乎可以忽略。
而且这次是和AMD深度合作,由AMD工程副总裁亲自主讲。
终于有一门课不是只讲CUDA了,终于有人开始讲硬件感知的优化了。
虽然会调用API的人已经满大街都是了,但能看穿模型内部。能诊断问题。能优化成本的人,才是未来三年最稀缺的。
我觉得这门课最大的价值,是它终于把Transformer从一个学术概念,变成了一个你可以摸得到,可以调试,可以优化的工程工具。
Video
New course: Transformers in Practice. You'll get a practical view of how transformer-based LLMs work, so you can reason about their behavior, diagnose problems like slow inference, and make smarter decisions about deployment. This course is built in partnership with @AMD and taught by @realSharonZhou. You'll see how transformers generate text one token at a time, how the model decides which earlier words matter most when predicting the next one, and how techniques like quantization speed up inference on GPUs. This is not a video-only course; interactive visualizations throughout let you play with these concepts and build intuition that sticks. Skills you'll gain: - Understand why LLMs hallucinate, and RAG and chain-of-thought shape what they generate - Look inside the model to see how attention and layers combine to predict the next token - Diagnose inference bottlenecks and learn the techniques that speed up transformers on GPUs Join and understand what's really happening inside your LLMs: deeplearning.ai/courses/tran… Video在 X 查看被引用的帖子
来源:@AYi_AInotes · x.com