AI 导读
作者发布深度博客《Inside the Transformer: The Life of a Token》,沿单个 token 的流动路径拆解现代稠密 Transformer 的内部设计。
正文
new in-depth blog post time: Inside the Transformer: The Life of a Token
a deep dive into a modern dense transformer, i cover YaRN (why does pairwise coordinate rotation induce positional information?), hybrid attention (getting to 160k context length), soft capping, QK normalization, etc. as the token flows through the transformer
bonus transformer math: FLOPs/token formula (and when is 6N formula broken), cluster sizing (how big of a cluster do you need given the model/data size and experiment throughput of interest), and more
来源:@gordic_aleksa · x.com