跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 24 天前AI 评分24
AI 导读

注意力机制本身没有时间概念。位置嵌入是语言模型判断 token 之间距离的方式。 Jane Street 和 3Blue1Brown 最近的一期 YouTube 视频,探讨了满足常见数学约束的嵌入空间。这个空间不仅在代数上有良好定义,而且在 ML 文献中已被充分挖掘。 如果我们放宽这些约束呢?是否存在一个更一般的可能函数空间?让我们来折腾一下(1/7)🧵

正文

Attention has no innate notion of time. Positional embeddings are how language models tell how far apart tokens are.

A recent YouTube video by Jane Street and 3Blue1Brown explores the space of embeddings that satisfy common-sense mathematical restrictions. This space is not only well-defined algebraically, but thoroughly exploited in the ML literature.

What if we relax those constraints? Is there an even more general space of possible functions? Let's mess around (1/7)🧵

来源:@SemiAnalysis_ · x.com