AI 导读
SemiAnalysis 提出进一步放宽注意力机制假设,不再要求可分离作用,得到 q(s)T A(t-s) k(t),其中 A 是输出方阵的平移不变函数,输出空间不再需要群结构。测试用高斯核 A(t-s) = exp(-(t-s)^2 / σ^2) I,嵌入函数等价于单调递减标量,效果逊于 RoPE 但仍优于对照组。
正文
What if we relax these assumptions even further? Let's stop requiring separable action. This leaves q(s)T A(t-s) k(t), where A is an arbitrary translation-invariant function that outputs a square matrix. This means we no longer require any group structure in the output space.
To test, we'll simply take a Gaussian: A(t-s) = \exp(-(t-s)^2 / \sigma^2) I. Our embedding function becomes equivalent to a monotonically decreasing scalar. A quick check shows that this function performs worse than RoPE but still beats the controls. (4/7)
来源:@SemiAnalysis_ · x.com