AI 导读
Cohere CEO Aidan Gomez 称,消费级 AI 工具产生的用户数据经合成改写后会被用于训练,这一说法来自多家大厂员工的传闻。他表示复杂的数学、商业、软件、生物等用例更可能被筛选和加权,即便在 ZDR 和承诺不训练用户的机制下,衍生数据通常也被排除在承诺之外。
正文
Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees.
In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn.
Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game.
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).在 X 查看被引用的帖子
来源:@aidangomez · x.com