Elvis Saravia 就 HuggingFace 与 OpenAI 事件一文提出三点:目前细节仍不完整,无法对该事件下结论;文章中的 AI 拟人化叙事值得警惕,可能引发无根据的恐慌并导致糟糕决策。他提醒 AI 工程师为即将到来的持久化模型与智能体做准备,重点关注 reward hacking、评测与沙箱技术,并认为通用前沿模型未必适用于所有场景,多数任务用更受限的定制模型可能更安全。
As I've been saying for months, we are truly not ready for persistent agents.
3 comments I want to make about this article:
1) We don't have the full picture
It's a great summary of what happened in the HuggingFace <> OpenAI incident. You have to read it, but do understand we are still missing lots of important details to make any meaningful conclusions about this event.
2) The danger of AI anthropomorphism
The anthropomorphism in this article is next level. I hope this doesn't become the new norm for writing about future AI capabilities. I prefer technical writeups with widely accepted terminology, etc. I admire Dwarkesh's desire to share AI trends, but we can all do better in how we communicate about AI. Everyone is paying attention, and so we have a responsibility to avoid AI anthropomorphism, which unfortunately has led to unfounded fear-based mongering, and as a result, terrible decision-making for our industry.
3) Prepare for persistent agents/models
For AI engineers, prepare for the next wave of persistent models and agents. They are fast approaching. On the research side, reward hacking is something to really pay attention to. On the technical side, develop a deep understanding and do deep research on evals and sandboxing. They are going to be key technology going forward. I don't think frontier models trained on general-purpose capabilities should be applicable everywhere. It might be interesting to use them for things like scientific discovery. You might be safer and better off using more constrained and custom models for the majority of tasks. I think it's good to take a few hours digesting the recent progress in AI and strategizing carefully.
来源:@omarsar0 · x.com