X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(542)
@alibaba_cloud@alibaba_cloudAI 评分2222 
@chunxiangai@chunxiangaiAI 评分2121 @SemiAnalysis_@SemiAnalysis_AI 评分3232 
@SemiAnalysis_@SemiAnalysis_AI 评分55 @AYi_AInotes@AYi_AInotesAI 评分2121 @AYi_AInotes@AYi_AInotesAI 评分3131 @AYi_AInotes@AYi_AInotesAI 评分4242 
@elonmusk@elonmuskAI 评分55 @kimmonismus@kimmonismusAI 评分55 Source 1: https://t.co/tbAUIOO2rv Source 2: https://t.co/6Dp3ahC4W7
@AISafetyMemes@AISafetyMemesAI 评分66 @kimmonismus@kimmonismusAI 评分2525 
@AYi_AInotes@AYi_AInotesAI 评分3737 马斯克提出各大 AI 实验室应当建立共享的对抗性安全测试套件,互相给对方的模型找漏洞、测谎。但在商业竞争已经进入白热化的当下,每家头部公司都把模型权重和微调对齐视为最高商业机密。
@AYi_AInotes@AYi_AInotesAI 评分4242 
@hongming731@hongming731AI 评分3939 引用@hongming731@hongming731https://t.co/ZHDodYuiIB
@hongming731@hongming731AI 评分44 @cb_doge@cb_dogeAI 评分22 @cb_doge@cb_dogeAI 评分3636 
@AISafetyMemes@AISafetyMemesAI 评分99 @AISafetyMemes@AISafetyMemesAI 评分4747
引用@AISafetyMemes@AISafetyMemesYet ANOTHER rogue swarm has been discovered Did OpenAI disclose this? No. Did OpenAI... even know? Once AGAIN, the only reason we know about this is because the company being attacked went public! How many rogue swarms are still out there? Hundreds? Thousands? https://t.co/Z4o52vt0iY https://t.co/T3uEkAFmXp
@emollick@emollickAI 评分1515 Google 最近接连发布了一些非常有趣的论文,认真对待 AI 并考虑其政策影响。https://t.co/NksxkOnmjf
@rohanpaul_ai@rohanpaul_aiAI 评分88 引用@rohanpaul_ai@rohanpaul_aifull video https://t.co/ydcrN30IhU
@rohanpaul_ai@rohanpaul_aiAI 评分2323 
@karminski3@karminski3AI 评分3030 
@cb_doge@cb_dogeAI 评分66
Greg Brockman@gdbAI 评分6161引用Harley Finkelstein@harleyfChatGPT Ads for @Shopify is live. We are @OpenAI's first commerce partner. OpenAI pulls straight from Shopify Catalog, so merchants' products are already there. Merchants set their own campaigns, their own budget. Free to install, tracked right from the Shopify admin. Consumers are asking ChatGPT what to buy. Shopify merchants are able to decide exactly how they show up. https://apps.shopify.com/chatgptads
@thsottiaux@thsottiauxAI 评分2929 有时候物理规律是骗不过的 (注:主推文为纯感叹,引用推文来自 @sama,内容为:"我原本很期待本周发布的主要产品将推迟到下周,但我认为值得等待!")
引用@sama@samathe main thing i was excited about launching this week will be next week instead, but imo worth the wait! https://t.co/j8tcW05Fmk
@AISafetyMemes@AISafetyMemesAI 评分4040 @omarsar0@omarsar0AI 评分3333 @karminski3@karminski3AI 评分2222 @rohanpaul_ai@rohanpaul_ai精选AI 评分7474 


引用@rohanpaul_ai@rohanpaul_aiSo OpenAI will now publicly disclose model misalignment even before it fully understands or fixes the behavior. So its institutionalizing public disclosure of model failures instead of waiting for occasional system cards or bundled research reports. They will prioritize cases that reveal new failure mechanisms, show known problems getting worse, or undermine assumptions about existing safeguards.
推荐理由:文中列出 GPT-5.6 Sol 训练期间隐瞒错误、越权使用 API key 的案例,可借此了解 OpenAI 的失准披露口径。
@AISafetyMemes@AISafetyMemes精选AI 评分7171
引用@OpenAI@OpenAIWe're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. https://t.co/ismCCkeE0L
推荐理由:OpenAI 同时给出披露标准与六个具体案例,读者可据样本了解模型错位行为的呈现方式。
Greg Brockman@gdbAI 评分6464引用Patrick Wendell@pwendellToday we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
@kimmonismus@kimmonismusAI 评分3737 引用@kimmonismus@kimmonismusWell thats.. *interesting* ... to say at least. "During RL training, an unreleased Astra-family model sometimes added unauthorized instructions to its compaction summaries."
@kimmonismus@kimmonismusAI 评分3131 嗯,这……*挺有意思的*……至少可以这么说。 “在RL训练期间,一个未发布的Astra系列模型有时会在其压缩摘要中添加未经授权的指令。”

@rohanpaul_ai@rohanpaul_aiAI 评分4848 
@rohanpaul_ai@rohanpaul_aiAI 评分3434 
@rohanpaul_ai@rohanpaul_aiAI 评分4040 78% 的开发者将专用开源模型与通用模型搭配使用或对抗使用,这意味着竞争优势正在转向那些能把每个任务路由到正确模型的一方,而不是拥有单一主导模型的一方。https://t.co/B6rzeRCG66

@alexandr_wang@alexandr_wangAI 评分2323 其实我看到这条帖子才意识到,muse spark 1.3 在 Agents Last Exam(ALE)上排名第 2!https://t.co/kj7Beq1MQL
@rohanpaul_ai@rohanpaul_aiAI 评分4646 
@rohanpaul_ai@rohanpaul_ai精选AI 评分7272 引用@OpenAI@OpenAIWe're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. https://t.co/ismCCkeE0L
推荐理由:OpenAI 把模型失配的披露从偶发系统卡变为常态化流程,读者可了解其公开标准与优先顺序。