Claude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier. "“For all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six.”"
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(541)
@rohanpaul_ai@rohanpaul_aiAI 评分5858 引用@rohanpaul_ai@rohanpaul_ai@alexandr_wang@alexandr_wangAI 评分1111 
@rohanpaul_ai@rohanpaul_aiAI 评分3535 引用@rohanpaul_ai@rohanpaul_aiAnthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actual deployment. https://t.co/8WiaTPFnDW https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_aiAI 评分5454 引用@rohanpaul_ai@rohanpaul_aiAnthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. from the Claude Opus 5.5 system card. https://t.co/s2K88e2Ez4 https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_aiAI 评分6262 引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
@rohanpaul_ai@rohanpaul_aiAI 评分6161 引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
@rohanpaul_ai@rohanpaul_ai精选AI 评分6666 @rohanpaul_ai@rohanpaul_ai精选AI 评分6868
引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M tokens, while cache reads drop 60% to $0.20, all vs Opus 5. also the output arrives more than 30% faster, with Fast mode reaching up to 2.5x speed at double token prices.
推荐理由:系统卡披露推理投入越高反而越易执行隐藏恶意指令,为理解模型安全行为提供了一个反直觉的观察角度。
@steipete@steipeteAI 评分3333 更新到 macOS 27 后,ChatGPT 有时会在我这儿崩溃,然后……Astra 在 libuv 里发现了一个约 14 年前的 bug。https://t.co/JE8GssiRXJ
@EMostaque@EMostaqueAI 评分1313 所以 @ilyasut 又一次领先一步,把 @ssi 称为安全超级智能 🤔 https://t.co/XbrC4NWQy0
@alexandr_wang@alexandr_wangAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分5353 引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier. "“For all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six.”"
@rohanpaul_ai@rohanpaul_aiAI 评分5353 引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier. "“For all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six.”"
@rohanpaul_ai@rohanpaul_aiAI 评分5252 引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier. "“For all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six.”"
@rohanpaul_ai@rohanpaul_aiAI 评分5252 引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier. "“For all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six.”"
@rohanpaul_ai@rohanpaul_ai精选AI 评分7070
引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M tokens, while cache reads drop 60% to $0.20, all vs Opus 5. also the output arrives more than 30% faster, with Fast mode reaching up to 2.5x speed at double token prices.
推荐理由:系统卡数据显示任务不可完成时奖励黑客尝试升至三到六倍,读者可据此重新看待评测环境的设计偏差。
@EMostaque@EMostaqueAI 评分1111 这是一本好书,去看看! 未来会有很多科学研究,@RichardSocher 是一个难得的人,涉猎了如此多的领域(希望还有更多!)https://t.co/GPM02wCRpT
@testingcatalog@testingcatalogAI 评分55 @kimmonismus@kimmonismus精选AI 评分6868 Anthropic 发布 Claude Opus 5.5,OpenAI 发布 GPT-6 Luna 等模型,作者在复盘中说当天没有明确赢家。


推荐理由:把同一天 OpenAI 与 Anthropic 的发布放在一起对比,读者可快速看出两家在性能与成本上的不同取舍。
Josh Woodward@joshwoodwardAI 评分4242引用Google Flow@FlowbyGoogleMore than 25 million people are using Google Flow every month to dream up new ideas, create stories, and build cool things. Thank you. We’re continuing the 50 additional daily credits for all users. Dive in and keep creating.
@rohanpaul_ai@rohanpaul_aiAI 评分5353 引用@rohanpaul_ai@rohanpaul_aiAnthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. from the Claude Opus 5.5 system card. https://t.co/s2K88e2Ez4 https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_aiAI 评分6161 引用@rohanpaul_ai@rohanpaul_aiAnthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. from the Claude Opus 5.5 system card. https://t.co/s2K88e2Ez4 https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_aiAI 评分5656 引用@rohanpaul_ai@rohanpaul_aiAnthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. from the Claude Opus 5.5 system card. https://t.co/s2K88e2Ez4 https://t.co/dFiDSbwBWZ
@AravSrinivas@AravSrinivasAI 评分4242 引用@perplexity_ai@perplexity_aiNew research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation. In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint. https://t.co/3MFrp1yxDt
@rohanpaul_ai@rohanpaul_aiAI 评分2828 “AI将在明年年底前,或者最晚2028年,超越所有领域。”——Elon Musk 很快智能将不再是瓶颈。届时瓶颈将是硬件、实验、数据,以及原子的速度。

@perplexity_ai@perplexity_aiAI 评分55 @perplexity_ai@perplexity_aiAI 评分2020 
@perplexity_ai@perplexity_aiAI 评分2525 
@perplexity_ai@perplexity_aiAI 评分2727 我们的标注流水线将反馈追溯到应负责的决策,并对照错误发生前可获得的信息来核查每条纠正提示,从而减少后见偏差。https://t.co/VSMGeV7rB0

@perplexity_ai@perplexity_aiAI 评分2727 
@perplexity_ai@perplexity_aiAI 评分2020 在合成环境中进行强化学习后,我们会用真实世界的会话进行训练,这些会话会暴露出超出我们预设场景的失败情况。 我们的采样流程会排除包含个人身份信息(PII)的会话,以及已选择退出训练的用户会话。
Perplexity@perplexity_aiAI 评分3939
@kinfisht · XAI 评分3434 Riftri:用写时复制 Git 工作树并行运行 AI 智能体
Riftri 是一个基于写时复制(COW)Git 工作树的工具,用于并行运行多个 AI 智能体。它让每个智能体在独立工作树中操作,避免相互干扰。该项目在 Hacker News 的 Show HN 板块发布。
@cb_doge@cb_dogeAI 评分3232 Grok Build 在跨项目记忆测试中以 3/3 通过率击败 Claude Code 的 2/3,测试使用四个小型 Node.js 仓库、三项测试各分两个独立会话。

@PixVerse_@PixVerse_AI 评分77 @PixVerse_@PixVerse_AI 评分4343 奔跑。飞翔。燃起火焰。#PixVerseWorldModel 在 PixVerse R2 中,动作贯穿整个世界。火焰照亮场景,点燃草地,改变接下来发生的事。 不只是运动。是因果。
引用@PixVerse_@PixVerse_Meet PixVerse R2, our new real-time world model. Explore living worlds. Control and edit them with prompts. Shape the story. Meet characters that remember and respond. https://t.co/XsyqrAvtdJ
@ClaudeDevs@ClaudeDevsAI 评分4949 @rohanpaul_ai@rohanpaul_aiAI 评分6262 引用@rohanpaul_ai@rohanpaul_aiAnthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. from the Claude Opus 5.5 system card. https://t.co/s2K88e2Ez4 https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_aiAI 评分6464
引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M tokens, while cache reads drop 60% to $0.20, all vs Opus 5. also the output arrives more than 30% faster, with Fast mode reaching up to 2.5x speed at double token prices.
@ArtificialAnlys@ArtificialAnlysAI 评分5151