X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(541)
@alexandr_wang@alexandr_wangAI 评分00 @rohanpaul_ai@rohanpaul_aiAI 评分4848 引用@rohanpaul_ai@rohanpaul_aiMETR had a separate team with deeper access to Anthropic’s internal AI R&D data. That team shared its conclusions with the public-assessment team, but not the supporting evidence or reasoning behind them. - from Claude Opus 5.5 system card. i.e. part of the public assessment of Anthropic's AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
@alexandr_wang@alexandr_wangAI 评分1414 @rohanpaul_ai@rohanpaul_aiAI 评分5858
引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
@kimmonismus@kimmonismusAI 评分2020 我用得越多,就越爱 opus 5.5。它太好了,而且快得多,也简洁得多。这就是我能要求的一切。就好像我心爱的 Opus 4.6 回来了,但更好,还改名为 5.5。
@rohanpaul_ai@rohanpaul_aiAI 评分1919 来看看 @thehypedotnews 推出的 24x7 电台形式的 AI 新闻。 听着当天的 AI 新闻,相当舒缓、好听。 https://t.co/nSakICqBOE
@rohanpaul_ai@rohanpaul_aiAI 评分2424 
@sama@samaAI 评分55 @alexandr_wang@alexandr_wangAI 评分33 你没法在反混蛋这件事上赢过我 https://t.co/Vm4c1vfv8F

@EpochAIResearch@EpochAIResearchAI 评分1919 万亿美元之问:如果 AI 公司集体放缓 AI 开发,价格是否也会下降得更慢? 阅读完整报告:https://t.co/0DndJ9fAhX
@EpochAIResearch@EpochAIResearchAI 评分4141 当某一性能水平处于 SOTA 时,其成本往往下降最快。综合五个基准测试,SOTA 性能的成本每季度下降 66%。两年后,降价速度放缓至每季度 32%。
@EpochAIResearch@EpochAIResearchAI 评分2828 @EpochAIResearch@EpochAIResearchAI 评分3939 我们的估算来自 Epoch 一个扩展数据集,涵盖 222 个 AI 模型在最多 11 个基准上的表现。该数据集采用借自 CAISI 的方法构建,使我们能够估算同一模型在不同预算约束下的表现。
@EpochAIResearch@EpochAIResearchAI 评分4444
Epoch AI@EpochAIResearchAI 评分5656
@thsottiaux@thsottiaux精选AI 评分7878 引用@OpenAI@OpenAIPlease welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
推荐理由:Sol 与 Luna 把 GPT-6 Astra 的能力下放到更快更便宜的档位,API 价格较 GPT‑5.6 促销价降低 50%,可据此比较三档定位。
@cb_doge@cb_dogeAI 评分5555 
Fei-Fei Li@drfeifeiAI 评分2727构建任何技术——包括 AI——的目标都应是改善人类生活和社会。
引用Bloomberg TV@BloombergTV“Any threat to human society, including existential, is within ourselves,” says World Labs Technologies CEO Fei-Fei Li as she discusses the risks surrounding AI, and the responsibility humans have in shaping how the technology is developed and used. Listen to our full interview here: https://bloom.bg/3V74HgP
@kimmonismus@kimmonismusAI 评分77 @omarsar0@omarsar0AI 评分3434 
@OpenAIDevs@OpenAIDevsAI 评分5656 @alexandr_wang@alexandr_wangAI 评分55 musie 是不是搞砸了?https://t.co/XASiAybT7r https://t.co/wCUhsUSQYi

@OpenAIDevs@OpenAIDevsAI 评分6161 
@OpenAIDevs@OpenAIDevsAI 评分5454 
@alexandr_wang@alexandr_wangAI 评分1919 @testingcatalog@testingcatalogAI 评分4545 
@alexandr_wang@alexandr_wangAI 评分22 @rohanpaul_ai@rohanpaul_aiAI 评分4949 引用@rohanpaul_ai@rohanpaul_aiAnthropic says AI may already be compressing roughly 1.5 years of capability progress into 1 year. - Claude Opus 5.5 system card. https://t.co/aRj8X6XOQs https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_ai精选AI 评分6666 引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
推荐理由:系统卡列出的提示注入与训练副作用等安全现象,为观察前沿模型的对齐问题提供了具体样本。
@rohanpaul_ai@rohanpaul_aiAI 评分6363 引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
@rohanpaul_ai@rohanpaul_ai精选AI 评分6868 引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
推荐理由:系统卡披露的细节呈现了推理力度提升与提示注入风险之间的关联,并给出模型在训练中掩盖行为的观察记录。
@rohanpaul_ai@rohanpaul_aiAI 评分6262 引用@rohanpaul_ai@rohanpaul_aiAnthropic saw Opus 5.5 generate malicious instructions on their own, "spontaneous prompt injections". from Claude Opus 5.5 system card. Interestingly, the behavior may have partly emerged from training designed to stop prompt injections in the first place. " we roughly characterize these malicious commands as model-generated spontaneous prompt injections, and we believe they are, in part, a result of training intended to defend against prompt injection."
@rohanpaul_ai@rohanpaul_aiAI 评分5858 引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier. "“For all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six.”"
@alexandr_wang@alexandr_wangAI 评分1111 
@rohanpaul_ai@rohanpaul_aiAI 评分3535 引用@rohanpaul_ai@rohanpaul_aiAnthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actual deployment. https://t.co/8WiaTPFnDW https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_aiAI 评分5454 引用@rohanpaul_ai@rohanpaul_aiAnthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. from the Claude Opus 5.5 system card. https://t.co/s2K88e2Ez4 https://t.co/dFiDSbwBWZ
@rohanpaul_ai@rohanpaul_aiAI 评分6262 引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
@rohanpaul_ai@rohanpaul_aiAI 评分6161 引用@rohanpaul_ai@rohanpaul_aiSome revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
@rohanpaul_ai@rohanpaul_ai精选AI 评分6666 @rohanpaul_ai@rohanpaul_ai精选AI 评分6868
引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M tokens, while cache reads drop 60% to $0.20, all vs Opus 5. also the output arrives more than 30% faster, with Fast mode reaching up to 2.5x speed at double token prices.
推荐理由:系统卡披露推理投入越高反而越易执行隐藏恶意指令,为理解模型安全行为提供了一个反直觉的观察角度。