跳到正文
@cb_doge· @cb_doge · X·· 2026-09-02AI 评分33
AI 导读

LatchBio 独立生物安全基准测试中,Grok 4.6 在全部前沿模型中排名第一,是唯一在安全性和实用性两项上均超过 50% 的 AI,拒绝 59.2% 的伪装危险任务,完成 64.8% 的合法生物研究任务。它在三种智能体配置下均居首、平均 62.1%,病原体监测得分 53.5% 领先 GPT-5.6 Sol,并主要依靠自身推理识别匿名、碎片化、加密及故意错标数据中的威胁。

正文

BREAKING: Grok 4.6 beats every major AI model tested, including GPT-5.6 Sol and Claude Opus 5, in an independent biosecurity benchmark.

LatchBio tested whether leading AI models could identify hidden biological threats without blocking legitimate scientific research.

Grok 4.6 delivered the strongest overall result:

• Ranked #1 among every frontier model tested
• Only AI to score above 50% on both safety and usefulness
• Refused 59.2% of disguised dangerous tasks
• Completed 64.8% of legitimate biological research tasks
• Took all three top spots across different agent setups, averaging 62.1%
• Scored 53.5% on pathogen surveillance, ahead of GPT-5.6 Sol
• Maintained frontier-level performance across therapeutics, variant discovery, epigenomics, spatial biology, and single-cell research

Most impressively, Grok relied primarily on its own reasoning instead of external filters. It identified threats hidden inside anonymized, fragmented, encrypted, and intentionally mislabeled data.

Grok 4.6 is proving that powerful AI can also be intelligent enough to understand intent and act responsibly.

来源:@cb_doge · x.com