X:阿里云 / Alibaba Cloud
@alibaba_cloud · X
切换来源
@alibaba_cloud@alibaba_cloudAI 评分4949 Qwen3.7-Max 以 77.3B tokens 的使用量登顶 @OpenRouter 趋势 LLM 榜单。 而这仅仅是个开始。 👇 int.alibabacloud.com/m/10004…

@alibaba_cloud@alibaba_cloudAI 评分5454 引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlysArtificial Analysis and IBM Research are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA’s SRE tasks benchmark model performance on Kubernetes incident response, where models must diagnose live systems by reading logs, tracing dependencies, and identifying root-cause entities across complex infrastructure. The underlying ITBench dataset has been developed by @IBM's Software Innovation Lab, leveraging IBM’s deep expertise in enterprise IT operations Artificial Analysis has worked closely with IBM over the last 6 months to develop a implementation of the dataset for frontier AI evaluation, beginning with Site Reliability Engineering (SRE) and expanding to Financial Operations (FinOps) and Chief Information Security Officer (CISO) tasks over time ITBench-AA SRE overview: ➤ 59 SRE tasks in total: 40 public tasks and 19 brand new, held-out tasks ➤ Each task provides a Kubernetes incident snapshot containing alerts, events, traces, metrics, logs, and application topology. The model must identify the minimal set of independent root-cause Kubernetes entities responsible for the incident ➤ Faults span typical SRE failure modes including infrastructure, service, application, and chaos-injected incidents, such as resource quota exhaustion, rollout failures, connection pool exhaustion, and network partitions Methodology details: ➤ Agentic harness: each task is solved by the model running in our open-source Stirrup reference harness, with shell access to a sandboxed file system containing the relevant logs and snapshots. 100-turn cap per task, 3 repeats per task ➤ Models submit a list of root-cause entities (Kubernetes Deployments, Services, Pods, etc.) they believe caused the incident. Each submission is compared against a ground-truth set of root causes provided by IBM Research ➤ Scoring uses average precision at full recall: if a model misses any of the ground-truth root causes, it scores 0.0 for that repeat. If it identifies all of them, it is awarded a score equal to its precision - the share of its submitted entities that are actual root causes, i.e. true positives / (true positives + false positives). The headline score is the average across 59 tasks × 3 repeats. ➤ The harness (Stirrup) is held constant across all evaluated models, allowing an apples-to-apples comparison between models. Key findings: ➤ Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads at 47%, followed by GPT-5.5 (xhigh) at 46% and Qwen3.7 Max at 42% ➤ All frontier models score below 50%, making ITBench-AA SRE one of the least saturated agentic benchmarks in our suite. For context, frontier models score considerably higher on Terminal-Bench ➤ Turn counts vary nearly 3x and longer trajectories do not translate to higher accuracy. GPT-5.5 (xhigh) averages 31 turns per task at 46%, while Gemini 3.1 Pro Preview averages 83 turns at 30%. Models that over-investigate tend to surface upstream fault-injection mechanisms or co-occurring symptoms as false positives ➤ GLM-5.1 (Reasoning) leads open weights models at 40%, effectively tied with Gemini 3.5 Flash (high). DeepSeek V4 Pro (Reasoning, Max Effort) follows at 38%, with Gemma 4 31B (Reasoning) at 37%, ahead of Gemini 3.1 Pro Preview at 30%
@alibaba_cloud@alibaba_cloudAI 评分5050 📢Qwen3.7-Max 刚刚在 ITbench-AA 上排名第 3——这是一个全新的基准测试,用来评估模型处理真实企业 IT 任务(智能体风格)的能力。 🔧智能体时代,选 Qwen。🏃🏃
引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlysArtificial Analysis and IBM Research are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA’s SRE tasks benchmark model performance on Kubernetes incident response, where models must diagnose live systems by reading logs, tracing dependencies, and identifying root-cause entities across complex infrastructure. The underlying ITBench dataset has been developed by @IBM's Software Innovation Lab, leveraging IBM’s deep expertise in enterprise IT operations Artificial Analysis has worked closely with IBM over the last 6 months to develop a implementation of the dataset for frontier AI evaluation, beginning with Site Reliability Engineering (SRE) and expanding to Financial Operations (FinOps) and Chief Information Security Officer (CISO) tasks over time ITBench-AA SRE overview: ➤ 59 SRE tasks in total: 40 public tasks and 19 brand new, held-out tasks ➤ Each task provides a Kubernetes incident snapshot containing alerts, events, traces, metrics, logs, and application topology. The model must identify the minimal set of independent root-cause Kubernetes entities responsible for the incident ➤ Faults span typical SRE failure modes including infrastructure, service, application, and chaos-injected incidents, such as resource quota exhaustion, rollout failures, connection pool exhaustion, and network partitions Methodology details: ➤ Agentic harness: each task is solved by the model running in our open-source Stirrup reference harness, with shell access to a sandboxed file system containing the relevant logs and snapshots. 100-turn cap per task, 3 repeats per task ➤ Models submit a list of root-cause entities (Kubernetes Deployments, Services, Pods, etc.) they believe caused the incident. Each submission is compared against a ground-truth set of root causes provided by IBM Research ➤ Scoring uses average precision at full recall: if a model misses any of the ground-truth root causes, it scores 0.0 for that repeat. If it identifies all of them, it is awarded a score equal to its precision - the share of its submitted entities that are actual root causes, i.e. true positives / (true positives + false positives). The headline score is the average across 59 tasks × 3 repeats. ➤ The harness (Stirrup) is held constant across all evaluated models, allowing an apples-to-apples comparison between models. Key findings: ➤ Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads at 47%, followed by GPT-5.5 (xhigh) at 46% and Qwen3.7 Max at 42% ➤ All frontier models score below 50%, making ITBench-AA SRE one of the least saturated agentic benchmarks in our suite. For context, frontier models score considerably higher on Terminal-Bench ➤ Turn counts vary nearly 3x and longer trajectories do not translate to higher accuracy. GPT-5.5 (xhigh) averages 31 turns per task at 46%, while Gemini 3.1 Pro Preview averages 83 turns at 30%. Models that over-investigate tend to surface upstream fault-injection mechanisms or co-occurring symptoms as false positives ➤ GLM-5.1 (Reasoning) leads open weights models at 40%, effectively tied with Gemini 3.5 Flash (high). DeepSeek V4 Pro (Reasoning, Max Effort) follows at 38%, with Gemma 4 31B (Reasoning) at 37%, ahead of Gemini 3.1 Pro Preview at 30%
@alibaba_cloud@alibaba_cloudAI 评分2121 GitHub🔗: github.com/alibaba/ANOLIS 快速入门指南🔗: int.alibabacloud.com/m/10004…
@alibaba_cloud@alibaba_cloudAI 评分5454 @alibaba_cloud@alibaba_cloudAI 评分1919 
@alibaba_cloud@alibaba_cloudAI 评分2222 
@alibaba_cloud@alibaba_cloudAI 评分2727 
@alibaba_cloud@alibaba_cloudAI 评分2020 
@alibaba_cloud@alibaba_cloudAI 评分3838 
@alibaba_cloud@alibaba_cloudAI 评分2424 GitHub🔗:github.com/alibaba/ANOLISA 快速入门指南🔗:int.alibabacloud.com/m/10004…
@alibaba_cloud@alibaba_cloudAI 评分5757 @alibaba_cloud@alibaba_cloudAI 评分4141 
@alibaba_cloud@alibaba_cloudAI 评分2626 
@alibaba_cloud@alibaba_cloudAI 评分2020 
@alibaba_cloud@alibaba_cloudAI 评分5555 阿里云在 QwenConference2026 上发布全开放的 AI 生态,除 Qwen 外,其他主流模型现可直接在 Model Studio 和 qwencloud.com 上访问。

@alibaba_cloud@alibaba_cloudAI 评分3434 用 Qwen3.7-Max 驱动 Hermes Agent。去看看 @NousResearch 🚀 (引用推文:Hermes Agent 现已支持 Qwen 3.7 Max)
引用Nous Research (@NousResearch)@NousResearchQwen 3.7 Max is now supported in Hermes Agent
@alibaba_cloud@alibaba_cloudAI 评分2222 阿里云被 Omdia 的 Agentic AI Market Radar 评为领导者。Omdia 重点肯定了阿里云在每一层的全栈能力,认可其为首个围绕 Agent 范式来构建整个平台的云服务商。

@alibaba_cloud@alibaba_cloudAI 评分3838 1M 上下文。更智能的推理。更多可能性。很高兴看到 Qwen3.7 Max 现已在 Go 中上线,支持 @opencode 🚀
引用OpenCode (@opencode)@opencodeQwen3.7 Max now available in Go - text only - 1M context - smartest model in the Qwen family to date
@alibaba_cloud@alibaba_cloudAI 评分3838 
@alibaba_cloud@alibaba_cloudAI 评分4343 引用PyTorch (@PyTorch)@PyTorchWe’re excited to welcome @alibaba_cloud as a Platinum Member of the PyTorch Foundation 🎉 Alibaba Cloud is a global leader in full-stack AI infrastructure and the force behind Qwen—one of the world’s most influential open-weight model families. Having run PyTorch at massive scale across diverse hardware, they bring invaluable, production-hardened engineering expertise to the upstream community.
@alibaba_cloud@alibaba_cloudAI 评分22 
@alibaba_cloud@alibaba_cloudAI 评分99 
@alibaba_cloud@alibaba_cloudAI 评分1818 
@alibaba_cloud@alibaba_cloudAI 评分55 
@alibaba_cloud@alibaba_cloudAI 评分77 
@alibaba_cloud@alibaba_cloudAI 评分1212 
@alibaba_cloud@alibaba_cloudAI 评分1111 看看创作者们用 AI 为 #DiveIntoYourWorld 打造了什么 👇🧵


@alibaba_cloud@alibaba_cloudAI 评分2626 
@alibaba_cloud@alibaba_cloudAI 评分1818 
@alibaba_cloud@alibaba_cloudAI 评分3030 
@alibaba_cloud@alibaba_cloudAI 评分3232 
@alibaba_cloud@alibaba_cloudAI 评分4141 
@alibaba_cloud@alibaba_cloudAI 评分2727 

@alibaba_cloud@alibaba_cloudAI 评分2727 

@alibaba_cloud@alibaba_cloudAI 评分2222 在 #QwenConference2026 上,总理公署高级政务部长、全国职工总会(NTUC)副秘书长 Desmond Tan 强调了技术生态的价值:🧵👇

@alibaba_cloud@alibaba_cloudAI 评分1212 @alibaba_cloud@alibaba_cloudAI 评分1717 “像阿里云这样的技术合作伙伴发挥着重要作用。十多年来,阿里云一直将新加坡作为其国际业务与运营的锚点,以此为基础在区域内不同市场开发和部署创新。”