跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-22AI 评分45
AI 导读

阿里与字节的论文指出,AI 智能体已无法按普通 LLM 请求方式服务,因为性能瓶颈常落在工具、记忆、环境与模型的协同上,并会随请求和部署变化转移到嵌入向量、数据库、沙箱或网络传输。

正文

New Alibaba ByteDance paper shows that AI agents can no longer be served like ordinary LLM requests.

Because most of the performance problem now sits across tools, memory, environments, and the model together.

Bottleneck can also move from the LLM to embeddings, databases, sandboxes, or network transfer as the request and deployment change.

This paper builds AgentSysBench around 10 agentic applications and finds that model inference is often no longer the main bottleneck.

They find task-aware serving cuts latency by 29–40%, communication-aware placement delivers up to a 4.5× speedup, state offloading cuts memory by 4.6×, and caching removes 35.2% of redundant search calls.

So optimizing tokens per second is no longer enough. Agent infrastructure has to schedule models, tools, memory, and communication as one workload.

– arxiv. org/abs/2608.15127

Title: "From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems"

来源:@rohanpaul_ai · x.com