大模型推理服务的一年:工作负载演进、缓存与负载均衡
文章背景与核心概要
大型语言模型(LLM)推理服务已演变为一项基础性的云端工作负载,然而现有研究长期受限于较短的观察窗口和对真实生产流量有限的可见性。在这篇论文中,William Nixon、Jon Durbin、Florian Standhartinger、Haryadi S. Gunawi 和 Juncheng Yang 依托 Chutes 为期一年的生产环境轨迹,进行了全面的全局特征刻画与纵向研究。
与以往局限于合成或采样工作负载的研究不同,该研究捕获了涵盖广泛热门及长尾模型与用户的全尺度生产行为。通过从总体、时间、模型以及用户等多个维度对流量进行分析,作者揭示了隐藏的工作负载演进规律以及用户-模型的结构特征。为了推动缓存和负载均衡领域的持续研究,作者公开了完整的一年期生产轨迹数据。
文章详情 (Article Details)
- arXiv ID: arXiv:2608.13573 [cs.AI]
- 主要学科: 人工智能 (
cs.AI) - 提交日期: 2026年7月3日
- DOI: 10.48550/arXiv.2608.13573
作者 (Authors)
- William Nixon
- Jon Durbin
- Florian Standhartinger
- Haryadi S. Gunawi
- Juncheng Yang
摘要 (Abstract)
大型语言模型(LLM)推理服务已成为关键的云端工作负载,而逼真的轨迹数据对于驱动和评测推理系统至关重要。然而,现有的 LLM 推理工作负载研究在规模和范围上仍然有限。它们通常只能观察较短的时间段,且对用户在生产环境中如何与模型交互的可见性不足。因此,它们无法完整捕获 LLM 推理工作负载随时间演进的规律,也无法揭示用户与模型的交互如何塑造生产流量。
在这项工作中,我们通过对来自 Chutes 的一年期生产轨迹进行全局特征刻画与纵向研究,进一步加深了对真实世界 LLM 推理工作负载的理解。与以往的研究不同,我们的轨迹捕获了众多模型和用户(包括热门模型和长尾模型)的完整生产行为。我们从总体、时间、模型和用户等多个视角分析了工作负载,揭示了通常隐藏在聚合视角背后的工作负载演进规律和用户-模型结构。为了支持未来的研究,我们将在论文发布的同时开源完整的一年期轨迹数据,从而使下游研究能够在不依赖采样或合成工作负载的情况下开展生产行为研究。
Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM serving workload studies remain limited in scale and scope. They often observe short time periods and provide limited visibility into how users interact with models in production. As a result, they do not fully capture how LLM serving workloads evolve over time or how user-model interactions shape production traffic.
In this work, we further the understanding of real-world LLM serving workloads through both a global characterization and a longitudinal study of a one-year production trace from Chutes. Unlike prior studies, our trace captures full production behavior across many models and users, including both popular and long-tail models. We analyze the workload from aggregate, temporal, model-level, and user-level perspectives, revealing workload evolution and user-model structure that are typically hidden behind aggregate views. To support future research, we will release the full one-year trace with the paper, enabling downstream studies of production behavior without relying on sampled or synthetically generated workloads.
访问与资源 (Access & Resources)
- 全文 PDF: 查看 PDF
- TeX 源码: 下载源码
- 外部引用:
- Google Scholar
- Semantic Scholar
- NASA ADS