文章背景与核心概要
在大语言模型(LLM)的实际应用中,统一分配推理预算往往会带来巨大的计算开销,甚至导致“过度思考(over-thinking)”带来的性能惩罚,这在依赖复杂视觉布局驱动难度的文档中心化任务中尤为明显。为了解决这一挑战,本文作者推出了 BudgetDoc,这是首个旨在为三种不同文档任务中的模型-预算-性能权衡提供显式监督的多模态基准。
借助 BudgetDoc,作者训练了 DRB(Document-Reasoning Balancer)——一个参数量约为 10 亿的预检(pre-flight)估计器,它结合了 SigLIP-2 与 Qwen3-0.6B。DRB 能够准确预测不同预算级别下模型的序数性能,加权 F1 分数达到了 0.753。在跨 5 个前沿模型和 3 个数据集动态分配推理预算时,DRB 在 15 种配置中的 9 种达到了甚至超越了全最大预算基线的 F1 分数,同时大幅削减了计算成本。此外,初步评估还凸显了 DRB 在泛化至跨模型选择方面的巨大潜力。
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference
Authors: Zishan Ahmad, Vishal Vaddina
Submitted: August 19, 2026
Primary Subject: Artificial Intelligence (cs.AI)
arXiv ID: arXiv:2608.18591 [cs.AI]
Summary
Uniformly allocating reasoning budgets to Large Language Models (LLMs) is often computationally expensive and can lead to over-thinking penalties, particularly in document-centric tasks where complex visual layouts drive the difficulty.
To overcome this challenge, the authors introduce BudgetDoc, the first multimodal benchmark designed to provide explicit supervision for model-budget-performance trade-offs across three distinct document tasks. Using BudgetDoc, they train DRB (Document-Reasoning Balancer)—an approximately 1-billion-parameter pre-flight estimator combining SigLIP-2 and Qwen3-0.6B. DRB accurately predicts ordinal model performance across different budget levels, achieving a 0.753 weighted F1 score.
When dynamically allocating reasoning budgets across five frontier models and three datasets, DRB matched or improved F1 scores compared to always-maximum-budget baselines in 9 out of 15 configurations while drastically reducing computational costs. Furthermore, preliminary evaluations highlight DRB's promising potential to generalize toward cross-model selection.
Metadata
| Field | Details |
|---|---|
| Subjects | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as | arXiv:2608.18591 [cs.AI] |
| DOI | 10.48550/arXiv.2608.18591 |
| License | Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International |
Links & Full-Text Access
- PDF: View PDF
- HTML: Experimental HTML Version
- Source: TeX Source
- External Citations: Google Scholar | Semantic Scholar | NASA ADS
