跳转至

通过人机交互实现高效测试时适应

文章背景与核心概要

现代AI智能体通常基于海量数据进行训练,以具备广泛的通用能力。然而,它们产出的成果往往难以满足专业人士日常工作中所要求的严苛个人标准。在成功标准多样化且文档记录不足的开放式实际任务中,真正的专业性正体现在超越“平均水平”的独特洞察上。

为了解决这一痛点,本文提出了通过人机交互实现的测试时适应(Test-Time Adaptation through Human-Agent Interaction, TAHI)这一全新框架。该框架将跨会话的人机交互过程捕获为信号,用以弥合通往个人专属专业知识的差距。TAHI通过演进的评分标准(rubric)模块将个性化标准具象化,并将用户交互数据直接整合到智能体的上下文和权重中。在涵盖写作和视觉创作领域的30名个人用户(共600项任务)的评估中,TAHI在仅需数十项任务的极短时间内,就将独立任务成功率提升了4.5%至20.9%,并将跨用户的个性化成功率提升了最多8.8%


📌 Executive Summary

Modern AI agents are trained on population-scale data to possess broad, generalized capabilities. However, the artifacts they produce frequently fall short of the rigorous personal standards required by professionals. On open-ended, realistic tasks where success criteria are heterogeneous and poorly documented, genuine expertise lies in departing from the average.

This paper introduces Test-Time Adaptation through Human-Agent Interaction (TAHI), a novel framework that captures cross-session human-AI interactions as a signal to bridge the gap toward individual expertise. TAHI integrates user interaction data directly into agent context and weights while crystallizing personalized criteria through an evolving rubric module. Evaluated across 30 individuals in writing and visual creation (600 tasks total), TAHI improves solo task success by 4.5% to 20.9% within just tens of tasks, and generalizes personalized success improvements up to 8.8% across different users.


📋 Abstract

AI智能体经过群体规模数据的训练,能够编码出涵盖众多从业者能力的广泛技能。然而,它们产出的成果极少能达到专业人士愿意为其声誉担保的个人标准。在那些成功标准异质且缺乏充分记录的现实开放式任务中,个人专业性恰恰存在于对平均水平的提升和超越之中。

在实践中,迭代式的人机交互会显现出用户无法在初期完全具细指定、却在跨任务中反复应用的各项标准。我们认为,这种跨会话的交互数据是一种丰富且未被充分利用的信号,可用于缩小通往个人专业知识的差距。

在这项工作中,我们提出了通过人机交互实现的测试时适应(TAHI),它将这些信号整合到智能体的上下文和权重中,并通过一个演进的评分标准模块将每个用户的训练与评估标准具象化。我们在写作和视觉创作这两个高实用性领域对30名个人进行了智能体适应性测试,总计涉及600项任务。我们的智能体在仅需数十项任务的时间内,将独立任务成功率提升了4.5–20.9%。同时,我们演进的评分标准模块充当了可扩展的标注工具,所创建的评估标准比仅靠大模型或人类单独评估多捕捉了16.0–22.3%的失败案例。尽管智能体是针对个人进行适应性调整的,但我们表明,这些个性化智能体在跨用户时也能带来高达8.8%的成功率提升。

AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average.

In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise.

In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5–20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0–22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.


🔗 Full-Text & Resources

🔗 Full-Text & Resources


📊 Additional Metadata

  • Primary Subject: Artificial Intelligence (cs.AI)
  • Bibliographic Tools: NASA ADS, Google Scholar, Semantic Scholar
  • Associated Platforms: Code, data, and community demos are accessible via platforms integrated with arXivLabs (including Hugging Face, Papers with Code, and alphaXiv).

📊 Additional Metadata

  • Primary Subject: Artificial Intelligence (cs.AI)
  • Bibliographic Tools: NASA ADS, Google Scholar, Semantic Scholar
  • Associated Platforms: Code, data, and community demos are accessible via platforms integrated with arXivLabs (including Hugging Face, Papers with Code, and alphaXiv).