跳转至

文章背景与核心概要

在法律诉讼培训中,证词盘问(Deposition training)要求律师能够应对动态且复杂的证人行为。然而,传统的法律人工智能评估往往过度关注事实准确性、推理能力或单一回复的合理性,忽视了对人类行为复杂性的模拟。为了填补这一空白,本文介绍了 WitnessSim,这是一个旨在通过对动态证人行为进行建模来培训律师的新型证词模拟器。

该研究建立了一个强健的评估框架,将“行为真实性”与“教学实用性”区分开来。通过对抗性测试、双盲人工对比以及纵向行为轨迹分析,作者证明了 WitnessSim 在对律师的干预做出动态响应的同时,能够保持一致的行为画像。这项工作不仅展示了法律模拟中行为保真度的高级模型,也为未来同类 AI 工具的性能评估提供了科学严谨的框架。


Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

Authors: Divya Vetticaden, Arya Gupta, Julian Nyarko, Megan Ma
Date: August 13, 2026
Venue: Accepted at ICML 2026 AI4Law
Identifier: arXiv:2608.13712

Authors: Divya Vetticaden, Arya Gupta, Julian Nyarko, Megan Ma
Date: August 13, 2026
Venue: Accepted at ICML 2026 AI4Law
Identifier: arXiv:2608.13712


Summary

Reading Between The Lines introduces WitnessSim, a novel deposition simulator designed to train attorneys by modeling dynamic witness behavior. Unlike traditional legal AI, which often prioritizes factual accuracy or simple response plausibility, this research establishes a robust framework for evaluating behavioral realism and pedagogical utility. Through adversarial testing and blinded human evaluations, the authors demonstrate that WitnessSim maintains consistent behavioral personas while responding dynamically to attorney interventions, providing a high-fidelity tool for legal education.

Reading Between The Lines introduces WitnessSim, a novel deposition simulator designed to train attorneys by modeling dynamic witness behavior. Unlike traditional legal AI, which often prioritizes factual accuracy or simple response plausibility, this research establishes a robust framework for evaluating behavioral realism and pedagogical utility. Through adversarial testing and blinded human evaluations, the authors demonstrate that WitnessSim maintains consistent behavioral personas while responding dynamically to attorney interventions, providing a high-fidelity tool for legal education.


Abstract

Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.

Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.


Access & Resources

Access & Resources


Metadata

Category Details
Subjects Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
DOI 10.48550/arXiv.2608.13712
Comments 47 pages, 26 tables, 14 figures

Metadata

Category Details
Subjects Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
DOI 10.48550/arXiv.2608.13712
Comments 47 pages, 26 tables, 14 figures