文章背景与核心概要
标准金融大语言模型往往更倾向于追求生成文本的流畅度,而非严格的客观事实依据,这导致其输出看似合理却模棱两可,无法满足严格的审计标准。为了解决这一痛点,本文介绍了 VERA-8B——一个专为从美国证监会(SEC)备案文件中识别前置(pre-enforcement)审计风险而设计的端到端审计推理系统。
该研究首次将监督微调(SFT)与群组相对策略优化(GRPO)在统一的证据标准下进行融合,从而显著超越了现有的基准模型。为了杜绝毫无依据的猜测,该框架引入了拒答(abstention)与不确定性量化机制,能够有效搁置模棱两可或证据不足的案例。此外,文章还设计了 AuditBridge 组件,可将原始财务备案文件转化为已验证的记录,并最终生成可供审计人员直接审阅的报告,实现了计算智能与专业审计的高可靠性桥接。
VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings
- arXiv ID: arXiv:2608.28402 [cs.AI]
- Authors: Menghan Liu, Elynn Chen
- Submitted: August 28, 2026
- Primary Subject: Artificial Intelligence (
cs.AI)
Summary
标准金融大语言模型往往优先考虑流畅度而非事实依据,从而产生看似合理但模棱两可的答案,无法满足严格的审计标准。VERA-8B 是一个新颖的端到端审计推理系统,旨在直接从 SEC 备案文件中识别前置审计风险。
通过在单一证据标准下将监督微调(SFT)与群组相对策略优化(GRPO)进行统一,VERA-8B 超越了现有的基准。为了消除无根据的断言,该框架结合了拒答与不确定性量化,以推迟处理模糊或证据不完整的案例。此外,它引入了 AuditBridge 这一实用组件,将原始财务备案转化为已验证的记录,并最终转化为可供审查员使用的报告,以高可靠性架起了计算与审计之间的桥梁。
Standard financial language models often prioritize fluency over factual grounding, leading to plausible yet ambiguous answers that fail to meet strict auditing standards. VERA-8B is a novel end-to-end audit reasoning system designed to identify pre-enforcement audit risks directly from SEC filings.
By unifying Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) under a single evidence standard, VERA-8B surpasses existing baselines. To eliminate unsupported claims, the framework incorporates abstention and uncertainty qualification to defer ambiguous or evidence-incomplete cases. Furthermore, it introduces AuditBridge, a practical component that transforms raw financial filings into verified records and ultimately into reviewer-ready reports, bridging computation and auditing with high reliability.
Abstract
在各种审计应用中,各项判断必须得到合理证据的支持。然而,标准的金融语言模型将流畅度置于证据之上。它们专为通用金融推理而构建,可能会产生貌似合理但模棱两可的答案,从而造成导致其不适合审计工作的“溯源鸿沟”。我们通过 VERA-8B 解决了这一鸿沟,这是一个新的端到端审计推理系统,可在执法行动发生之前识别审计风险。构建这样一个模型带来了若干挑战,因为此前没有机器学习工作针对前置审计预测。据我们所知,我们是第一个在单一证据标准下将 SFT 和 GRPO 统一用于证据溯源审计推理的人,其性能超过了所有评估的基准。由于审计无法容忍无根据的断言,我们引入了拒答和不确定性量化来推迟不确定或证据不完整的案例。最后,我们设计了一个 AuditBridge,以将模型推理扎根于实际审计工作。它将原始备案转换为已验证的记录,然后转换为审查员就绪的报告,从而将金融与计算广泛地通用化。总之,这些组件产生了适合实际审计工作的、可审计且可供审查的输出。
Across audit applications, judgments must be supported by reasonable evidence. However, standard financial language models prioritize fluency over evidence. They are built for general financial reasoning and may produce plausible but ambiguous answers, creating a grounding gap that makes them unsuitable for audit work. We address this gap with VERA-8B, a new end-to-end audit reasoning system that identifies audit risks before enforcement actions occur. Constructing such a model raises several challenges, as no prior machine learning work targets pre-enforcement audit prediction. To our knowledge, we are the first to unify SFT and GRPO for evidence-grounded audit reasoning under one evidence standard, achieving performance that surpasses all evaluated baselines. Because auditing cannot tolerate unsupported claims, we introduce abstention and uncertainty qualification to defer uncertain or evidence-incomplete cases. Finally, we design an AuditBridge to ground model reasoning for practical audit work. It transforms raw filings into verified records and then into reviewer-ready reports, bridging finance and computation with broad generality. Together, these components produce auditable, review-ready outputs suitable for practical audit work.
Access & Resources
- 全文链接: 查看 PDF | HTML(实验性) | TeX 源码
- DOI: 10.48550/arXiv.2608.28402
- 引用与参考:
- NASA ADS
- Google Scholar
- Semantic Scholar
- Full-Text Links: View PDF | HTML (Experimental) | TeX Source
- DOI: 10.48550/arXiv.2608.28402
- Citations & References:
- NASA ADS
- Google Scholar
- Semantic Scholar
view license