无法追责的授权、日渐衰退的技能:绘制职场AI代理的风险图谱
文章背景与核心概要
随着组织加速将AI代理(AI Agents)引入工作场所,现有的宏观风险分类法往往无法捕捉特定于岗位层面的漏洞。本文提出了一个全面的框架,用于绘制和评估职场AI代理的社会技术风险。
研究人员开发了一个建模代理、目标与环境的多层框架,并将其应用于ONET数据库中的2,078个工作任务,生成了8,356个经过验证的风险场景,最终演化出全新的15类职场AI风险分类法*。核心洞察表明:增强并不比自动化天然更安全(因为过度依赖会侵蚀工人技能),代理的错误行为在人机交互边界处集中了最高密度的严重风险,而以人为中心的部署设计对于打造更安全的工作场所至关重要。
arXiv: arXiv:2608.08601 [cs.AI]
提交时间: 2026年8月9日
一级学科: 人工智能 (cs.AI)
二级学科: 多智能体系统 (cs.MA)
作者: Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka, Jie M. Zhang, Michael Luck, Elizabeth Black, Daniele Quercia
执行摘要 (Executive Summary)
As organizations increasingly integrate AI agents into the workplace, existing broad risk taxonomies fail to capture specific, job-level vulnerabilities. This paper introduces a comprehensive framework to map and evaluate the socio-technical risks of workplace AI agents.
The researchers developed a multi-layer framework modeling agents, goals, and environments, applying it to 2,078 job tasks from the ONET database to generate 8,356 validated risk scenarios. These scenarios culminated in a new 15-category workplace AI risk taxonomy*. Key insights reveal that augmentation is not inherently safer than automation (as overreliance erodes worker skills), erroneous agent actions represent the highest concentration of severe risks at the human-agent boundary, and human-centric deployment design is essential for safer workplaces.
随着组织越来越多地将AI代理整合到工作场所中,现有的广泛风险分类法未能捕捉到具体的工作层面的漏洞。本文引入了一个全面的框架来映射和评估职场AI代理的社会技术风险。
研究人员开发了一个多层框架,对代理、目标和环境进行建模,将其应用于ONET数据库中的2,078个工作任务,以生成8,356个经过验证的风险场景。这些场景最终形成了一个新的15类职场AI风险分类法*。核心见解表明,增强并不比自动化天生更安全(因为过度依赖会侵蚀工人的技能),错误的代理行为在人机边界处代表了最高集中度的严重风险,并且以人为中心的部署设计对于更安全的工作场所至关重要。
元数据与快速链接 (Metadata & Quick Links)
- View PDF: arXiv:2608.08601 PDF
- HTML Version: arXiv HTML (experimental)
- DOI: 10.48550/arXiv.2608.08601
- 查看 PDF: arXiv:2608.08601 PDF
- HTML 版本: arXiv HTML (实验性)
- DOI: 10.48550/arXiv.2608.08601
摘要 (Abstract)
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job-specific risks introduced by agents. To address this gap, we make three main contributions.
- Framework Development: We developed a multi-layer framework from a literature review of AI agents, modeling three core components and their interactions: agents, goals, and environment.
- Scenario Generation & Validation: We embedded this framework in a structured prompt and applied it to descriptions of 2,078 job tasks from the O*NET database, producing 8,356 risk scenarios labeled by severity and deployment mode (automation or augmentation). We validated these scenarios with 45 workers across 10 job roles and an independent LLM judge, confirming their plausibility and alignment with job tasks.
- Taxonomy Creation: We extended an existing taxonomy to create a 15-category taxonomy of workplace AI agent risks that covers all our risk scenarios.
为了预测来自AI代理的社会技术风险,组织需要分类法对其进行分类。然而,现有的AI风险分类法侧重于广泛的风险,并没有捕捉到代理引入的特定工作岗位风险。为了填补这一空白,我们作出了三项主要贡献。
- 框架开发: 通过对AI代理文献的回顾,我们开发了一个多层框架,对三个核心组件及其相互作用进行建模:代理、目标和环境。
- 场景生成与验证: 我们将此框架嵌入到结构化提示词中,并将其应用于O*NET数据库中2,078个工作任务的描述,生成了8,356个按严重程度和部署模式(自动化或增强)标记的风险场景。我们通过10个工作角色的45名工人和一个独立的LLM裁判验证了这些场景,确认了它们的合理性与工作任务的一致性。
- 分类法创建: 我们扩展了现有分类法,创建了一个涵盖所有风险场景的15类职场AI代理风险分类法。
核心发现 (Key Findings)
Our analysis highlights four primary findings regarding the deployment of workplace AI agents:
- Augmentation Harbors Hidden Risks: Augmentation is not inherently safe. Overreliance on AI agents can gradually erode workers' core skills and diminish effective oversight.
- Erroneous Agent Actions Dominate Severity: Erroneous Agent Actions account for the largest share of risk scenarios and feature the highest concentration of severe risks, many of which emerge directly at the human-agent boundary.
- Deployment Modes Dictate Risk Targets: Automation is primarily associated with organizational-level risks, whereas augmentation is predominantly tied to risks impacting individual workers.
- Superior Usability: Workers found our new taxonomy significantly easier to use for risk classification tasks compared to two existing frameworks, preferring it in 64% of non-tied comparisons against a recent generative AI risk taxonomy.
我们的分析强调了关于职场AI代理部署的四个主要发现:
- 增强隐藏着隐患: 增强并不天生安全。对AI代理的过度依赖会逐渐侵蚀工人的核心技能,并削弱有效的监督。
- 错误的代理行为主导了严重性: 代理的错误行为占风险场景的最大份额,并具有最高密度的严重风险,其中许多直接出现在人机交互边界。
- 部署模式决定风险目标: 自动化主要与组织层面的风险相关,而增强则主要与影响个体工人的风险挂钩。
- 卓越的可用性: 与现有的两个框架相比,工人们发现我们的新分类法在风险分类任务中明显更容易使用,在与近期生成式AI风险分类法的非平局比较中,有 64% 的人更青睐它。
Takeaway: Workplace AI agent risks do not stem from the technology alone; they depend fundamentally on how people work with agents and how those agents are deployed. Ensuring safer workplaces requires both inherently safer agents and carefully designed human-AI collaboration.
启示: 职场AI代理的风险并非完全源于技术本身;它们根本上取决于人们如何与代理一起工作以及如何部署这些代理。确保更安全的工作场所既需要本质上更安全的代理,也需要精心设计的人机协作。
引用 (Citation)
If you use or reference this work, you can cite it via BibTeX:
如果您使用或参考本工作,可以通过BibTeX对其进行引用:
@article{lamalfa2026unaccountable,
title={Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents},
author={La Malfa, Gabriele and Meegahapola, Lakmal and Bogucka, Edyta and Zhang, Jie M. and Luck, Michael and Black, Elizabeth and Quercia, Daniele},
journal={arXiv preprint arXiv:2608.08601},
year={2026}
}