文章背景与核心概要
大语言模型(LLM)可能潜藏着隐蔽或具欺骗性的行为,例如后门触发器、潜伏代理部署线索、性能压制(sandbagging)或特定主题审查机制,这些行为通常只在特定条件下才会激活。由于这些隐藏行为极其稀疏且难以预测,在缺乏先验知识的情况下,标准检测方法往往束手无策。
为了解决这一问题,研究人员引入了一种名为激活匹配微调(Activation-Matched Finetuning, AMFT)的无监督检测方法,该方法无需预先了解触发器或目标行为。其核心方法是:选取一个可疑模型和一个公共参考模型(即“锚点模型”),然后在少量良性语料库上微调锚点模型,使其能够复现可疑模型的激活状态。由于良性语料库自然缺乏稀疏的触发条件,参考模型在学习可疑模型良性计算方式的同时,并不会习得其隐藏行为。因此,当在触发器提示词(及其语义邻近词)上进行评估时,两个模型之间的分歧(即残差)会出现显著飙升,从而可靠地向防御者发出异常隐藏行为存在的警报。针对第三方模型和自定义模型的实验验证了该方法的可靠性,即使面对旨在逃避检测的具防御意识的攻击,该方法依然表现出色。
Metadata
- arXiv ID: arXiv:2609.00351 [cs.CL]
- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI) - Authors: Robin Haselhorst, Lucie Flek, Florian Mai
- Submission Date: May 29, 2026
- Status: Under review
Access Links
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
Summary
Summary
Large language models can harbor latent or deceptive behaviors—such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship—that only activate under specific conditions. Because these hidden behaviors are extremely sparse and difficult to anticipate, standard detection methods often fail without prior knowledge of what to look for.
Large language models can harbor latent or deceptive behaviors—such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship—that only activate under specific conditions. Because these hidden behaviors are extremely sparse and difficult to anticipate, standard detection methods often fail without prior knowledge of what to look for.
To address this, researchers introduced Activation-Matched Finetuning (AMFT), an unsupervised detection method requiring no prior knowledge of triggers or target behaviors. The core methodology involves taking a suspect model and a public reference model ("anchor"), then finetuning the anchor to replicate the suspect model's activations on a small, benign corpus. Because a benign corpus naturally lacks sparse trigger conditions, the reference model learns the benign computations of the suspect model without adopting its hidden behaviors. Consequently, when evaluated on trigger prompts (and their semantic neighbors), the divergence—or residual—between the two models spikes significantly, reliably alerting defenders to the presence of unusual, hidden behaviors. Experiments across both third-party and custom models demonstrate the method's reliability, even when tested against defense-aware attacks designed to evade detection.
To address this, researchers introduced Activation-Matched Finetuning (AMFT), an unsupervised detection method requiring no prior knowledge of triggers or target behaviors. The core methodology involves taking a suspect model and a public reference model ("anchor"), then finetuning the anchor to replicate the suspect model's activations on a small, benign corpus. Because a benign corpus naturally lacks sparse trigger conditions, the reference model learns the benign computations of the suspect model without adopting its hidden behaviors. Consequently, when evaluated on trigger prompts (and their semantic neighbors), the divergence—or residual—between the two models spikes significantly, reliably alerting defenders to the presence of unusual, hidden behaviors. Experiments across both third-party and custom models demonstrate the method's reliability, even when tested against defense-aware attacks designed to evade detection.
Metadata
Metadata
- arXiv ID: arXiv:2609.00351 [cs.CL]
- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI) - Authors: Robin Haselhorst, Lucie Flek, Florian Mai
- Submission Date: May 29, 2026
- Status: Under review
- arXiv ID: arXiv:2609.00351 [cs.CL]
- Subjects: Computation and Language (
cs.CL); Artificial Intelligence (cs.AI)- Authors: Robin Haselhorst, Lucie Flek, Florian Mai
- Submission Date: May 29, 2026
- Status: Under review
Access Links
Access Links