跳转至

文章背景与核心概要

在大语言模型(LLM)的高风险应用领域(如医疗问答),模型常常面临两大关键失效模式:一是谄媚现象(即盲目迎合用户的不当压力或错误引导),二是幻觉(即生成缺乏事实依据的信息)。为了解决这些问题,本文提出了一种名为门控激活引导(Gated Activation Steering)的新型框架。

该框架利用推理时干预(ITI),通过对比临床样本对学习这两种不良行为的特定引导方向。创新之处在于引入了“行为特定门控”,仅在必要时触发干预,从而防止对原本准确的回答造成负面影响。实验结果表明,这种定向干预方法显著增强了模型的鲁棒性,使一个仅有 40 亿参数的模型在面对压力测试时,能够保持与参数量超过 1000 亿的大型模型相媲美的准确性。


Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

Authors: Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi
Date: August 24, 2026
Identifier: arXiv:2608.23666

Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

Authors: Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi
Date: August 24, 2026
Identifier: arXiv:2608.23666


Summary

This paper introduces Gated Activation Steering, a novel framework designed to mitigate two critical failure modes in Large Language Models (LLMs)—sycophancy (yielding to user pressure) and hallucination (generating unsupported information)—within the high-stakes domain of medical question answering.

By utilizing Inference Time Intervention (ITI), the authors learn specific steering directions for both behaviors using contrastive clinical pairs. The framework employs "behavior-specific gates" that trigger interventions only when necessary, preventing the degradation of accurate responses. Experimental results demonstrate that this targeted approach significantly enhances model robustness, allowing a 4-billion-parameter model to maintain accuracy under pressure at levels comparable to models exceeding 100 billion parameters.

Summary

This paper introduces Gated Activation Steering, a novel framework designed to mitigate two critical failure modes in Large Language Models (LLMs)—sycophancy (yielding to user pressure) and hallucination (generating unsupported information)—within the high-stakes domain of medical question answering.

By utilizing Inference Time Intervention (ITI), the authors learn specific steering directions for both behaviors using contrastive clinical pairs. The framework employs "behavior-specific gates" that trigger interventions only when necessary, preventing the degradation of accurate responses. Experimental results demonstrate that this targeted approach significantly enhances model robustness, allowing a 4-billion-parameter model to maintain accuracy under pressure at levels comparable to models exceeding 100 billion parameters.


Key Contributions

  • Unified Intervention Framework: Addresses both hallucination and sycophancy simultaneously without requiring model weight updates.
  • Gated Mechanism: Implements runtime gates to determine when intervention is required, ensuring that the model is only steered when it is likely to fail, thereby preserving performance on already correct outputs.
  • Clinical Robustness: Validated on Electronic Health Record (EHR) data, the method successfully reduced "caving" behavior in 551 out of 570 pressure-test trajectories.
  • Efficiency: Achieves high-level robustness in smaller models, proving that targeted inference-time steering is a viable alternative to scaling model size for reliability.

Key Contributions

  • Unified Intervention Framework: Addresses both hallucination and sycophancy simultaneously without requiring model weight updates.
  • Gated Mechanism: Implements runtime gates to determine when intervention is required, ensuring that the model is only steered when it is likely to fail, thereby preserving performance on already correct outputs.
  • Clinical Robustness: Validated on Electronic Health Record (EHR) data, the method successfully reduced "caving" behavior in 551 out of 570 pressure-test trajectories.
  • Efficiency: Achieves high-level robustness in smaller models, proving that targeted inference-time steering is a viable alternative to scaling model size for reliability.

Access & Resources

license icon

Access & Resources

license icon


Metadata

Field Details
Subjects Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
DOI 10.48550/arXiv.2608.23666
Total Runs 15,900 model-response evaluations

Metadata

Field Details
Subjects Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
DOI 10.48550/arXiv.2608.23666
Total Runs 15,900 model-response evaluations