PertMind:通过细胞扰动数据强化学习激发大模型的涌现生物推理能力
文章背景与核心概要
尽管大语言模型(LLM)在描述分子和细胞机制方面表现出色,但其可扩展的后训练过程通常依赖于昂贵且需人工标注的生物推理轨迹。为了解决这一瓶颈,本文提出了 PertMind 框架,该框架将细胞扰动图谱转化为强化学习(RL)环境,利用实测的基因响应数据为生物推理提供可计算的奖励信号。
PertMind 通过结合“可信轨迹监督初始化”与“基因、通路及格式层面的强化信号”,在未见过的细胞背景下展现了卓越的扰动响应预测能力,同时保持了通用的语言处理能力。该模型无需针对特定任务进行后训练,即可在反向扰动识别、双重扰动推理、表型筛选优先级排序及生物过程解释等任务中实现零样本迁移,证明了基于实验终点的强化学习是构建通用生物推理模型的一种可扩展路径。
📌 摘要 (Summary)
虽然大语言模型(LLM)能够描述分子和细胞机制,但可扩展的后训练通常依赖于昂贵且需人工整理的生物推理轨迹。PertMind 提出了一种新范式:将细胞扰动图谱转化为强化学习(RL)环境,其中测得的基因响应为生物推理提供可计算的奖励。
通过结合可信轨迹监督初始化与基因、通路及格式层面的强化信号,PertMind 在未见过的细胞背景下的前向扰动响应预测中表现优异,同时保留了通用的语言能力。此外,它在无需特定任务后训练的情况下,零样本迁移至多种复杂的生物学任务——如反向扰动识别、双重扰动推理、表型筛选优先级排序以及生物过程解释——确立了基于扰动的强化学习作为通用生物推理的可扩展框架。
While large language models (LLMs) are capable of describing molecular and cellular mechanisms, scalable post-training typically relies on expensive, manually curated biological reasoning traces. PertMind presents a novel paradigm: transforming cellular perturbation atlases into reinforcement learning (RL) environments where measured gene responses supply computable rewards for biological reasoning.
By combining trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals, PertMind achieves superior performance in forward perturbation-response prediction across unseen cellular contexts while retaining general language capabilities. Furthermore, it generalizes zero-shot to various complex biological tasks—such as reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation—establishing perturbation-derived reinforcement learning as a scalable framework for general-purpose biological reasoning.
👥 作者 (Authors)
- Zhenchao Tang
- Xiaogang Xu
- Tianxu Lv
- Jiahui Guan
- Jiale Zhou
- Haohuai He
- Zhi Song
- Hanbo Huang
- Jiehui Huang
- Jiafei Wu
- Zhe Liu
📖 论文摘要 (Abstract)
大语言模型能够描述生物机制,但可扩展的后训练仍依赖于昂贵且需人工整理的生物推理轨迹。本文展示了细胞扰动图谱可以转化为强化学习环境,其中测得的基因响应为生物推理提供了可计算的奖励。我们引入了 PertMind,它结合了可信轨迹监督初始化与基因、通路及格式层面的强化信号。仅通过前向扰动响应预测进行训练,PertMind 在未见过的细胞背景下改进了响应推理,同时保留了通用的语言能力。它还在无需特定任务后训练的情况下,迁移至反向扰动识别、双重扰动推理、表型筛选优先级排序和生物过程解释等任务。PertMind 进一步生成了生物学特征,支持了多尺度下游任务中具有竞争力的基因、细胞和供体表征。这些结果支持了以下假设:基于实验终点的强化可以集中预训练模型中已有的可重用生物学策略。更广泛地说,基于扰动的强化学习为将不断扩展的实验图谱转化为通用生物推理的训练环境提供了一条可扩展的途径。
Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Trained only on forward perturbation-response prediction, PertMind improved response inference in unseen cellular contexts while retaining general language capabilities. It also transferred without task-specific post-training to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generated biological profiles that supported competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.
🔗 项目资源 (Project Resources)
- 项目主页: Shapsider - PertMind
- 代码仓库: GitHub - shapsider/PertMind
- 模型权重: Hugging Face - tzcfly/PertMind
- 全文 PDF: arXiv:2608.16419 PDF