论文概要
研究领域: ML 作者: Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon 发布时间: 2026-08-19 arXiv: 2608.19121
中文摘要
改善分子性质,如药物相似性或结合亲和力,是早期药物发现中的经常性任务。然而,在无约束化学空间中优化的分子如果不能合成,则实用价值有限。前向合成的策略梯度(PGFS)是一种合成感知的强化学习方法,用于分子改进,但其使用反应物嵌入预测使反应物选择间接,正如我们所示,这限制了学习效果。我们首先开发PGFS+,其中反应模板和第二反应物由可训练嵌入查找表表示。结合更有效的评分函数和RL算法,PGFS+显著改善了期望性质。然而,它暴露了一种奖励黑客失败模式:强大的反应物搜索可以将多样的输入分子映射到相同的高奖励磁分子,改善奖励同时崩溃输出多样性。因此,我们引入PGFS++,一种用于输入特定分子改进的合成感知强化学习框架。给定输入分子,PGFS++将其视为前向合成轨迹的起点,应用学习的反应模板与兼容的库存构建块,并产生具有改善目标性质、显式合成路线和与输入结构相似性的分子。分子改进任务的实验表明,PGFS++在保持高输出多样性的同时改善目标性质。
原文摘要
Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. However, molecules optimized in an unconstrained chemical space have limited practical value if they cannot be synthesized. Policy Gradient for Forward Synthesis (PGFS) is a synthesis-aware reinforcement learning method for molecular improvement, but its use of reactant embedding prediction makes reactant selection indirect, which, as we show, limits learning effectiveness. We first develop PGFS+, in which reaction templates and second reactants are represented by trainable embedding lookup tables. Combined with a more effective scoring function and RL algorithm, PGFS+ significantly improves the desired property. However, it exposes a reward-hacking failure mode: a …
— 自动采集于 2026-08-21
#论文 #arXiv #ML #小凯
