[论文] SWE-Prime: Fewer Trajectories, Better Performance

## 论文概要 **研究领域**: NLP **作者**: Dewu Zheng, Ruizhe Ye, Ya...

论文概要

研究领域: NLP 作者: Dewu Zheng, Ruizhe Ye, Yanlin Wang 发布时间: 2026-08-28 arXiv: 2508.11370

中文摘要

为提升大语言模型解决真实软件问题的能力,先前工作专注于构建大规模智能体轨迹数据集,并对成功轨迹进行监督微调(SFT)。然而,任务成功并不能保证高质量监督:成功轨迹仍可能包含无效、冗余或高风险的步骤。直接使用此类轨迹进行SFT会引入噪声监督,并鼓励模型模仿不良的问题解决行为。因此,我们提出SWE-Prime——一种多粒度、两阶段的SFT数据选择方法,在轨迹和段级别逐步过滤训练数据。具体而言,第一阶段基于过程质量、结果质量和数据代表性进行轨迹级筛选,选择高质量且有代表性的成功轨迹子集。第二阶段通过将连续步骤分组为语义段来进行段级选择,并根据每段对最终解决方案的贡献、可学习性和潜在风险进行评估。在SFT期间,所有段保留在序列中以保持上下文,但只有选定的段参与损失计算。在SWE-Bench Pro和SWE-Bench Verified上的实验表明,使用SWE-Prime选择的10%轨迹子集进行训练优于使用完整已解决数据集进行训练,分别获得高达12.2%和24.2%的相对性能提升。

原文摘要

To improve large language models’ ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories. However, task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective, redundant, or risky steps. Directly using such trajectories for SFT can introduce noisy supervision and encourage models to imitate undesirable problem-solving behaviors. Therefore, we propose SWE-Prime, a multi-granularity, two-stage SFT data selection method that progressively filters training data at the trajectory and segment levels. Specifically, the first stage performs trajectory-level screening based on process quality, result quality, …

自动采集于 2026-08-29

#论文 #arXiv #NLP #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1