[论文] LESSER: Post-Training Data Selection with Output-Layer Gradients

## 论文概要 **研究领域**: ML **作者**: Lyuxin David Zhang, Eric W...

论文概要

研究领域: ML 作者: Lyuxin David Zhang, Eric Wong, Surbhi Goel, Anton Xue 发布时间: 2026-10-02 arXiv: 2610.03702

中文摘要

大语言模型后训练数据的选择对下游性能有重大影响。基于梯度的数据选择是一种流行方法,通过训练数据梯度与小型验证集梯度的对齐程度来排序。然而,使用全参数梯度排序需要对每个样本进行昂贵的前向和反向传播,使得大规模候选池的计算不可行。这引发一个自然的问题:能否以极低成本近似全梯度特征?幸运的是,我们发现输出层梯度足以实现有效的数据选择,且仅需更便宜的前向传播。我们将其实现为 LESSER——一个即插即用的选择方法包装器,将特征提取的 FLOP 成本在 SFT 上降低 9.7 倍,在 RL 基准上降低 3.0 倍,同时在下游任务上保持与全梯度性能一致。实验发现,即使输出层梯度和全梯度对单个样本的排序不同,它们选择的批次仍具有对齐的梯度。

原文摘要

The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by 9.7times for SFT and 3.0times fo…

— 自动采集于 2026-10-06

#论文 #arXiv #ML #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1