[论文] Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

## 论文概要 **研究领域**: ML **作者**: Perry Dong, Ron Polonsky, ...

论文概要

研究领域: ML 作者: Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Fin 发布时间: 2026-07-29 arXiv: 2607.27203

中文摘要

预训练+微调已成为学习高性能策略的主流方法。在基于价值的强化学习(RL)中,一个自然的问题是:给定一个预训练策略,Q函数是否也应该在离线数据上预训练?传统观点认为应该,但近期研究表明,在线RL中使用随机初始化的Q函数也能获得高性能且可靠的策略。本文系统研究了在预训练基础策略之上微调时,Q函数预训练是否真的有帮助。令人惊讶的是,我们发现朴素的Q函数预训练相比随机初始化往往几乎没有优势。这源于一个根本性的不匹配:预训练阶段学到的Q函数针对的是预训练策略的Q函数,而非在线微调收敛到的Q函数,这一差距即使在离线价值最大化后依然存在。基于此,我们提出IPE(策略集成初始化),一种简单方法:训练多个多样化的策略,用它们的汇总rollout来启动在线RL中的Q函数学习。

原文摘要

Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function. In this paper, we systematically study whether pretraining the Q-function actually helps when fine-tuning on top of a pretrained base policy. We find, surprisingly, that naive Q-function pretraining often provides little benefit over random initialization. We show this stems from a fundamental mismatch: the Q-function…

自动采集于 2026-07-31

#论文 #arXiv #ML #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1