论文概要
研究领域: ML 作者: Zhuoyi Yang, Ian G. Harris, Salar Hashemitaheri, Cassie Huang, Yuangang Li, Hyunwoo Oh, Paul Dourish, Tony Givargis, Mohsen Imani, Li Zhang 发布时间: 2026-08-21 arXiv: 2608.21345
中文摘要
自精炼通常结构化为生成、批评和修订,是改进LLM生成的广泛采用的范式,也是许多LLM代理的核心机制。虽然三个阶段涉及不同的认知需求,但大多数现有方法方便地将模型大小视为实现细节而非研究对象,这可能导致资源浪费。很少有工作系统研究模型大小如何影响每个阶段,或者有效的自精炼是否需要生成、批评和修订同等能力的模型。我们首次在自精炼流程上进行了分阶段模型大小研究,使用6种大小的Qwen3和4种大小的Gemma 3,在来自不同领域的5个基准上进行。我们得出结论:更大的生成器和精炼器通常能改进流程,而过小的精炼器甚至可能损害性能。其次,性能对批评器的大小高度不敏感,尽管即使包含一个小的批评器也始终优于完全省略批评。我们的发现表明,模型容量不应在自精炼流程中均匀分配。相反,不同阶段表现出不同的大小缩放特性,为设计计算更高效的多阶段语言模型系统提供了实用指导。
原文摘要
Self-refinement, typically structured as generation, critique, and revision, is a widely adopted paradigm for improving LLM generation and serves as a core mechanism in many LLM agents. While the three stages involve different cognitive demands, most existing approaches conveniently treat the model size as an implementation detail rather than a subject of study, which may lead to a waste of resources. Little work has systematically examined how model size affects each stage or whether effective self-refinement requires equally capable models for generation, critique, and revision. We present the first stage-wise model size study of the self-refinement pipeline on 5 benchmarks from different domains using 6 model sizes of Qwen3 and 4 model sizes of Gemma 3. We conclude that larger generator…
— 自动采集于 2026-08-25
#论文 #arXiv #ML #小凯
