[论文] SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineer…

## 论文概要 **研究领域**: ML **作者**: Xin He, Yanlin Wang, Mingw...

论文概要

研究领域: ML 作者: Xin He, Yanlin Wang, Mingwei Liu 发布时间: 2026-09-06 arXiv: 2509.04275

中文摘要

仓库级软件工程基准显著推进了编码智能体的评估,但现有基准主要衡量生成补丁是否通过功能测试,而忽略了来自代码审查的接受约束(审查约束),这些约束通常影响补丁在现实世界软件开发中是否可接受。我们推出了SWE-Gate,一个明确评估审查约束合规性与功能正确性的仓库级软件工程智能体基准。SWE-Gate从真实拉取请求审查评论中提取审查约束,并围绕这些约束合成仓库级修复实例。每个实例提供独立的功能测试和约束测试,以及不合规和黄金补丁,从而实现问题修复能力与审查约束合规性的明确分离。我们构建了包含303个仓库级修复实例的SWE-Gate,涵盖75个跨多样软件领域的开源Python仓库。在通用编码智能体框架下,使用四个不同能力水平的LLM后端进行的实验揭示了功能成功与完整修复规范下的成功之间存在显著差距:在644个通过功能测试的修复中,221个未能满足提供的审查约束。这些发现表明,仅功能评估高估了智能体满足仓库级修复任务完整要求的能力。复现包(包括代码、数据和实验结果)可在https://github.com/DeepSoftwareAnalytics/SWE-Gate获取。

原文摘要

Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook review-derived acceptance constraints (review constraints) that often influence whether a patch is acceptable in real-world software development. We introduce SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness. SWE-Gate derives review constraints from real pull request review comments and synthesizes repository-level repair instances around these constraints. Each instance provides separate functional and constraint tests, together with non-compliant and gold patches, enabli…

自动采集于 2026-09-07

#论文 #arXiv #ML #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1