论文概要
研究领域: CV 作者: Weiliang Chen, Haowen Sun, Jun Gao et al. (43 authors) 发布时间: 2026-08-17 arXiv: 2608.16859
中文摘要
基准测试应提供的不仅仅是标量分数:使评估可信的是证明该分数的推理。这对世界模型尤为关键,判断一次展开需要理解物理、因果性和世界状态是否正确演化。人类能自然地发现此类违反,但没有现有基准能自动化这一能力:指标是暴力计算的,没有可检查或验证的推理链。我们引入HarnessEval-W,一个智能体化的评估管道,将LLM生态系统的harness范式带入世界模型基准测试。不是应用固定评分标准,HarnessEval-W解释每个评估案例的上下文,将评估问题分解为可测量的子问题,并生成专门的子智能体,每个配备定制的上下文和诊断工具来推理其自己的子问题。父智能体然后验证收集的证据并将其总结为最终裁决。这种层次化工作流将每次评估转化为透明的证据树,其完整推理链证明了结果。我们将完整管道开源为实时基准测试。
原文摘要
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns spec…
— 自动采集于 2026-08-19
#论文 #arXiv #CV #小凯
