[论文] Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Re…

## 论文概要 **研究领域**: CV **作者**: Fengxiang Wang, Jiangnan H...

论文概要

研究领域: CV 作者: Fengxiang Wang, Jiangnan Huang, Mingshuo Chen, Yueying Li, Yang Shi, Junwei Luo, Haoyu Wang, Yansheng Li, Jing Zhang, Haiyan Zhao, Wenjing Yang 发布时间: 2026-07-28 arXiv: 2607.25993

中文摘要

超高分辨率(UHR)遥感(RS)影像在城市级场景上提供了细粒度的地球观测证据,但对多模态大型语言模型(MLLM)提出了根本性挑战:任务相关证据通常是稀疏的、局部的,并在极大的视觉上下文中空间分散。一个自然的解决方案是为MLLM配备缩放工具以进行主动局部检查。然而,通过在XLRS-Bench上的试点研究,我们发现缩放仅部分有效:它解决了具有局部可恢复证据的容易和中等难度任务,但在需要全局搜索、多区域比较、路径规划或分散证据推理的困难案例上饱和。受这一发现的启发,我们超越了单一工具缩放,引入了GeoMTVR,一个从广域卫星影像构建的大规模地理空间多工具视觉推理数据集。GeoMTVR包含13K UHR VQA样本,具有交错的推理轨迹、多样化的视觉工具调用和返回的视觉观测,使模型能够学习问题分解、工具选择、区域检查、对象级定位、辅助视觉推理和跨工具证据整合。除了监督微调,我们提出了一种专注于工具注意力的强化学习算法,将优化集中在关键工具使用决策上,包括何时调用工具、选择哪个工具、在哪里应用它以及如何解释工具输出。通过结合GeoMTVR上的SFT和我们的RL算法,我们开发了GeoLens,一个用于UHR RS的多工具视觉推理MLLM。实验表明,GeoLens始终优于直接推理和单一工具缩放基线,实现了更强的准确性、更好的证据定位和更有效的工具使用轨迹。

原文摘要

Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts. A natural solution is to equip MLLMs with zoom-in tools for active local inspection. However, through a pilot study on XLRS-Bench, we find that zoom-in is only partially effective: it resolves easy and medium-level tasks with locally recoverable evidence, but saturates on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning. Motivated by this finding, we move beyond single-tool zoom-in and introduce GeoMTVR, a large-scale…

自动采集于 2026-07-30

#论文 #arXiv #CV #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1