Evaluating Tool-Augmented Agents in Remote Sensing Platforms
Simranjit Singh, Michael Fore, Dimitrios Stamoulis CoStrategist R、D Group, Microsoft Corporation, Redmond, WA, USA {simsingh, mifore,
Abstract
Tool-augmented Large Language Models (LLMs) have shown impressive capabilities in remote sensing (RS) applications. However, existing benchmarks assume question-answering input templates over predefined image-text data pairs. These standalone instructions neglect the intricacies of realistic user-grounded tasks. Consider a geospatial analyst: they zoom in a map area, they draw a region over which to collect satellite imagery, and they succinctly ask “Detect all objects here”. Where is here, if it is not explicitly hardcoded in the image-text template, but instead is implied by the system state, e.g., the live map positioning? To bridge this gap, we present GeoLLM-QA, a benchmark designed to capture long sequences of verbal, visual, and click-based actions on a real UI platform. Through in-depth evaluation of state-of-the-art LLMs over a diverse set of 1,000 tasks, we offer insights towards stronger agents for RS applications.
中文速览
现有的遥感大模型评测基准大多只考虑固定的"图片+文字"问答对,根本无法反映真实用户在地图平台上缩放、框选区域、随口提问等连续交互场景。为此,研究团队构建了GeoLLM-QA——一个包含1000个任务的基准,完整记录用户在真实Web地图界面上的语言、视觉与点击操作序列,并配套117个工具及覆盖光学与合成孔径雷达(SAR)影像的三大公开数据集。他们用成功率、函数调用正确率、ROUGE得分和目标检测召回率等多维指标,对GPT-4/3.5结合CoT、ReAct、Chameleon等提示策略进行了系统评测,发现"漏调工具"是最普遍的错误类型,占所有错误的一半以上,且传统文本相似度指标难以准确反映智能体的真实能力。这项工作填补了遥感领域"现实交互式任务评测"的空白,为开发更强大的地理空间智能体提供了可复现的基础平台与基准参考。
原文 arXiv:2405.00709;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2405.00709v1