GUI-CC:评估 GUI 世界模型作为智能体环境的上下文一致性基准
原标题:GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
AI 摘要
GUI-CC 是一个用于评估 GUI 世界模型作为智能体环境时上下文一致性的基准测试。它包含离线轨迹和在线智能体循环两个轨道,基于 GUIOdyssey 构建了 500 个离线任务和 200 个在线任务。实验表明,当前模型在单步生成上表现合理,但在多步环境中难以保持任务相关上下文,导致智能体陷入看似合理但无效的状态。该基准强调了从单步预测到多步环境模拟评估的转变。
正文节选
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments Abstract GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a key requirement under-tested: generated states must remain contextually consistent when they are repeatedly reused for future interaction. We introduce GUI-CC, a benchmark that evaluates contextual consistency of GUI world mode