OmniAssistBench:全模态LLM助手式交互基准测试
原标题:Paper page - OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
AI 摘要
OmniAssistBench是一个用于评估全模态大语言模型(Omni-LLMs)作为实时视频助手能力的基准测试。该基准通过逆向工程现有互联网视频构建多轮交互数据集,要求模型根据预定义路径引导用户。测试结果显示,专有模型Gemini-3-Pro得分66.4,开源模型Qwen3-Omni-Instruct得分51.2,当前模型普遍存在视觉提示理解、上下文保持和响应时机把握方面的不足。该数据集构建耗时超过1000个专家工时,表明全模态模型在成为可靠助手前仍有较大改进空间。
正文节选
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs Abstract OmniAssistBench evaluates real-time interactive video assistants by reverse-engineering multi-turn interaction videos, revealing that current omni-modal models struggle with visual prompts, context retention, and timely responses. Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unl