返回全部动态

OmniAssistBench:全模态LLM助手式交互基准测试

原标题:Paper page - OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

OmniAssistBench是一个用于评估全模态大语言模型(Omni-LLMs)作为实时视频助手能力的基准测试。该基准通过逆向工程现有互联网视频构建多轮交互数据集,要求模型根据预定义路径引导用户。测试结果显示,专有模型Gemini-3-Pro得分66.4,开源模型Qwen3-Omni-Instruct得分51.2,当前模型普遍存在视觉提示理解、上下文保持和响应时机把握方面的不足。该数据集构建耗时超过1000个专家工时,表明全模态模型在成为可靠助手前仍有较大改进空间。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs Abstract OmniAssistBench evaluates real-time interactive video assistants by reverse-engineering multi-turn interaction videos, revealing that current omni-modal models struggle with visual prompts, context retention, and timely responses. Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unl


发布时间:—
抓取时间:2026-08-24 12:28
来源机构:Hugging Face
阅读原文huggingface.co