GazeAnywhere:基于概念提示的任意场景凝视目标估计
原标题:Gaze Target Estimation Anywhere with Concepts
AI 摘要
研究人员提出了一种新的可提示凝视目标估计(PGE)任务,通过文本或视觉提示识别主体并预测凝视目标,无需多阶段流水线。他们开发了Gaze-Co数据集(包含12万张带提示标注的图像对)和GazeAnywhere模型,该模型使用基于Transformer的检测器融合冻结编码器特征,同时解决主体定位、帧内/外存在性和凝视目标热图估计。GazeAnywhere在多个PGE基准上达到最先进性能,并已在GitHub上开源,论文被CVPR 2026接收。
正文节选
Gaze Target Estimation Anywhere with Concepts Abstract A new promptable paradigm integrates subject localization and gaze estimation into an end-to-end transformer model that uses text or visual prompts to identify subjects and predict gaze targets without multi-stage pipelines. Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and hu