返回全部动态

GazeAnywhere:基于概念提示的任意场景凝视目标估计

原标题:Gaze Target Estimation Anywhere with Concepts

Hugging Face Daily Papers一手来源研究质量 81

AI 摘要

研究人员提出了一种新的可提示凝视目标估计(PGE)任务,通过文本或视觉提示识别主体并预测凝视目标,无需多阶段流水线。他们开发了Gaze-Co数据集(包含12万张带提示标注的图像对)和GazeAnywhere模型,该模型使用基于Transformer的检测器融合冻结编码器特征,同时解决主体定位、帧内/外存在性和凝视目标热图估计。GazeAnywhere在多个PGE基准上达到最先进性能,并已在GitHub上开源,论文被CVPR 2026接收。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Gaze Target Estimation Anywhere with Concepts Abstract A new promptable paradigm integrates subject localization and gaze estimation into an end-to-end transformer model that uses text or visual prompts to identify subjects and predict gaze targets without multi-stage pipelines. Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and hu


发布时间:—
抓取时间:2026-08-14 05:08
来源机构:Hugging Face
阅读原文huggingface.co