mmMind:姿态引导的毫米波雷达行为理解模型
原标题:Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding
AI 摘要
arXiv 论文提出 mmMind,一种雷达-语言模型,利用同步 3D 姿态作为训练监督,使基础模型能理解毫米波雷达数据。该模型通过姿态引导预训练学习人体结构和运动,推理时仅需雷达输入,并用于行为描述和时空问答。作者还发布了包含 17.9 小时真实数据的 mmMind-Bench 基准,实验表明 mmMind 在多个任务上优于现有基线。
正文节选
Computer Science > Computer Vision and Pattern Recognition Title:Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding View PDF HTML (experimental) Abstract:Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and contactless sensing modality, but radar observations are difficult to align with language. Existing radar-language methods