可解释的 MEG 语音解码:皮层源与驱动检索的刺激特征
原标题:Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
AI 摘要
Hugging Face 每日论文展示了一项关于可解释的 MEG 语音解码研究。研究者重新设计了 MEG 到音频检索架构,将解码器参数减少约 20 倍,同时保持与最先进系统相当的性能(Top-1 准确率 39.75%)。通过将空间注意力替换为球谐函数、减少分支并添加时间滤波,模型权重可映射到皮层源,揭示语音感知网络,并发现左侧分支携带更高频的节律成分。配对 MEG 遮挡实验表明,静音、声音强度、元音和声学起始等 15 个刺激特征对检索有贡献,而随机词列表的替换实验显示叙事结构对可恢复信息的重要性。
正文节选
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval Abstract Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. We build on a high-performing MEG-to-audio ret