KITE:用稀疏旗舰校准扩展 Jev 人群实验
原标题:KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration
AI 摘要
研究者提出 KITE 架构,用类型化行为内核(TypeSafe 的 Jev,jev-1.13.0)对每个唯一状态查询一次,再以表格化执行任意规模人群实验,并用稀疏旗舰模型锚点估计干预效应。在 Epstein 的 9,070 名参与者实验中,覆盖 1.7% 状态的锚点使效应误差降低 41%;在 37 个 SocSci210 实验中,0.5–1.5% 锚点覆盖率将捕获决策增益从 0.27 提升至 0.39。该方法将人机差异作为共享误差传播,使不确定性由关于人的证据而非蒙特卡洛噪声决定,可用于人体试验前的干预筛选与多国内容审计。
正文节选
KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration Abstract KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human–model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states a