BeTaL-GBI:自我审计的几何信念接口验证框架
原标题:BeTaL-GBI: Admission-Aware Benchmark Tuning and Full-Stack Verification of Geometric Belief Interfaces
AI 摘要
BeTaL-GBI 研究通过三个版本迭代,构建了可自我审计的几何信念接口验证框架。v0.2 通过基准调优将平均目标差距从 13.61% 降至 2.87%;v2 引入参考无关的见证状态,检测出全部 116 个严重矛盾;v3 将 99 条声明映射为可执行证据类,并发现原始论文中的 Fisher 条件数值错误,将其标记为勘误而非静默修正。该框架强调模型输出仅为证据而非权威,但明确不声称具备临床安全性或生产就绪性。
正文节选
BeTaL-GBI: Admission-Aware Benchmark Tuning and Full-Stack Verification of Geometric Belief Interfaces A Companion Empirical Study of BoundaryBench v0.2, GBI v2, and GBI-DCSE v3 Abstract A verification substrate is more credible when it can expose errors in claims about itself, not only errors in model output. In extending GBI-DCSE into executable claim-level verification, the v3 harness falsified a numerical implication in the original architecture manuscript: the previously reported Fisher-con