返回全部动态

BeTaL-GBI:自我审计的几何信念接口验证框架

原标题:BeTaL-GBI: Admission-Aware Benchmark Tuning and Full-Stack Verification of Geometric Belief Interfaces

arXiv cs.SE一手来源研究质量 82

AI 摘要

BeTaL-GBI 研究通过三个版本迭代,构建了可自我审计的几何信念接口验证框架。v0.2 通过基准调优将平均目标差距从 13.61% 降至 2.87%;v2 引入参考无关的见证状态,检测出全部 116 个严重矛盾;v3 将 99 条声明映射为可执行证据类,并发现原始论文中的 Fisher 条件数值错误,将其标记为勘误而非静默修正。该框架强调模型输出仅为证据而非权威,但明确不声称具备临床安全性或生产就绪性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

BeTaL-GBI: Admission-Aware Benchmark Tuning and Full-Stack Verification of Geometric Belief Interfaces A Companion Empirical Study of BoundaryBench v0.2, GBI v2, and GBI-DCSE v3 Abstract A verification substrate is more credible when it can expose errors in claims about itself, not only errors in model output. In extending GBI-DCSE into executable claim-level verification, the v3 harness falsified a numerical implication in the original architecture manuscript: the previously reported Fisher-con


发布时间:2026-08-25 12:00
抓取时间:2026-08-25 12:20
来源机构:arXiv
阅读原文arxiv.org