德国联邦宪法法院句子级解释准则分类基准
原标题:Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court
AI 摘要
该论文为大型语言模型(LLM)的司法推理分析构建了一个句子级基准,用于分类德国联邦宪法法院判决中的解释准则。作者将拉伦茨(Larenz)的解释理论操作化为分类标准,提供了句子级标注的数据集,并对来自三个模型家族的四个LLM进行了基线评估。结果显示,七个二元子任务的平均得分在70.4至79.2之间,语法解释通常最易识别,系统解释通常最难;在测试配置下,GEPA优化提示并未系统性地优于专家手写提示。
正文节选
Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court Abstract Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Sec