Vibe Patenting:评估LLM裁判在专业专利撰写智能体中的表现
原标题:Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
AI 摘要
该研究提出 Vibe Patenting 测试平台,用于评估 LLM 裁判在专业专利撰写任务中的可靠性。研究发现,裁判引导的迭代修订能持续提升裁判评估的草稿质量,而无引导修订趋于饱和;迭代反馈还能让低推理能力智能体接近高推理能力智能体的表现。但将 LLM 裁判与专业专利律师的独立评估对比后发现,两者一致性高度依赖具体指标,存在系统性校准差异,说明自动裁判下的改进未必等同于专业专家判断下的改进。
正文节选
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents Abstract LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and dra