返回全部动态

smevals:小型评估套件,用于模型、提示词和测试框架评估

原标题:smevals - a small eval suite for evaluating models, prompts, and harnesses

Simon Willison's Weblog开源质量 75

AI 摘要

Simon Willison 与 Jesse Vincent 的 Prime Radiant 实验室合作开发了 smevals,这是一个用于评估模型、提示词和测试框架的小型评估套件。该工具允许用户通过 YAML 文件定义评估任务,运行模型测试,并独立进行评分,最终生成静态 HTML 报告。smevals 是 Willison 对评估方法的第三次迭代,旨在帮助回答关于不同模型能力的问题。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

31st July 2026 - Link Blog smevals - a small eval suite for evaluating models, prompts, and harnesses. I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models. The result is smevals, a new tool for running small eval suites across different model configurations and grading the results. The blog entry describes the tool in detail. Here's the 10 second version: - Tell your cod


发布时间:2026-08-01 05:15
抓取时间:2026-08-02 00:23
来源机构:Simon Willison
阅读原文simonwillison.net