BudgetBench:本地LLM智能体记忆策略的预算分层评估协议
原标题:BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
AI 摘要
研究者提出 BudgetBench,一个以每次调用输入 token 预算为自变量的记忆策略评估协议与参考工具,在 2K 至 32K 预算下测量质量、预算利用率、延迟和预算违规率。初步实验使用本地 qwen2.5:1.5b、托管的 Qwen3 30B-A3B 以及 LongMemEval 研究,暴露了预算合规失败和非单调质量曲线,但预算与全上下文方向仍未定论。该协议、工具和失败报告规范已开源,旨在推动固定预算记忆策略评估成为社区基准。
正文节选
BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents Abstract For local large language model agents, active context is a scarce operational resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an active-budget evaluation protocol and reference harness that treats the per-call input-token budget as the independent vari