返回全部动态
AgentStream:评估自进化LLM代理在流式任务中的表现
原标题:AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
AI 摘要
AgentStream 是一个统一框架,用于评估自进化 LLM 代理在流式任务中的表现,通过将基准组织为可配置任务流,并实例化隔离、顺序和交错三种流式场景。研究发现自进化可靠性随场景变化,其收益受模型能力限制且非单调,没有单一方法在所有模型和场景中占优。该研究倡导在现实任务流而非孤立单任务设置中评估自进化代理。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Abstract Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified fr
发布时间:—
抓取时间:2026-08-05 09:05
来源机构:Hugging Face