返回全部动态
Xinference 发布 v3.4.0:新增 vLLM 多进程路由与 prefill/decode 分离
原标题:v3.4.0
AI 摘要
Xinference 发布 v3.4.0,新增 vLLM 原生多进程执行器路由、Gemma-4 Transformers 后端批处理支持、prefill/decode 分离、多 worker 副本扩缩容等特性,并加入 MonkeyOCR、dots OCR、Fish Audio S1-mini/S2-Pro、MiniCPM5-2B 等模型支持。该版本还增强了缓存管理、TTS 流式播放与权限控制,并修复了流式工具调用、批处理推理隔离、安全与多设备缓存等问题。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
# What's new in 3.4.0 (2026-09-11) These are the changes in inference v3.4.0. ## New features * feat(vllm): support native multiprocessing executor routing by @m199369309 in https://github.com/xorbitsai/inference/pull/5447 * feat: Add Gemma-4 Transformers backend with batching support by @maoyuehui in https://github.com/xorbitsai/inference/pull/5430 * feat(model): add download-only flow by @Minamiyama in https://github.com/xorbitsai/inference/pull/5463 * feat(model): support dots ocr model by @l
发布时间:2026-09-11 20:40
抓取时间:2026-09-11 21:27
来源机构:Xorbits