消费级GPU本地LLM能源效率基准:架构与量化策略影响显著
原标题:Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware
AI 摘要
一项针对消费级GPU上本地部署LLM的能源效率基准研究,在RTX 4060Ti上测试了9个开源模型(1B-7B参数),发现模型架构和量化策略比参数数量更影响能效,其中gemma3:1b和llama3.2:1b能耗最低,而7B-Mistral能耗最高达4.4倍,qwen3.5:2b因内部推理导致每提示能耗异常高。
正文节选
Computer Science > Artificial Intelligence Title:Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware View PDF HTML (experimental) Abstract:The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. This paper presents a reproduc