Llama-3.2-3B 去审查变体发布,拒绝率大幅降低
原标题:saidutta69/Llama-3.2-3B-Instruct-heretic
AI 摘要
Hugging Face 上发布了 Llama-3.2-3B-Instruct-heretic,这是一个基于 Meta 的 Llama-3.2-3B-Instruct 的去审查变体,通过 Heretic 工具进行方向消融(abliteration)抑制拒绝行为,而非微调。该模型在保持低 KL 散度(0.03)的同时,将拒绝率从 97/100 降至 2/100,并提供了多种 GGUF 量化版本,适用于本地代理、角色扮演等场景。开发者需注意其缺乏安全过滤,部署时需谨慎。
正文节选
--- language: - en - de - fr - it - pt - hi - es - th library_name: transformers pipeline_tag: text-generation tags: - facebook - meta - pytorch - llama - llama-3 - heretic - uncensored - decensored - abliterated license: llama3.2 --- # Llama-3.2-3B-Instruct-heretic <div align="center"> <img src="https://res.cloudinary.com/cmazqjs6/image/upload/racer_is_op_banner_branded_pu7zud.png" alt="RACER IS OP" width="100%"> </div> <br> A decensored variant of [meta-llama/Llama-3.2-3B-Instruct](https