返回全部动态

K-Search 将 CUDA 内核专业知识迁移至 Apple Silicon

原标题:From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Berkeley AI Research Blog一手来源研究质量 87

AI 摘要

Berkeley AI Research 博客介绍了 K-Search 框架的扩展,该框架利用 AI 进化搜索优化 GPU 内核,并新增了针对 Apple Silicon 的 MLX 后端。通过结构化的 CUDA 到 MLX 翻译层,K-Search 能将 CUDA 内核知识库自动适配为高质量 MLX 内核,在注意力内核上达到 0.97 倍原生性能,在 Mamba SSM 内核上实现最高 20 倍预填充加速。该方法不限于 MLX,可推广至其他硬件生态。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

We face a new epoch in computing. Hardware is changing rapidly — not just faster GPUs, but a growing range of chips from different vendors, each with its own architecture and often tailored to specific AI workloads. Software is changing just as fast, and AI coding tools now generate in minutes what took months of effort a few years ago. With so much of computing now centered on AI, GPU kernels are a crucial component of its success. These are the low-level programs that run inside the GPU, and w


发布时间:2026-07-29 17:00
抓取时间:2026-08-02 00:24
来源机构:BAIR
阅读原文bair.berkeley.edu