返回全部动态

Together AI 内核团队:从 FlashAttention 到 Blackwell 的突破

原标题:Inside the Together AI kernels team

Together AI Blog一手来源研究质量 84

AI 摘要

Together AI 的 kernels 团队在 2022 年通过 FlashAttention 实现了 2-3 倍加速,证明了 GPU 优化仍有巨大潜力。2025 年 3 月,该团队在获得 NVIDIA Blackwell GPU 后的一周内,利用 ThunderKittens 库开发出最快的 FP4/FP8 GEMM 内核,性能比 H100 上的 cuBLAS 快 2 倍。团队采用学术界与工业界协同的模式,与 UCSD、Princeton、Stanford 等合作,推动内核研究。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

The breakthrough came on a holiday weekend. Memorial Day 2022. While most of Silicon Valley was at barbecues, Dan Fu, Tri Dao, and their colleagues were about to prove the AI establishment wrong. The conventional wisdom was settled: transformer attention was already optimized. GPU experts had squeezed every drop of performance from the hardware. There wasn't much left to gain. Then Dan, Tri, and their colleagues published FlashAttention. Andrej Karpathy, then Senior Director of AI at Tesla, twee


发布时间:
抓取时间:2026-08-03 01:12
来源机构:Together AI
阅读原文together.ai