突破 FP64 限制:AdaptiveGEMM 透過 INT8 Tensor Core 在消費級 GPU 逼近 A100 效能
- 時間
- 2026-08-08 12:00 ~ 12:30
- 講者
- Tsai,Ming-Han
- 位置
- TR212
議程簡介
高精度矩陣乘法 (GEMM) 在科學與 AI 運算中是核心算子。然而,消費級 GPU (如 RTX 40 系列) 的 FP64 雙精度算力受硬體限制,成為許多開發者與量化交易的效能瓶頸。
本議程將分享開源專案 AdaptiveGEMM:探討如何利用實作 Ozaki Scheme 將高精度運算降維拆解。我們將深入底層,利用 CUDA PTX 與 mma.sync 指令極限壓榨 INT8 Tensor Core,解決 Shared Memory Bank Conflict,成功在 RTX 4060 上大幅提升效能,達到逼近 A100 等級的 FLOPS 表現。
【難易度:進階】適合尋求突破硬體極限的 HPC 與 AI 底層架構工程師。 【先備知識】建議具備基礎 C++ 能力,了解 GEMM 運作原理,並對 GPU 記憶體架構與 Tensor Core 有初步認識。 Breaking the FP64 Bottleneck: AdaptiveGEMM with INT8 Tensor Cores on Consumer GPUs
FP64 performance on consumer GPUs is... not exactly great.
In this talk, I’ll share AdaptiveGEMM, a small experimental project inspired by prior work on the Ozaki Scheme. The idea is to break high-precision matrix multiplication into lower-precision operations and make use of the much faster INT8 Tensor Cores available on consumer GPUs.
Rather than focusing on the algorithm itself, this talk is mostly about the engineering journey: turning ideas from papers into CUDA code, working with PTX and mma.sync, dealing with shared-memory bank conflicts, and figuring out why something that looks fast on paper is not always fast on the GPU.
Most of the experiments were done on an RTX 4060.
If you enjoy CUDA, GPU optimization, HPC, or simply making hardware do things it was not particularly designed to do, this talk might be fun for you.
講者
Tsai,Ming-Han
是個對複雜的事物都很喜歡且很有自己想法的一個人。
ps 我在找工作 有需要都歡迎聯繫! [email protected] I’m an engineer who enjoys digging into complicated systems, especially when performance, hardware, and low-level software are involved. my linkedin: https://www.linkedin.com/in/aloha1357/