Breaking FP64 Limits: AdaptiveGEMM Achieves Near-A100 Performance on RTX 4060 using INT8 Tensor Cores
- Time
- 2026-08-08 12:00 ~ 12:30
- Speaker
- Tsai,Ming-Han
- Room
- TR212
- Co-write
Abstract
High-precision GEMM is vital for scientific computing and AI. However, consumer GPUs (like the RTX 40 series) are hindered by hardware limitations in FP64 performance.
This session introduces "AdaptiveGEMM," an open-source project employing the Ozaki Scheme to algorithmically decompose high-precision operations. We will dive into CUDA architecture, utilizing PTX and mma.sync instructions to maximize INT8 Tensor Core utilization while resolving Shared Memory Bank Conflicts. Discover how this project bypasses hardware restrictions to achieve near A100-level FLOPS on an RTX 4060.
Designed for advanced HPC engineers, this talk requires basic C++ proficiency and an understanding of GPU memory hierarchy.
Speaker
Tsai,Ming-Han
是個對複雜的事物都很喜歡且很有自己想法的一個人。
ps 我在找工作 有需要都歡迎聯繫! [email protected]