4 results found Sort:
📚200+ Tensor/CUDA Cores Kernels, ⚡️flash-attn-mma, ⚡️hgemm with WMMA, MMA and CuTe (98%~100% TFLOPS of cuBLAS/FA2 🎉🎉).
Created
2022-12-17
499 commits to main branch, last one 6 days ago
Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruction.
Created
2023-06-22
1 commits to master branch, last one 5 months ago
Several optimization methods of half-precision general matrix vector multiplication (HGEMV) using CUDA core.
Created
2023-10-09
1 commits to master branch, last one 5 months ago
⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.
Created
2024-11-30
43 commits to main branch, last one 7 days ago