4 results found Sort:
- Filter by Primary Language:
- C++ (2)
- Python (2)
- +
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created
2020-07-21
3,610 commits to master branch, last one a day ago
row-major matmul optimization
Created
2018-10-28
158 commits to master branch, last one about a year ago
An innovative library for efficient LLM inference via low-bit quantization
This repository has been archived
(exclude archived)
Created
2023-11-20
345 commits to main branch, last one 2 months ago
Advanced Quantization Algorithm for LLMs. This is official implementation of "Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs"
Created
2024-01-04
305 commits to main branch, last one a day ago