4 results found Sort:

258
2.3k
apache-2.0
33
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created 2020-07-21
3,699 commits to master branch, last one 2 days ago
row-major matmul optimization
Created 2018-10-28
158 commits to master branch, last one about a year ago
38
349
apache-2.0
8
An innovative library for efficient LLM inference via low-bit quantization
This repository has been archived (exclude archived)
Created 2023-11-20
345 commits to main branch, last one 3 months ago
25
284
apache-2.0
12
Advanced Quantization Algorithm for LLMs/VLMs. This is official implementation of "Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs"
Created 2024-01-04
372 commits to main branch, last one 5 days ago