2 results found Sort:

248
2.1k
apache-2.0
34
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created 2020-07-21
3,480 commits to master branch, last one a day ago
33
303
apache-2.0
8
An innovative library for efficient LLM inference via low-bit quantization
Created 2023-11-20
344 commits to main branch, last one 4 days ago