1 result found Sort:

263
2.4k
apache-2.0
32
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created 2020-07-21
3,733 commits to master branch, last one 17 hours ago