2 results found Sort:

257
2.2k
apache-2.0
33
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created 2020-07-21
3,619 commits to master branch, last one a day ago
74
265
mit
12
Low Precision Arithmetic Simulation in PyTorch
Created 2018-07-30
246 commits to master branch, last one 6 months ago