3 results found Sort:

258
2.3k
apache-2.0
33
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created 2020-07-21
3,709 commits to master branch, last one 4 days ago
43
375
apache-2.0
10
[EMNLP 2024 Industry Track] This is the official PyTorch implementation of "LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit".
Created 2024-03-06
432 commits to main branch, last one 23 hours ago
25
297
apache-2.0
12
Advanced Quantization Algorithm for LLMs/VLMs.
Created 2024-01-04
380 commits to main branch, last one 18 hours ago