Search Results - RepositoryStats

263

2.4k

apache-2.0

32

SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime

awq fp4 gptq int4 int8 pruning mxformat sparsity sparsegpt auto-tuning smoothquant quantization low-precision large-language-models knowledge-distillation post-training-quantization quantization-aware-training

Created 2020-07-21

3,730 commits to master branch, last one 3 days ago

llmc ModelTC

50

436

apache-2.0

10

[EMNLP 2024 Industry Track] This is the official PyTorch implementation of "LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit".

Created 2024-03-06

472 commits to main branch, last one 11 days ago

auto-round intel

30

397

apache-2.0

11

Advanced Quantization Algorithm for LLMs/VLMs.

awq gptq int4 rounding quantization neural-compressor

Created 2024-01-04

424 commits to main branch, last one 4 days ago