13 results found Sort:
- Filter by Primary Language:
- Python (10)
- Jupyter Notebook (3)
- +
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created
2020-07-21
3,619 commits to master branch, last one a day ago
micronet, a model compression and deploy lib. compression: 1、quantization: quantization-aware-training(QAT), High-Bit(>2b)(DoReFa/Quantization and Training of Neural Networks for Efficient Integer-Ari...
bnn
twn
onnx
dorefa
pruning
pytorch
tensorrt
xnor-net
quantization
network-slimming
group-convolution
model-compression
network-in-network
tensorrt-int8-python
convolutional-networks
neuromorphic-computing
integer-arithmetic-only
batch-normalization-fuse
post-training-quantization
quantization-aware-training
Created
2019-12-04
295 commits to master branch, last one 3 years ago
TinyNeuralNetwork is an efficient and easy-to-use deep learning model compression framework.
Created
2021-11-02
816 commits to main branch, last one 28 days ago
[ICML 2024] SqueezeLLM: Dense-and-Sparse Quantization
Created
2023-06-12
50 commits to main branch, last one 8 months ago
[ICCV 2023] Q-Diffusion: Quantizing Diffusion Models.
Created
2023-03-24
12 commits to master branch, last one 8 months ago
A model compression and acceleration toolbox based on pytorch.
Created
2022-07-21
134 commits to main branch, last one about a year ago
[EMNLP 2024 Industry Track] This is the official PyTorch implementation of "LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit".
Created
2024-03-06
365 commits to main branch, last one 22 hours ago
[IJCAI 2022] FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer
Created
2021-11-24
20 commits to main branch, last one about a year ago
This repository contains notebooks that show the usage of TensorFlow Lite for quantizing deep neural networks.
Created
2020-04-29
143 commits to master branch, last one about a year ago
[NeurIPS 2024 Oral🔥] DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs.
Created
2024-05-25
5 commits to main branch, last one about a month ago
Notes on quantization in neural networks
Created
2023-11-24
15 commits to main branch, last one 11 months ago
[CVPR 2024 Highlight] This is the official PyTorch implementation of "TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models".
Created
2024-03-09
32 commits to main branch, last one 3 months ago
[ICLR 2024] This is the official PyTorch implementation of "QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models"
Created
2024-02-21
9 commits to main branch, last one 8 months ago