8 results found Sort:
- Filter by Primary Language:
- Python (5)
- C++ (3)
- +
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created
2020-07-21
3,610 commits to master branch, last one a day ago
An innovative library for efficient LLM inference via low-bit quantization
This repository has been archived
(exclude archived)
Created
2023-11-20
345 commits to main branch, last one 2 months ago
tensorrt int8 量化yolov5 onnx模型
Created
2021-01-31
4 commits to master branch, last one 3 years ago
TensorRT int8 量化部署 yolov5s 模型,实测3.3ms一帧!
Created
2021-01-31
3 commits to master branch, last one 3 years ago
RepVGG TensorRT int8 量化,实测推理不到1ms一帧!
Created
2021-02-04
8 commits to master branch, last one 3 years ago
a simple pipline of int8 quantization based on tensorrt.
Created
2022-08-22
23 commits to main branch, last one 2 years ago
👀 Apply YOLOv8 exported with ONNX or TensorRT(FP16, INT8) to the Real-time camera
Created
2024-05-14
73 commits to main branch, last one 5 months ago
nanodet int8 量化,实测推理2ms一帧!
Created
2021-02-09
4 commits to master branch, last one 3 years ago