9 results found Sort:

263
2.4k
apache-2.0
32
SOTA low-bit LLM quantization (INT8/FP8/INT4/FP4/NF4) & sparsity; leading model compression techniques on TensorFlow, PyTorch, and ONNX Runtime
Created 2020-07-21
3,732 commits to master branch, last one a day ago
Large Language Models for All, 🦙 Cult and More, Stay in touch !
Created 2023-03-30
27 commits to main branch, last one about a year ago
31
407
apache-2.0
11
Advanced Quantization Algorithm for LLMs/VLMs.
Created 2024-01-04
424 commits to main branch, last one 6 days ago
21
400
apache-2.0
4
🦖 X—LLM: Cutting Edge & Easy LLM Finetuning
Created 2023-11-10
62 commits to main branch, last one about a year ago
54
384
apache-2.0
4
Production ready LLM model compression/quantization toolkit with hw accelerated inference support for both cpu/gpu via HF, vLLM, and SGLang.
Created 2024-06-17
2,112 commits to main branch, last one a day ago
27
168
mit
4
Run any Large Language Model behind a unified API
Created 2023-04-02
31 commits to main branch, last one about a year ago
🪶 Lightweight OpenAI drop-in replacement for Kubernetes
This repository has been archived (exclude archived)
Created 2023-05-23
154 commits to main branch, last one about a year ago
A guide about how to use GPTQ models with langchain
Created 2023-05-11
8 commits to main branch, last one about a year ago
zero零训练llm调参
Created 2023-05-17
40 commits to main branch, last one about a year ago