1 project found for "quantization"
High-performance C/C++ implementation for running LLM inference locally on consumer hardware, including CPUs, with minimal dependencies.