Search open source projects
4 projects found for "llm-inference"
Related category: AI & Machine Learning (4)
vLLM
Apache-2.0High-throughput, memory-efficient inference and serving engine for large language models, built for running LLMs in production at scale rather than on a single local machine.
- AI & Machine Learning
llama.cpp
MITHigh-performance C/C++ implementation for running LLM inference locally on consumer hardware, including CPUs, with minimal dependencies.
- AI & Machine Learning
SGLang
Apache-2.0Fast serving framework for large language models and vision-language models, with a structured generation language for complex LLM programs.
- AI & Machine Learning
KServe
Apache-2.0Kubernetes-native platform for serving machine learning models at scale, standardizing model deployment across frameworks.
- AI & Machine Learning