ggml-org/llama.cpp
llama.cpp
llama.cpp runs large language models locally, from command-line inference to an OpenAI-compatible API server.
Catalog projects marked with #inference. Tags work as dedicated landing pages, so related tools are easier to find and connect.
This collection holds 4 projects with a combined 296,600 GitHub stars. Main languages: Python, C++.
llama.cpp runs large language models locally, from command-line inference to an OpenAI-compatible API server.
vLLM is a high-performance engine for LLM inference and serving with an OpenAI-compatible API, batching, and efficient memory management.
Llama is Meta’s repository with inference code for earlier Llama models and links to newer family repositories.