Desktop NVIDIA
Ollama to start → llama.cpp when you need exact quants and context → vLLM / SGLang only on a dedicated machine.
Local coding
Change the runtime cheaply. Use the component map to see IDE vs RAG vs runtime vs weights vs hardware—then match platform.
Local coding is five layers. Missing any required layer = “nothing works.” Optional RAG only if you need repo retrieval beyond the model context.
VS Code / Cursor extensions (Continue, Cline), JetBrains plugins (Continue and similar), plus CLI agents like Aider. Point them at a local OpenAI-compatible endpoint.
The host is the IDE or terminal—the plugin is not the model. Browse Agent / Toolchain boards for repos.
Most clients ship an index (Continue/Cline codebase, Aider repo map). Standalone layer: AnythingLLM, RAGFlow, or LlamaIndex; vector stores Chroma, Qdrant, or FAISS; run embeddings locally.
Skip if short-context complete is enough; required for large monorepos with small local models. This is not the inference runtime.
Ollama, llama.cpp server, MLX, vLLM, or SGLang—exposes localhost chat/completions.
This is the “stack” decision on this page. Community stars below measure popularity of this layer.
GGUF / MLX / safetensors from the Models allowlist—size locks your hardware band.
API-only MoE coders never land here.
NVIDIA CUDA, Apple unified + Metal/MLX, or limited AMD/Intel paths—see Hardware.
Weights fit ≠ context fit (KV tax).
Quick “what am I missing?”
Ollama to start → llama.cpp when you need exact quants and context → vLLM / SGLang only on a dedicated machine.
Prefer MLX for native unified memory. Use llama.cpp or Ollama when you need cross-platform GGUF workflows.
Point a VS Code / Cursor extension (Continue, Cline), a JetBrains plugin (Continue and similar), or Aider at localhost. Browse Agent and Toolchain boards for open repos—this page does not invent a second leaderboard.
Stars from Trends segments inference / serving / llm_lib. This is popularity of layer 3 (local runtime)—not SWE skill, not IDE quality, not model quality.
Use it to sanity-check “is this runtime still alive?”, not to pick a coding model.
| Repository | Trends segment | Stars |
|---|---|---|
| ggml-org/llama.cpp | inference | 124,446 |
| earendil-works/pi | inference | 92,672 |
| BerriAI/litellm | inference | 56,603 |
| sgl-project/sglang | inference | 31,995 |
| microsoft/onnxruntime | inference | 21,399 |
| huggingface/transformers | llm_lib | 164,211 |
| fighting41love/funNLP | llm_lib | 82,521 |
| unslothai/unsloth | llm_lib | 73,350 |
| ComposioHQ/awesome-claude-skills | llm_lib | 72,700 |
| labmlai/annotated_deep_learning_paper_implementations | llm_lib | 67,312 |
| open-webui/open-webui | serving | 149,090 |
| Mintplex-Labs/anything-llm | serving | 64,859 |
| janhq/jan | serving | 44,038 |
| lm-sys/FastChat | serving | 39,513 |
| chatchat-space/Langchain-Chatchat | serving | 38,546 |
Hardware bands are classes with USD reference ranges (not street quotes). VRAM estimates use curated GGUF footprints plus size-class KV factors. Coding scores are hand-curated SWE/Aider/LiveCode/HumanEval references—not harness re-runs. API job cost uses OpenRouter primary listing when present.