Local coding

Component map (what you need)

Local coding is five layers. Missing any required layer = “nothing works.” Optional RAG only if you need repo retrieval beyond the model context.

  1. 1

    1. IDE / Agent client

    VS-Code-/Cursor-Erweiterungen (Continue, Cline), JetBrains-Plugins (Continue und ähnliche) sowie CLIs wie Aider. Alle auf einen lokalen OpenAI-kompatiblen Endpunkt.

    Der Wirt ist IDE oder Terminal — das Plugin ist nicht das Modell. Repos auf Agent / Toolchain.

  2. 2 Optional

    2. RAG / repo index

    Die meisten Clients bringen einen Index mit (Continue/Cline-Codebase, Aider-Repo-Map). Eigene Schicht: AnythingLLM, RAGFlow oder LlamaIndex; Stores Chroma, Qdrant oder FAISS; Embeddings lokal.

    Weglassen, wenn kurzer Complete reicht; bei kleinem Modell und großem Monorepo fast Pflicht. Das ist nicht die Inferenz-Runtime.

  3. 3

    3. Local runtime (OpenAI-compatible server)

    Ollama, llama.cpp server, MLX, vLLM, or SGLang—exposes localhost chat/completions.

    This is the “stack” decision on this page. Community stars below measure popularity of this layer.

  4. 4

    4. Model weights on disk

    GGUF / MLX / safetensors from the Models allowlist—size locks your hardware band.

    API-only MoE coders never land here.

  5. 5

    5. Hardware (VRAM / unified memory)

    NVIDIA CUDA, Apple unified + Metal/MLX, or limited AMD/Intel paths—see Hardware.

    Weights fit ≠ context fit (KV tax).

Quick “what am I missing?”

  • Have an IDE but no localhost runtime → install Ollama / llama.cpp / MLX first.
  • Runtime up but empty models → download an allowlisted coder GGUF (Models page).
  • Model loads but agent is blind in a big repo → add RAG / repo index (optional layer).
  • OOM or tiny context → Hardware band is wrong, or KV budget is too high.

Decision tree

Desktop NVIDIA

Ollama to start → llama.cpp when you need exact quants and context → vLLM / SGLang only on a dedicated machine.

Apple Silicon

Prefer MLX for native unified memory. Use llama.cpp or Ollama when you need cross-platform GGUF workflows.

IDE wiring

Zeigen Sie eine VS-Code-/Cursor-Erweiterung (Continue, Cline), ein JetBrains-Plugin (Continue u. ä.) oder Aider auf localhost. Offene Repos auf Agent / Toolchain — diese Seite erfindet kein zweites Ranking.

Agent · Toolchain

Community signal — runtime popularity (GitHub OSS)

Stars from Trends segments inference / serving / llm_lib. This is popularity of layer 3 (local runtime)—not SWE skill, not IDE quality, not model quality.

Use it to sanity-check “is this runtime still alive?”, not to pick a coding model.

Repository Trends segment Stars
ggml-org/llama.cpp inference 130,674
earendil-works/pi inference 113,430
BerriAI/litellm inference 60,357
sgl-project/sglang inference 36,869
microsoft/onnxruntime inference 22,037
huggingface/transformers llm_lib 167,053
fighting41love/funNLP llm_lib 83,755
unslothai/unsloth llm_lib 77,463
ComposioHQ/awesome-claude-skills llm_lib 76,698
asgeirtj/system_prompts_leaks llm_lib 69,122
open-webui/open-webui serving 154,212
Mintplex-Labs/anything-llm serving 66,820
janhq/jan serving 44,852
lm-sys/FastChat serving 39,549
chatchat-space/Langchain-Chatchat serving 38,675

← Local coding hub

Assumptions and sources

Hardware bands are classes, not street quotes. VRAM estimates use curated GGUF footprints plus size-class KV factors. API job cost uses OpenRouter primary listing when present. choose-ai-infrastructure-guide covers API hosts—not desk GPUs.