Desktop NVIDIA
Ollama to start → llama.cpp when you need exact quants and context → vLLM / SGLang only on a dedicated machine.
Local coding
Change the runtime cheaply. Use the component map to see IDE vs RAG vs runtime vs weights vs hardware—then match platform.
Local coding is five layers. Missing any required layer = “nothing works.” Optional RAG only if you need repo retrieval beyond the model context.
Расширения VS Code / Cursor (Continue, Cline), плагины JetBrains (Continue и аналоги) и CLI вроде Aider. Все — на локальный OpenAI-совместимый endpoint.
Хост — IDE или терминал; плагин не модель. Репозитории на Agent / Toolchain.
У большинства клиентов уже есть индекс (codebase Continue/Cline, repo map Aider). Отдельный слой: AnythingLLM, RAGFlow или LlamaIndex; хранилища Chroma, Qdrant или FAISS; эмбеддинги локально.
Можно пропустить при коротком complete; почти обязателен с маленькой моделью и большим monorepo. Это не runtime инференса.
Ollama, llama.cpp server, MLX, vLLM, or SGLang—exposes localhost chat/completions.
This is the “stack” decision on this page. Community stars below measure popularity of this layer.
GGUF / MLX / safetensors from the Models allowlist—size locks your hardware band.
API-only MoE coders never land here.
NVIDIA CUDA, Apple unified + Metal/MLX, or limited AMD/Intel paths—see Hardware.
Weights fit ≠ context fit (KV tax).
Quick “what am I missing?”
Ollama to start → llama.cpp when you need exact quants and context → vLLM / SGLang only on a dedicated machine.
Prefer MLX for native unified memory. Use llama.cpp or Ollama when you need cross-platform GGUF workflows.
Направьте расширение VS Code / Cursor (Continue, Cline), плагин JetBrains (Continue и аналоги) или Aider на localhost. Открытые репозитории — на досках Agent / Toolchain; эта страница не делает второй рейтинг.
Stars from Trends segments inference / serving / llm_lib. This is popularity of layer 3 (local runtime)—not SWE skill, not IDE quality, not model quality.
Use it to sanity-check “is this runtime still alive?”, not to pick a coding model.
| Repository | Trends segment | Stars |
|---|---|---|
| ggml-org/llama.cpp | inference | 124,446 |
| earendil-works/pi | inference | 92,672 |
| BerriAI/litellm | inference | 56,603 |
| sgl-project/sglang | inference | 31,995 |
| microsoft/onnxruntime | inference | 21,399 |
| huggingface/transformers | llm_lib | 164,211 |
| fighting41love/funNLP | llm_lib | 82,521 |
| unslothai/unsloth | llm_lib | 73,350 |
| ComposioHQ/awesome-claude-skills | llm_lib | 72,700 |
| labmlai/annotated_deep_learning_paper_implementations | llm_lib | 67,312 |
| open-webui/open-webui | serving | 149,090 |
| Mintplex-Labs/anything-llm | serving | 64,859 |
| janhq/jan | serving | 44,038 |
| lm-sys/FastChat | serving | 39,513 |
| chatchat-space/Langchain-Chatchat | serving | 38,546 |
Hardware bands are classes, not street quotes. VRAM estimates use curated GGUF footprints plus size-class KV factors. API job cost uses OpenRouter primary listing when present. choose-ai-infrastructure-guide covers API hosts—not desk GPUs.