Lokales Coding-Modell: Stack, Gewichte, VRAM
Der Stack ist billig zu wechseln; Modellgröße sperrt Hardware. Passende Gewichte heißen nicht passender Kontext.
Reihenfolge: Stack → Modell → Hardware
Hardware ist teuer, der Text geht vom Billigen zum Teuren. Wer schon Mac / 24GB hat, nutzt den Assistenten rückwärts. Hub: [lokales Coding](/de/local-coding/).
| Platform | Stack |
|---|---|
| Desktop NVIDIA | Ollama → llama.cpp → vLLM/SGLang (dedicated box) |
| Apple Silicon | MLX first; llama.cpp / Ollama for GGUF portability |
| Laptop | Never start with vLLM |
| IDE | Point Continue / Aider / Cline at localhost |
Kein vLLM auf dem Laptop. Details: [Stack](/de/local-coding/stack/).
Nur herunterladbare Coding-Gewichte, die in Consumer-Bänder passen
Claude/GPT nicht per Namensheuristik lokal machen. Allowlist: offene Lizenz, Download, Q4 für den Schreibtisch. 480B/284B-MoE = API oder Rack. Das Komposit ist kein SWE. [Modelle](/de/local-coding/models/). [Qwen3.8-27B](/de/insights/qwen3-8-27b-review/).
VRAM-Bänder: die Sperre ist Speicher, nicht TFLOPS
Bild vergrößern | Band | Typical class | Max params (Q4 class) | Workloads |
|---|---|---|---|
| 16GB | RTX 4060 / 5060 Ti 16GB class | 8B Q4 class | complete |
| 24GB | RTX 4090 24GB | 32B Q4 class | complete, agent |
| 32GB | RTX 5090 32GB class | 40B Q4 class | complete, agent, long |
| 48GB | 48GB workstation / dual-24 class | 70B Q4 class | complete, agent, long |
| 80GB | 80GB-class pro card | 120B Q4 class | complete, agent, long |
| 128GB | Mac Studio 128–192GB | 180B Q4 class | complete, agent, long |
| 512GB | Mac Studio 512GB | 480B Q4 class | complete, agent, long |
Passende Gewichte heißen nicht passender Kontext. Details: [Hardware](/de/local-coding/hardware/).
TCO: Token-Sticker null, Strom und Abschreibung nicht
Bild vergrößern Kein 70B auf 16GB, kein 480B auf 4090, kein H100 für Autocomplete. Der Infra-Guide betrifft API-Hosts. Assistent: [Hub](/de/local-coding/).