From f5a23b06ac9f62f9ab4618132140cc958cad227c Mon Sep 17 00:00:00 2001 From: Claude Date: Sat, 22 Aug 2026 03:41:25 +0000 Subject: [PATCH] docs(ai-stack): note MoE models as a coding option for Pascal/Volta MoE VRAM tracks total params but compute tracks active params, so a model like qwen3-coder:30b-a3b (30B total, ~3.3B active) runs light enough that old GPUs' weak/missing tensor cores barely matter. --- services/ai-stack.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/services/ai-stack.md b/services/ai-stack.md index b267821..1847092 100644 --- a/services/ai-stack.md +++ b/services/ai-stack.md @@ -82,6 +82,21 @@ architectures the *next* major toolkit will support — existing CUDA 12.x builds keep working, but factor this in before buying used Pascal/Volta hardware today. +**MoE models are the exception that gives Pascal/Volta real life for coding.** +The "coding" column above assumes dense models, where token speed tracks the +full parameter count — exactly where Pascal/Volta's missing or first-gen +tensor cores hurt most. A mixture-of-experts model breaks that link: VRAM is +still set by *total* params (every expert has to be resident — no memory +saving from sparsity), but compute per token is set by *active* params only. +`qwen3-coder:30b-a3b` in `ollama pull` is the concrete case — 30B total, only +~3.3B active per token (128 experts, 8 routed) — so it needs the same ~19GB +VRAM (Q4_K_M) as a dense 30B model but computes like a dense ~3B one. That's +light enough that Pascal/Volta's weak tensor cores barely matter, making it +the best coding model to put on a P40 24GB or a V100 — a dense 32B coder on +the same card would be noticeably slower for no quality gain. Mixtral 8x7B +(46.7B total / ~13B active, ~24-26GB at Q4) is the same trade at a larger +size — fits Volta 32GB or Ampere, with the same active-vs-total gap. + ## Cloud LLM providers (Open WebUI) Open WebUI uses an OpenAI-compatible connection list. The local RAG server is the first entry; any cloud providers added at install follow it. Two semicolon-separated