Reframe guide around Drop Max / Keep Pro strategy

The real use case: drop Claude Max ($100/mo), keep Pro ($20/mo), offload
bulk codebase work to local. Local handles the 80% (reading 10K-line
projects, routine fixes, boilerplate) with no rate limits. Pro handles
the hard 20% where Opus quality matters. GPU pays for itself in 13 months,
saves $1,596 over 3 years vs Max.

Also adds: NVLink speed reality check, image gen capabilities (SDXL/Flux
included at no extra cost), hardware longevity estimate (3-5 years),
and updated cost comparison tables.

https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
This commit is contained in:
Claude
2026-03-22 17:07:01 +00:00
parent 1e7b087a5d
commit 5cfd8b2696
+91 -15
View File
@@ -48,9 +48,34 @@ than traditional transformers.
| **Qwen3-Coder-Next 80B** (MoE) | needs 64GB+ RAM offload | Beats Claude Opus 4.6 on SWE-rebench | Hybrid attention, 256K context | | **Qwen3-Coder-Next 80B** (MoE) | needs 64GB+ RAM offload | Beats Claude Opus 4.6 on SWE-rebench | Hybrid attention, 256K context |
### Honest Assessment: Local vs Claude Code ### Honest Assessment: Local vs Claude Code
Nothing local approaches Claude Opus 4.6 quality for complex multi-file agentic coding. Nothing local approaches Claude Opus 4.6 quality for complex multi-file agentic coding.
These 32B models are competitive with **GPT-4o** — a tier below Claude Sonnet, two tiers below Opus. These 32B models are competitive with **GPT-4o** — a tier below Claude Sonnet, two tiers below Opus.
Best strategy: use local models for routine tasks, save Claude credits for hard problems.
**The real strategy: Drop Max ($100/mo), keep Pro ($20/mo), offload bulk work to local.**
The problem with Pro for large projects: rate limits. A 10,000-line codebase needs the model
to read, understand, and hold context across many files. On Pro you'll hit usage caps mid-session
on complex multi-file work. Max ($100/mo) removes those limits — but that's $80/mo extra.
Local AI eliminates this problem differently:
- **Local model (262K context)**: Reads your entire 10K-line project at once. No rate limits,
no usage caps, runs 24/7. Handles the bulk work — understanding codebase structure, routine
bug fixes, simple refactors, code explanation, test writing, boilerplate generation.
- **Claude Pro ($20/mo)**: Reserved for the hard problems — complex multi-file architectural
changes, subtle bugs that need Opus-level reasoning, code review on critical paths.
Pro limits are fine when you're only sending Claude the *hard* 20% instead of everything.
This is the unlock: local doesn't replace Claude, it **reduces your Claude usage enough
that Pro limits stop being a problem.** The 80% of routine work that was burning through
your Max quota now runs locally with zero limits.
| Plan | Monthly | What You Get | Limit Problem |
|------|---------|--------------|---------------|
| Max only | $100 | Opus unlimited | Paying $80/mo for unlimited when you don't need it |
| Pro only | $20 | Opus with rate limits | **Hits caps on 10K-line projects** |
| **Pro + Local GPU** | **$29** | Opus for hard stuff + unlimited local | **No caps — bulk work is local** |
| Local only (no Claude) | $9 | A- quality only | Stuck on hard problems with no escape hatch |
## 48GB GPU Market (March 22, 2026 — Real Prices) ## 48GB GPU Market (March 22, 2026 — Real Prices)
@@ -222,6 +247,44 @@ Even a single card can run the Qwen3.5-35B-A3B MoE model — just with a smaller
NVLink matters most at large context windows where KV cache spans both cards. NVLink matters most at large context windows where KV cache spans both cards.
At short contexts that fit on one card, the second GPU adds less benefit. At short contexts that fit on one card, the second GPU adds less benefit.
**Speed reality check:** NVLink doesn't make it faster — it prevents the slowdown you'd get
from PCIe when the model spans both cards. The base speed is still Turing (2018 silicon).
10-35 tok/s is fast enough for coding (you read slower than that), but it's not instant.
The MoE architecture (only 3B active params at inference) is what makes it viable on older
hardware — NVLink just removes the inter-GPU bottleneck for 262K context.
### Image Generation (Included — No Extra Cost)
The RTX 5000 has **384 Tensor Cores with native FP16** — full SDXL/Flux support, no hacks.
| Workload | VRAM Needed | Where It Runs |
|----------|-------------|---------------|
| SDXL (1024x1024) | ~8-10GB | Either card alone |
| Flux Dev | ~12-14GB | Single card (16GB) |
| Flux Dev (high-res / batched) | ~18-24GB | Both cards via NVLink (32GB) |
| ComfyUI / InvokeAI | Works natively | No `--force-fp32` needed |
**Run both workloads simultaneously:**
```bash
# Option A: Dedicated cards (no model swapping)
CUDA_VISIBLE_DEVICES=0 # Ollama — coding LLM on GPU 0
CUDA_VISIBLE_DEVICES=1 # ComfyUI/InvokeAI — image gen on GPU 1
# Option B: Both cards unified for whichever task you're doing
# Switch between LLM (32GB, 262K context) and image gen (32GB, high-res batches)
```
### Hardware Longevity: 3-5 Years Realistic
- **2026-2027**: Sweet spot. MoE + linear attention models are getting smaller active params.
32GB unified handles the best coding models at full context. Peak value.
- **2028-2029**: Still useful. The trend is more efficient models, not bigger ones.
32GB likely still runs the best ~35-70B MoE coding models of that era.
- **2030+**: Questionable. New architectures may need FP8, newer tensor core ops that
Turing lacks. But VRAM is VRAM — something useful will always run on 32GB.
- **The cards themselves won't die** — Quadro-grade, designed for 24/7 data center use.
They'll be outclassed before they fail.
### Power Consumption & Cost ### Power Consumption & Cost
| Config | Idle | Load | Monthly (8hr/day @ $0.09/kWh) | Annual | | Config | Idle | Load | Monthly (8hr/day @ $0.09/kWh) | Annual |
@@ -290,24 +353,37 @@ watch -n 0.5 nvidia-smi
| Item | Cost | | Item | Cost |
|------|------| |------|------|
| 1x Quadro RTX 5000 (Phase 1) | $400 | | 2x Quadro RTX 5000 | ~$800 |
| 1x Quadro RTX 5000 (Phase 2) | $400 | | NVLink HB Bridge 2-slot (P4934) | ~$50 |
| NVLink HB Bridge 3-slot | ~$50 | | Dell cables/riser (R720/R730) | ~$60 |
| **Hardware total** | **$850** | | Dell 1100W PSUs (if needed) | ~$60-100 |
| **Hardware total** | **~$960** |
| Monthly power (2 cards, 8hr/day) | $9/mo | | Monthly power (2 cards, 8hr/day) | $9/mo |
| Claude Pro subscription | $20/mo | | Claude Pro subscription (keep) | $20/mo |
| **Monthly operating cost** | **~$29/mo** | | Claude Max subscription (drop) | -$100/mo saved |
| **3-year total cost of ownership** | **$850 + $1,044 = $1,894** | | **Net monthly cost** | **$29/mo (was $100/mo)** |
### The Math: Drop Max, Keep Pro, Add Local
| | Year 1 | Year 2 | Year 3 | **3-Year Total** |
|---|--------|--------|--------|-----------------|
| **Claude Max (current)** | $1,200 | $1,200 | $1,200 | **$3,600** |
| **Pro + Local GPU** | $960 + $348 | $348 | $348 | **$2,004** |
| **Savings** | | | | **$1,596** |
You save ~$71/mo after hardware payoff. The GPU pays for itself in **13 months**.
After that, you're saving $80/mo vs Max with no usage limits on bulk work.
### Comparison: This Build vs Alternatives ### Comparison: This Build vs Alternatives
| Setup | Cost (3yr) | Best Model | Max Context | Quality | | Setup | Monthly | 3yr Total | Limits? | Quality |
|-------|-----------|------------|-------------|---------| |-------|---------|-----------|---------|---------|
| **2x RTX 5000 + $20 Pro** | **$1,894** | Qwen3.5-35B-A3B + Opus | 262K local | **A- local, A+ cloud** | | **Pro + 2x RTX 5000** | **$29** | **$2,004** | **Unlimited local, Pro limits for Opus** | **A- local, A+ cloud** |
| 1x RTX 3090 + $20 Pro | $1,420 | Qwen3.5-35B-A3B + Opus | ~128K local | A- local, A+ cloud | | Pro + 1x RTX 3090 | $25 | $1,620 | Unlimited local (128K ctx), Pro limits | A- local, A+ cloud |
| RTX 8000 (48GB) + $20 Pro | $3,220+ | Qwen3.5-35B-A3B + Opus | 262K+ local | A- local, A+ cloud | | Pro + RTX 8000 (48GB) | $29 | $3,040+ | Unlimited local, Pro limits | A- local, A+ cloud |
| Claude Max only (no GPU) | $3,600 | Opus 4.6 | 200K | A+ cloud only | | **Claude Max (no GPU)** | **$100** | **$3,600** | **Unlimited Opus** | **A+ cloud only** |
| API-only (Opus heavy use) | $18,000+ | Opus 4.6 | 200K | A+ cloud only | | Claude Pro only (no GPU) | $20 | $720 | **Hits caps on large projects** | A+ cloud, limited |
| API-only (Opus heavy use) | $500+ | $18,000+ | Pay per token | A+ cloud only |
### Phased Build Plan ### Phased Build Plan