Reframe guide around Drop Max / Keep Pro strategy

The real use case: drop Claude Max ($100/mo), keep Pro ($20/mo), offload
bulk codebase work to local. Local handles the 80% (reading 10K-line
projects, routine fixes, boilerplate) with no rate limits. Pro handles
the hard 20% where Opus quality matters. GPU pays for itself in 13 months,
saves $1,596 over 3 years vs Max.

Also adds: NVLink speed reality check, image gen capabilities (SDXL/Flux
included at no extra cost), hardware longevity estimate (3-5 years),
and updated cost comparison tables.

https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
This commit is contained in:
Claude
2026-03-22 17:07:01 +00:00
parent 1e7b087a5d
commit 5cfd8b2696
+91 -15
View File
@@ -48,9 +48,34 @@ than traditional transformers.
| **Qwen3-Coder-Next 80B** (MoE) | needs 64GB+ RAM offload | Beats Claude Opus 4.6 on SWE-rebench | Hybrid attention, 256K context |
### Honest Assessment: Local vs Claude Code
Nothing local approaches Claude Opus 4.6 quality for complex multi-file agentic coding.
These 32B models are competitive with **GPT-4o** — a tier below Claude Sonnet, two tiers below Opus.
Best strategy: use local models for routine tasks, save Claude credits for hard problems.
**The real strategy: Drop Max ($100/mo), keep Pro ($20/mo), offload bulk work to local.**
The problem with Pro for large projects: rate limits. A 10,000-line codebase needs the model
to read, understand, and hold context across many files. On Pro you'll hit usage caps mid-session
on complex multi-file work. Max ($100/mo) removes those limits — but that's $80/mo extra.
Local AI eliminates this problem differently:
- **Local model (262K context)**: Reads your entire 10K-line project at once. No rate limits,
no usage caps, runs 24/7. Handles the bulk work — understanding codebase structure, routine
bug fixes, simple refactors, code explanation, test writing, boilerplate generation.
- **Claude Pro ($20/mo)**: Reserved for the hard problems — complex multi-file architectural
changes, subtle bugs that need Opus-level reasoning, code review on critical paths.
Pro limits are fine when you're only sending Claude the *hard* 20% instead of everything.
This is the unlock: local doesn't replace Claude, it **reduces your Claude usage enough
that Pro limits stop being a problem.** The 80% of routine work that was burning through
your Max quota now runs locally with zero limits.
| Plan | Monthly | What You Get | Limit Problem |
|------|---------|--------------|---------------|
| Max only | $100 | Opus unlimited | Paying $80/mo for unlimited when you don't need it |
| Pro only | $20 | Opus with rate limits | **Hits caps on 10K-line projects** |
| **Pro + Local GPU** | **$29** | Opus for hard stuff + unlimited local | **No caps — bulk work is local** |
| Local only (no Claude) | $9 | A- quality only | Stuck on hard problems with no escape hatch |
## 48GB GPU Market (March 22, 2026 — Real Prices)
@@ -222,6 +247,44 @@ Even a single card can run the Qwen3.5-35B-A3B MoE model — just with a smaller
NVLink matters most at large context windows where KV cache spans both cards.
At short contexts that fit on one card, the second GPU adds less benefit.
**Speed reality check:** NVLink doesn't make it faster — it prevents the slowdown you'd get
from PCIe when the model spans both cards. The base speed is still Turing (2018 silicon).
10-35 tok/s is fast enough for coding (you read slower than that), but it's not instant.
The MoE architecture (only 3B active params at inference) is what makes it viable on older
hardware — NVLink just removes the inter-GPU bottleneck for 262K context.
### Image Generation (Included — No Extra Cost)
The RTX 5000 has **384 Tensor Cores with native FP16** — full SDXL/Flux support, no hacks.
| Workload | VRAM Needed | Where It Runs |
|----------|-------------|---------------|
| SDXL (1024x1024) | ~8-10GB | Either card alone |
| Flux Dev | ~12-14GB | Single card (16GB) |
| Flux Dev (high-res / batched) | ~18-24GB | Both cards via NVLink (32GB) |
| ComfyUI / InvokeAI | Works natively | No `--force-fp32` needed |
**Run both workloads simultaneously:**
```bash
# Option A: Dedicated cards (no model swapping)
CUDA_VISIBLE_DEVICES=0 # Ollama — coding LLM on GPU 0
CUDA_VISIBLE_DEVICES=1 # ComfyUI/InvokeAI — image gen on GPU 1
# Option B: Both cards unified for whichever task you're doing
# Switch between LLM (32GB, 262K context) and image gen (32GB, high-res batches)
```
### Hardware Longevity: 3-5 Years Realistic
- **2026-2027**: Sweet spot. MoE + linear attention models are getting smaller active params.
32GB unified handles the best coding models at full context. Peak value.
- **2028-2029**: Still useful. The trend is more efficient models, not bigger ones.
32GB likely still runs the best ~35-70B MoE coding models of that era.
- **2030+**: Questionable. New architectures may need FP8, newer tensor core ops that
Turing lacks. But VRAM is VRAM — something useful will always run on 32GB.
- **The cards themselves won't die** — Quadro-grade, designed for 24/7 data center use.
They'll be outclassed before they fail.
### Power Consumption & Cost
| Config | Idle | Load | Monthly (8hr/day @ $0.09/kWh) | Annual |
@@ -290,24 +353,37 @@ watch -n 0.5 nvidia-smi
| Item | Cost |
|------|------|
| 1x Quadro RTX 5000 (Phase 1) | $400 |
| 1x Quadro RTX 5000 (Phase 2) | $400 |
| NVLink HB Bridge 3-slot | ~$50 |
| **Hardware total** | **$850** |
| 2x Quadro RTX 5000 | ~$800 |
| NVLink HB Bridge 2-slot (P4934) | ~$50 |
| Dell cables/riser (R720/R730) | ~$60 |
| Dell 1100W PSUs (if needed) | ~$60-100 |
| **Hardware total** | **~$960** |
| Monthly power (2 cards, 8hr/day) | $9/mo |
| Claude Pro subscription | $20/mo |
| **Monthly operating cost** | **~$29/mo** |
| **3-year total cost of ownership** | **$850 + $1,044 = $1,894** |
| Claude Pro subscription (keep) | $20/mo |
| Claude Max subscription (drop) | -$100/mo saved |
| **Net monthly cost** | **$29/mo (was $100/mo)** |
### The Math: Drop Max, Keep Pro, Add Local
| | Year 1 | Year 2 | Year 3 | **3-Year Total** |
|---|--------|--------|--------|-----------------|
| **Claude Max (current)** | $1,200 | $1,200 | $1,200 | **$3,600** |
| **Pro + Local GPU** | $960 + $348 | $348 | $348 | **$2,004** |
| **Savings** | | | | **$1,596** |
You save ~$71/mo after hardware payoff. The GPU pays for itself in **13 months**.
After that, you're saving $80/mo vs Max with no usage limits on bulk work.
### Comparison: This Build vs Alternatives
| Setup | Cost (3yr) | Best Model | Max Context | Quality |
|-------|-----------|------------|-------------|---------|
| **2x RTX 5000 + $20 Pro** | **$1,894** | Qwen3.5-35B-A3B + Opus | 262K local | **A- local, A+ cloud** |
| 1x RTX 3090 + $20 Pro | $1,420 | Qwen3.5-35B-A3B + Opus | ~128K local | A- local, A+ cloud |
| RTX 8000 (48GB) + $20 Pro | $3,220+ | Qwen3.5-35B-A3B + Opus | 262K+ local | A- local, A+ cloud |
| Claude Max only (no GPU) | $3,600 | Opus 4.6 | 200K | A+ cloud only |
| API-only (Opus heavy use) | $18,000+ | Opus 4.6 | 200K | A+ cloud only |
| Setup | Monthly | 3yr Total | Limits? | Quality |
|-------|---------|-----------|---------|---------|
| **Pro + 2x RTX 5000** | **$29** | **$2,004** | **Unlimited local, Pro limits for Opus** | **A- local, A+ cloud** |
| Pro + 1x RTX 3090 | $25 | $1,620 | Unlimited local (128K ctx), Pro limits | A- local, A+ cloud |
| Pro + RTX 8000 (48GB) | $29 | $3,040+ | Unlimited local, Pro limits | A- local, A+ cloud |
| **Claude Max (no GPU)** | **$100** | **$3,600** | **Unlimited Opus** | **A+ cloud only** |
| Claude Pro only (no GPU) | $20 | $720 | **Hits caps on large projects** | A+ cloud, limited |
| API-only (Opus heavy use) | $500+ | $18,000+ | Pay per token | A+ cloud only |
### Phased Build Plan