Research and document the $850 budget build: 2x Quadro RTX 5000 (16GB each) connected via NVLink for 32GB unified VRAM. With Qwen 3.5's Gated Delta Network architecture (Feb 2026), this setup runs the 35B-A3B MoE model with full 262K context in ~25GB — the best price-to-capability ratio for local AI coding available. Includes: - Exact NVLink bridge part numbers (RTX 5000 uses unique smaller connector) - Motherboard/PSU requirements and slot spacing guidance - llama.cpp and Ollama multi-GPU configuration - VRAM budget calculations for all Qwen 3.5 model sizes - Phased build plan (start with 1 card at $400, add second later) - Updated model table with full Qwen 3.5 family specs - Cost comparison vs RTX 8000, RTX 3090, Claude Max, and API pricing https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
20 KiB
GPU Setup Research: Rack Server AI Workloads
Last updated: March 22, 2026
Goal
Cost-efficient rack-mountable GPU setup for:
- LLM coding inference — Run 32B+ parameter coding models with maximum context windows
- Image generation — ComfyUI / InvokeAI with Stable Diffusion SDXL / Flux
Target servers: Dell R720/R730 or HP DL380 equivalent (2U rack)
Why 48GB VRAM is the Right Target
The Problem with 24GB
32B coding models at Q4_K_M quantization use ~20GB of weights, leaving only ~4GB for KV cache on a 24GB card. This severely limits context window size — the key ingredient for complex coding sessions where the model needs to understand your entire codebase.
What 48GB Unlocks
- 32B models at higher quantization (Q6_K/Q8_0) = better output quality
- 28GB+ free for KV cache = massive context windows (32K+ tokens)
- 70B models in aggressive quantization (~12 t/s but functional)
- Simultaneous model loading — coding model + image gen model at once
- Room for future larger models without hardware changes
Best Local Coding Models (2026)
Qwen 3.5 Family (February 2026 — Gated Delta Networks)
Architecture breakthrough: 3 of every 4 layers use linear attention (O(n) scaling), drastically reducing KV cache memory. These models need far less VRAM for long contexts than traditional transformers.
| Model | Type | Active Params | Size at Q4_K_M | Max Context | Quality | Notes |
|---|---|---|---|---|---|---|
| Qwen3.5-35B-A3B | MoE | 3B | ~12GB | 262K | A- | Best bang for buck — 35B model, 3B active, fits 262K ctx in 25GB |
| Qwen3.5-27B | Dense | 27B | ~17GB | 262K | A- | 72.4% SWE-bench, ties GPT-5 mini |
| Qwen3.5-122B-A10B | MoE | 10B | ~76GB | 262K | A | Matches GPT-5 mini across the board |
| Qwen3.5-9B | Dense | 9B | ~6GB | 262K | B+ | Fits on any modern GPU |
| Qwen3.5-4B | Dense | 4B | ~3GB | 262K | B | Tiny but capable |
Previous Generation (Still Relevant)
| Model | Size at Q4_K_M | Quality | Notes |
|---|---|---|---|
| Qwen2.5-Coder 32B | ~20GB | 73.7 Aider (≈ GPT-4o) | FIM king, 92.7% HumanEval |
| Qwen3-Coder 30B-A3B (MoE) | ~18GB | #1 SWE-rebench (64.6%) | Only 3.3B active, very fast |
| Qwen3-Coder-Next 80B (MoE) | needs 64GB+ RAM offload | Beats Claude Opus 4.6 on SWE-rebench | Hybrid attention, 256K context |
Honest Assessment: Local vs Claude Code
Nothing local approaches Claude Opus 4.6 quality for complex multi-file agentic coding. These 32B models are competitive with GPT-4o — a tier below Claude Sonnet, two tiers below Opus. Best strategy: use local models for routine tasks, save Claude credits for hard problems.
48GB GPU Market (March 22, 2026 — Real Prices)
| GPU | Arch | Used Price | TDP | Cooling | Tensor Cores | Mem BW |
|---|---|---|---|---|---|---|
| Quadro RTX 8000 | Turing (2018) | $2,000–2,900 | 260W | Passive variant | Yes (576) | 672 GB/s |
| A40 | Ampere (2020) | ~$5,050+ | 300W | Passive | Yes (336 3rd-gen) | 696 GB/s |
| RTX A6000 | Ampere (2020) | ~$5,400+ | 300W | Active (blower) | Yes (336 3rd-gen) | 768 GB/s |
| L40 | Ada (2022) | ~$6,500+ | 300W | Passive | Yes (568 4th-gen) | 864 GB/s |
| RTX 6000 Ada | Ada (2022) | ~$6,500+ | 300W | Active | Yes (568 4th-gen) | 960 GB/s |
Sources: eBay active/sold listings, GPUPoet price tracking, Pangoly, CamelCamelCamel (all March 2026) Note: One outlier RTX 8000 listing at ~$750 exists but is not representative of the market.
Cheapest 48GB Option: Quadro RTX 8000 Passive ($2,000–2,900)
The RTX 8000 is still the cheapest 48GB card — roughly half the price of an A40 and a third of an A6000. The passive variant is purpose-built for rack servers — no fan, relies on chassis airflow, designed for 24/7 operation in 2U/4U systems.
Key advantages over the P40:
- 48GB vs 24GB — room for models + massive context
- Has Tensor Cores (576 Turing) — native FP16, no
--force-fp32hacks for image gen - NVLink support — pair two for 96GB combined (100 GB/s bidirectional)
- 10W idle power draw
Cost Reality Check
At $2,000–2,900 the RTX 8000 is a significant investment. The key question: is unified 48GB VRAM worth 4–6x the cost of dual P40s ($400–500)?
Yes, if you need large context windows (32K+) for complex coding — KV cache can't be split across two GPUs without NVLink (which P40s don't have).
No, if you're mostly doing short-prompt coding tasks and image gen — dual P40s give you 48GB total (split) at a fraction of the cost, and each card can handle its own workload.
RTX 8000 Performance Benchmarks
LLM Inference (Exllama, 5.0 bpw quantization)
| Model | Context | Prompt Processing | Generation |
|---|---|---|---|
| Qwen3 30B-A3B (MoE) | 8K | 950 t/s | 34 t/s |
| Qwen3 30B-A3B (MoE) | 16K | 673 t/s | 21 t/s |
| Qwen3 30B-A3B (MoE) | 32K | 345 t/s | 11 t/s |
| Llama 3.3 70B | short | 36 t/s | 13 t/s |
| Llama 3.1 8B | — | — | 72 t/s |
Compared to P40 (24GB)
| Metric | P40 (24GB) | RTX 8000 (48GB) |
|---|---|---|
| Used price | $150–320 | $2,000–2,900 |
| 32B model fit | Barely (~2GB free) | Comfortable (~28GB free) |
| 32B generation speed | ~5-12 t/s (est.) | ~20-34 t/s |
| Max practical context | ~4K tokens | 32K+ tokens |
| Image gen (SDXL) | ~49s (--force-fp32) |
Faster (native FP16) |
| Rack server ready | Yes (passive) | Yes (passive variant) |
Image Generation
The RTX 8000 has Turing Tensor Cores with native FP16 support. Unlike the P40, it does NOT
need --force-fp32 workarounds. Image gen performance is significantly better than the P40,
though still behind Ampere/Ada cards.
Budget Build: 2x Quadro RTX 5000 + NVLink ($850 Total)
The best price-to-capability ratio for local AI coding in 2026.
Why This Works Now
Qwen 3.5 (February 2026) introduced Gated Delta Networks — 3 out of 4 layers use linear attention (O(n) scaling) instead of quadratic. KV cache memory usage is dramatically lower than traditional transformers. A 35B MoE model with 262K context now fits in ~25GB VRAM.
Hardware
GPU: NVIDIA Quadro RTX 5000 (Turing, TU104)
| Spec | Value |
|---|---|
| VRAM | 16GB GDDR6 |
| CUDA Cores | 3072 |
| Tensor Cores | 384 (Gen 2, FP16) |
| TDP | ~230W |
| NVLink | Yes — 50 GB/s bidirectional |
| Form Factor | Dual-slot, blower cooler (rack-friendly) |
| PCIe | 3.0 x16 |
| Used Price | ~$400 |
| Part Number | VCQRTX5000-PB |
NVLink Bridge (CRITICAL: RTX 5000 uses a unique smaller connector)
The Quadro RTX 5000 has a shorter NVLink connector than all other Quadro RTX cards. Bridges from the RTX 6000/8000 will NOT physically fit. You must buy the RTX 5000-specific bridge.
| Detail | Value |
|---|---|
| Product | NVIDIA Quadro RTX 5000 NVLink HB Bridge 2-Slot |
| SKU | NVLINKX8-2SLOT-PB |
| Part Numbers | 1JF3K, 699-54934-0500-000, 900-54934-0100-000, P4934, 6FY12AA, L55997-001 |
| Price | ~$30-80 (eBay, Amazon) |
| Bandwidth | 50 GB/s total (25 GB/s per direction) |
| Sizing | 2-slot (cards adjacent) or 3-slot (one slot gap — better thermals) |
Where to buy:
- eBay: search "Quadro RTX 5000 NVLink" or part numbers P4934 / L55997-001 / 1JF3K
- Amazon: search part number 6FY12AA or 1JF3K
WARNING: The 3-slot bridge is recommended over 2-slot. With a 2-slot bridge the cards sit directly adjacent — the top card's blower intake gets blocked by the bottom card. A 3-slot bridge leaves an air gap for proper cooling.
Motherboard Requirements
| Requirement | Details |
|---|---|
| PCIe slots | Two x16 slots (x8 electrical is fine — LLM inference is VRAM-bound, not PCIe-bound) |
| Slot spacing | Must match your NVLink bridge size (2-slot or 3-slot gap) |
| Power supply | 650W+ minimum (80 PLUS Gold recommended), 850W+ for headroom |
| Power connectors | 2x 8-pin PCIe power (one per card). Do NOT daisy-chain — use separate cables |
| CPU platform | Any modern platform works. Threadripper/Xeon not required |
Recommended motherboards (workstation/server):
- Any board with 2x PCIe x16 slots spaced 2-3 slots apart
- Server: Dell R730/R740 with GPU riser (but verify 3-slot bridge clearance in 2U)
- Workstation: MSI X399 Creation, ASUS WS series, Supermicro X11/X12 boards
- Desktop: Most ATX boards with 2 full-length x16 slots work
Rack server note: The Quadro RTX 5000's blower cooler exhausts out the bracket — this works well in rack airflow. If using a 2U server, measure clearance for the NVLink bridge sitting on top of the cards. A 4U chassis gives the most room.
What Runs on 32GB Unified (2x RTX 5000 + NVLink)
| Model | Arch | Quant | Weights | Context | Total VRAM | Quality |
|---|---|---|---|---|---|---|
| Qwen3.5-35B-A3B | MoE (3B active) | Q4_K_M | ~12GB | 262K | ~25GB | A- |
| Qwen3.5-27B | Dense | Q4_K_M | ~17GB | 128K+ | ~25GB | A- |
| Qwen3-Coder-Next (80B/3B active) | MoE | Q4 | ~20GB | 128K | ~28GB | A |
| Qwen2.5-Coder-14B | Dense | Q4_K_M | ~10GB | 128K | ~22GB | B+ |
| Qwen2.5-Coder-14B | Dense | Q8 | ~16GB | 64K | ~28GB | A- |
| Qwen2.5-Coder-32B | Dense | Q4_K_M | ~20GB | 16-24K | ~28GB | A- |
The sweet spot: Qwen3.5-35B-A3B at Q4_K_M with 262K context. This is a 35B parameter model with only 3B active at inference (MoE). The Gated Delta Network architecture slashes KV cache memory. The entire model + full 262K context fits in ~25GB — well within 32GB.
What Runs on 16GB (Single RTX 5000 — Phase 1)
| Model | Quant | Context | Quality |
|---|---|---|---|
| Qwen3.5-35B-A3B | Q4_K_L | ~64-128K | A- |
| Qwen3.5-9B | Q8 | 128K+ | B+ |
| Qwen3.5-4B | Q8 | 262K | B |
| Qwen2.5-Coder-7B | Q8 | 128K | B |
| Qwen2.5-Coder-14B | Q4_K_M | 16-32K | B+ |
Even a single card can run the Qwen3.5-35B-A3B MoE model — just with a smaller context window.
Estimated Inference Speed
| Model | 1x RTX 5000 | 2x RTX 5000 (NVLink) |
|---|---|---|
| Qwen3.5-35B-A3B Q4 (short ctx) | ~25-35 tok/s | ~25-35 tok/s |
| Qwen3.5-35B-A3B Q4 (128K ctx) | ~10-18 tok/s | ~15-25 tok/s |
| Qwen3.5-35B-A3B Q4 (262K ctx) | Won't fit | ~10-18 tok/s |
| Qwen2.5-Coder-14B Q4 | ~20-30 tok/s | ~25-35 tok/s |
NVLink matters most at large context windows where KV cache spans both cards. At short contexts that fit on one card, the second GPU adds less benefit.
Power Consumption & Cost
| Config | Idle | Load | Monthly (8hr/day @ $0.09/kWh) | Annual |
|---|---|---|---|---|
| 1x Quadro RTX 5000 | ~15W | ~210W | ~$4.50 | ~$54 |
| 2x Quadro RTX 5000 | ~30W | ~420W | ~$9.00 | ~$108 |
Software Setup
llama.cpp (Recommended — Best Multi-GPU Support)
# Build with CUDA support
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j$(nproc)
# Download Qwen3.5-35B-A3B GGUF (Q4_K_M)
# Get from: https://huggingface.co/unsloth/Qwen3.5-35B-A3B-GGUF
# Run on dual GPU with NVLink
./build/bin/llama-server \
-m Qwen3.5-35B-A3B-Q4_K_M.gguf \
-ngl 999 \
-c 262144 \
--host 0.0.0.0 \
--port 8080
# llama.cpp auto-detects NVLink and splits layers across both GPUs
# Use -ts 1,1 to manually set equal split if needed
Ollama
# Requires Ollama v0.17+ for Qwen3.5 support
# NOTE: As of March 2026, some Qwen3.5 GGUFs have compatibility issues
# with Ollama due to mmproj vision files. llama.cpp may be more reliable.
# Environment variables for multi-GPU
export OLLAMA_GPU_SPLIT=16,16 # Equal split across both 16GB cards
export OLLAMA_KV_CACHE_TYPE=q8_0 # Halves KV cache VRAM with minimal quality loss
export OLLAMA_KEEP_ALIVE=24h # Keep model loaded in VRAM
export OLLAMA_FLASH_ATTENTION=1 # Enable flash attention for VRAM savings
# Pull and run
ollama pull qwen3.5:35b-a3b-q4_K_M
ollama run qwen3.5:35b-a3b-q4_K_M
Verify NVLink Is Working
# Check NVLink status
nvidia-smi nvlink --status
# Check NVLink bandwidth
nvidia-smi nvlink -gt d
# Monitor both GPUs during inference
watch -n 0.5 nvidia-smi
Total Cost Summary
| Item | Cost |
|---|---|
| 1x Quadro RTX 5000 (Phase 1) | $400 |
| 1x Quadro RTX 5000 (Phase 2) | $400 |
| NVLink HB Bridge 3-slot | ~$50 |
| Hardware total | $850 |
| Monthly power (2 cards, 8hr/day) | $9/mo |
| Claude Pro subscription | $20/mo |
| Monthly operating cost | ~$29/mo |
| 3-year total cost of ownership | $850 + $1,044 = $1,894 |
Comparison: This Build vs Alternatives
| Setup | Cost (3yr) | Best Model | Max Context | Quality |
|---|---|---|---|---|
| 2x RTX 5000 + $20 Pro | $1,894 | Qwen3.5-35B-A3B + Opus | 262K local | A- local, A+ cloud |
| 1x RTX 3090 + $20 Pro | $1,420 | Qwen3.5-35B-A3B + Opus | ~128K local | A- local, A+ cloud |
| RTX 8000 (48GB) + $20 Pro | $3,220+ | Qwen3.5-35B-A3B + Opus | 262K+ local | A- local, A+ cloud |
| Claude Max only (no GPU) | $3,600 | Opus 4.6 | 200K | A+ cloud only |
| API-only (Opus heavy use) | $18,000+ | Opus 4.6 | 200K | A+ cloud only |
Phased Build Plan
Phase 1 — Start with one card ($400)
- Buy Quadro RTX 5000 (VCQRTX5000-PB) — ~$400 on eBay
- Install in any PCIe x16 slot
- Install llama.cpp or Ollama v0.17+
- Run Qwen3.5-35B-A3B at Q4_K_L with 64-128K context
- Already A- quality for coding — test if local inference fits your workflow
Phase 2 — Add second card + NVLink ($450)
- Buy matching Quadro RTX 5000 — ~$400
- Buy NVLink HB Bridge 3-slot (part: P4934 / 1JF3K / 6FY12AA) — ~$50
- Install second card in adjacent/nearby x16 slot
- Connect NVLink bridge
- Verify with
nvidia-smi nvlink --status - Now running 32GB unified — Qwen3.5-35B-A3B at Q4_K_M with full 262K context
24GB GPU Options (Previous Research — Still Valid for Tighter Budgets)
| GPU | VRAM | Price Range | Best Deals | Notes |
|---|---|---|---|---|
| Tesla P40 | 24GB | $150-320 | Newegg refurb $219-270; eBay used $150-200 | Best VRAM/$ at 24GB |
| RTX A2000 12GB | 12GB | $250-535 | eBay used ~$250-350; one listing at $490 | Can't run 32B models |
| Tesla T4 | 16GB | $150-350 | eBay used $150-250 | Great power efficiency |
| RTX A4000 | 16GB | $700-750+ | eBay used ~$700; new $720+ | Too expensive for 16GB |
Rack Server Compatibility
Quadro RTX 8000 Passive in R720/R730
- Physical fit: Full-length, dual-slot — fits in GPU riser slots
- Power: 260W, requires 8-pin aux power + GPU enablement kit
- Cooling: Passive — relies on server chassis fans (same as P40)
- Requirement: Dual CPUs, redundant 1100W PSUs recommended
- NVLink: Can pair two RTX 8000s for 96GB combined VRAM
- Very similar physical/power requirements to the Tesla P40
RTX A2000 in R720/R730
- Physical fit: Yes. Dual-slot, low-profile, 167mm length
- Power: 70W bus-powered, no aux cable needed. Must use 75W slots (slots 4-7 on R720)
- Cooling: Blower-style fan exhausts out bracket — ideal for rack airflow
- Requirement: Dual CPUs needed for GPU PCIe slots
- Confirmed working in Dell R740XD (similar architecture)
Tesla P40 in R720/R730
- Physical fit: Yes. Full-length, single-slot, designed for rack servers
- Power: 250W, requires 8-pin aux power. Needs GPU enablement kit
- Cooling: Passive — relies on server chassis fans
- Requirement: Dual CPUs, redundant 1100W PSUs recommended
- Natively supported in these servers
R720 vs R730
- R720: PCIe Gen2 (not a bottleneck for LLM inference, which is VRAM-bound)
- R730: PCIe Gen3, generally preferred
- Both support up to 2x double-wide or 4x single-wide GPUs
Recommended Setups
If budget allows ($2,000–2,900): RTX 8000 Passive
Single card handles both coding and image gen. 48GB VRAM fits 32B models with massive context windows (32K+). Passive cooling is rack-native. Tensor cores handle FP16 image gen properly. One card, one slot, simple setup. The premium buys you unified VRAM = big context.
If budget allows + dedicated image gen ($2,300–3,250): RTX 8000 + A2000
RTX 8000 for coding with full 48GB dedicated to LLM context. A2000 for image gen (3x faster than Turing, 70W, bus-powered, blower cooled). Best separation of concerns — no model swapping needed.
Best value ($400–500): Dual P40
Two P40s for 48GB total, but split across cards (can't combine for one model without NVLink, which P40s lack). One for 32B coding (tight fit, ~4K context), one for image gen (slow, needs --force-fp32). 5x cheaper than RTX 8000 but with significant context limitations.
Cheapest entry ($200–300): Single P40
Run 32B coding model with very limited context (~4K tokens). Swap to image gen when needed. Good for testing whether local LLM coding works for your workflow before investing more.
Configuration Notes for local-ai stack
For 48GB RTX 8000
# Ollama — take advantage of the full 48GB
OLLAMA_NUM_GPU=999
OLLAMA_NUM_CTX=32768 # Large context window — 48GB can handle it
OLLAMA_KEEP_ALIVE=24h
# Pull best coding models
ollama pull qwen2.5-coder:32b-instruct-q4_K_M # ~20GB, leaves 28GB for context
ollama pull qwen3.5:27b # ~16GB at Q4, even more context room
ollama pull qwen3-coder:30b # MoE, very fast inference
# Higher quantization for better quality (48GB allows this)
# Look for Q6_K or Q8_0 variants on Ollama for better output quality
For 32B models on P40 (24GB — tight fit)
OLLAMA_NUM_GPU=999
OLLAMA_NUM_CTX=4096 # Keep context small to fit in remaining VRAM
OLLAMA_KEEP_ALIVE=24h
ollama pull qwen2.5-coder:32b-instruct-q4_K_M
For dual-GPU setup (RTX 8000 + A2000 or P40 + anything)
# Assign GPU 0 to Ollama (coding), GPU 1 to InvokeAI (image gen)
# In docker-compose.yml for Ollama:
CUDA_VISIBLE_DEVICES=0
# In docker-compose.yml for InvokeAI:
CUDA_VISIBLE_DEVICES=1
For image gen on P40 (no tensor cores)
# InvokeAI
INVOKEAI_PRECISION=float32
# ComfyUI launch args
--force-fp32
For image gen on RTX 8000 / A2000 / T4 (has tensor cores)
# InvokeAI — native FP16 works fine
INVOKEAI_PRECISION=float16
# ComfyUI — no special flags needed
Sources
- Quadro RTX 8000 for Local LLMs — Hardware Corner
- RTX 8000 Passive — Network Outlet
- LLM Benchmarks on Turing/Ampere GPUs — Stefandroid
- NVIDIA A40 Price Tracking — GPUPoet
- NVIDIA L40 Price Tracking — GPUPoet
- RTX A6000 Price History — CamelCamelCamel
- RTX A6000 Price History — Pangoly
- NVIDIA RTX A2000 Datasheet
- Dell R730 Owner's Manual — Expansion Cards
- Dell R720 Owner's Manual — Expansion Cards
- ComfyUI GPU Benchmarks Discussion
- ComfyUI P40 FP32 Issue
- Best Local LLMs for 24GB VRAM 2026
- Best Coding Models 2026
- Ollama VRAM Requirements Guide
- Local LLMs That Can Replace Claude Code
- 7 Local LLM Families to Replace Claude/Codex
- Qwen2.5-Coder 32B on Ollama
- Qwen3-Coder — How to Run Locally