Files
local-ai/docs/gpu-setup-research.md
T
Claude 7fa91e9816 Add Quadro RTX 5000 NVLink budget build guide and Qwen 3.5 models
Research and document the $850 budget build: 2x Quadro RTX 5000 (16GB each)
connected via NVLink for 32GB unified VRAM. With Qwen 3.5's Gated Delta
Network architecture (Feb 2026), this setup runs the 35B-A3B MoE model
with full 262K context in ~25GB — the best price-to-capability ratio
for local AI coding available.

Includes:
- Exact NVLink bridge part numbers (RTX 5000 uses unique smaller connector)
- Motherboard/PSU requirements and slot spacing guidance
- llama.cpp and Ollama multi-GPU configuration
- VRAM budget calculations for all Qwen 3.5 model sizes
- Phased build plan (start with 1 card at $400, add second later)
- Updated model table with full Qwen 3.5 family specs
- Cost comparison vs RTX 8000, RTX 3090, Claude Max, and API pricing

https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
2026-03-22 15:57:24 +00:00

20 KiB
Raw Blame History

GPU Setup Research: Rack Server AI Workloads

Last updated: March 22, 2026

Goal

Cost-efficient rack-mountable GPU setup for:

  1. LLM coding inference — Run 32B+ parameter coding models with maximum context windows
  2. Image generation — ComfyUI / InvokeAI with Stable Diffusion SDXL / Flux

Target servers: Dell R720/R730 or HP DL380 equivalent (2U rack)

Why 48GB VRAM is the Right Target

The Problem with 24GB

32B coding models at Q4_K_M quantization use ~20GB of weights, leaving only ~4GB for KV cache on a 24GB card. This severely limits context window size — the key ingredient for complex coding sessions where the model needs to understand your entire codebase.

What 48GB Unlocks

  • 32B models at higher quantization (Q6_K/Q8_0) = better output quality
  • 28GB+ free for KV cache = massive context windows (32K+ tokens)
  • 70B models in aggressive quantization (~12 t/s but functional)
  • Simultaneous model loading — coding model + image gen model at once
  • Room for future larger models without hardware changes

Best Local Coding Models (2026)

Qwen 3.5 Family (February 2026 — Gated Delta Networks)

Architecture breakthrough: 3 of every 4 layers use linear attention (O(n) scaling), drastically reducing KV cache memory. These models need far less VRAM for long contexts than traditional transformers.

Model Type Active Params Size at Q4_K_M Max Context Quality Notes
Qwen3.5-35B-A3B MoE 3B ~12GB 262K A- Best bang for buck — 35B model, 3B active, fits 262K ctx in 25GB
Qwen3.5-27B Dense 27B ~17GB 262K A- 72.4% SWE-bench, ties GPT-5 mini
Qwen3.5-122B-A10B MoE 10B ~76GB 262K A Matches GPT-5 mini across the board
Qwen3.5-9B Dense 9B ~6GB 262K B+ Fits on any modern GPU
Qwen3.5-4B Dense 4B ~3GB 262K B Tiny but capable

Previous Generation (Still Relevant)

Model Size at Q4_K_M Quality Notes
Qwen2.5-Coder 32B ~20GB 73.7 Aider (≈ GPT-4o) FIM king, 92.7% HumanEval
Qwen3-Coder 30B-A3B (MoE) ~18GB #1 SWE-rebench (64.6%) Only 3.3B active, very fast
Qwen3-Coder-Next 80B (MoE) needs 64GB+ RAM offload Beats Claude Opus 4.6 on SWE-rebench Hybrid attention, 256K context

Honest Assessment: Local vs Claude Code

Nothing local approaches Claude Opus 4.6 quality for complex multi-file agentic coding. These 32B models are competitive with GPT-4o — a tier below Claude Sonnet, two tiers below Opus. Best strategy: use local models for routine tasks, save Claude credits for hard problems.

48GB GPU Market (March 22, 2026 — Real Prices)

GPU Arch Used Price TDP Cooling Tensor Cores Mem BW
Quadro RTX 8000 Turing (2018) $2,0002,900 260W Passive variant Yes (576) 672 GB/s
A40 Ampere (2020) ~$5,050+ 300W Passive Yes (336 3rd-gen) 696 GB/s
RTX A6000 Ampere (2020) ~$5,400+ 300W Active (blower) Yes (336 3rd-gen) 768 GB/s
L40 Ada (2022) ~$6,500+ 300W Passive Yes (568 4th-gen) 864 GB/s
RTX 6000 Ada Ada (2022) ~$6,500+ 300W Active Yes (568 4th-gen) 960 GB/s

Sources: eBay active/sold listings, GPUPoet price tracking, Pangoly, CamelCamelCamel (all March 2026) Note: One outlier RTX 8000 listing at ~$750 exists but is not representative of the market.

Cheapest 48GB Option: Quadro RTX 8000 Passive ($2,0002,900)

The RTX 8000 is still the cheapest 48GB card — roughly half the price of an A40 and a third of an A6000. The passive variant is purpose-built for rack servers — no fan, relies on chassis airflow, designed for 24/7 operation in 2U/4U systems.

Key advantages over the P40:

  • 48GB vs 24GB — room for models + massive context
  • Has Tensor Cores (576 Turing) — native FP16, no --force-fp32 hacks for image gen
  • NVLink support — pair two for 96GB combined (100 GB/s bidirectional)
  • 10W idle power draw

Cost Reality Check

At $2,0002,900 the RTX 8000 is a significant investment. The key question: is unified 48GB VRAM worth 46x the cost of dual P40s ($400500)?

Yes, if you need large context windows (32K+) for complex coding — KV cache can't be split across two GPUs without NVLink (which P40s don't have).

No, if you're mostly doing short-prompt coding tasks and image gen — dual P40s give you 48GB total (split) at a fraction of the cost, and each card can handle its own workload.

RTX 8000 Performance Benchmarks

LLM Inference (Exllama, 5.0 bpw quantization)

Model Context Prompt Processing Generation
Qwen3 30B-A3B (MoE) 8K 950 t/s 34 t/s
Qwen3 30B-A3B (MoE) 16K 673 t/s 21 t/s
Qwen3 30B-A3B (MoE) 32K 345 t/s 11 t/s
Llama 3.3 70B short 36 t/s 13 t/s
Llama 3.1 8B 72 t/s

Compared to P40 (24GB)

Metric P40 (24GB) RTX 8000 (48GB)
Used price $150320 $2,0002,900
32B model fit Barely (~2GB free) Comfortable (~28GB free)
32B generation speed ~5-12 t/s (est.) ~20-34 t/s
Max practical context ~4K tokens 32K+ tokens
Image gen (SDXL) ~49s (--force-fp32) Faster (native FP16)
Rack server ready Yes (passive) Yes (passive variant)

Image Generation

The RTX 8000 has Turing Tensor Cores with native FP16 support. Unlike the P40, it does NOT need --force-fp32 workarounds. Image gen performance is significantly better than the P40, though still behind Ampere/Ada cards.

The best price-to-capability ratio for local AI coding in 2026.

Why This Works Now

Qwen 3.5 (February 2026) introduced Gated Delta Networks — 3 out of 4 layers use linear attention (O(n) scaling) instead of quadratic. KV cache memory usage is dramatically lower than traditional transformers. A 35B MoE model with 262K context now fits in ~25GB VRAM.

Hardware

GPU: NVIDIA Quadro RTX 5000 (Turing, TU104)

Spec Value
VRAM 16GB GDDR6
CUDA Cores 3072
Tensor Cores 384 (Gen 2, FP16)
TDP ~230W
NVLink Yes — 50 GB/s bidirectional
Form Factor Dual-slot, blower cooler (rack-friendly)
PCIe 3.0 x16
Used Price ~$400
Part Number VCQRTX5000-PB

The Quadro RTX 5000 has a shorter NVLink connector than all other Quadro RTX cards. Bridges from the RTX 6000/8000 will NOT physically fit. You must buy the RTX 5000-specific bridge.

Detail Value
Product NVIDIA Quadro RTX 5000 NVLink HB Bridge 2-Slot
SKU NVLINKX8-2SLOT-PB
Part Numbers 1JF3K, 699-54934-0500-000, 900-54934-0100-000, P4934, 6FY12AA, L55997-001
Price ~$30-80 (eBay, Amazon)
Bandwidth 50 GB/s total (25 GB/s per direction)
Sizing 2-slot (cards adjacent) or 3-slot (one slot gap — better thermals)

Where to buy:

  • eBay: search "Quadro RTX 5000 NVLink" or part numbers P4934 / L55997-001 / 1JF3K
  • Amazon: search part number 6FY12AA or 1JF3K

WARNING: The 3-slot bridge is recommended over 2-slot. With a 2-slot bridge the cards sit directly adjacent — the top card's blower intake gets blocked by the bottom card. A 3-slot bridge leaves an air gap for proper cooling.

Motherboard Requirements

Requirement Details
PCIe slots Two x16 slots (x8 electrical is fine — LLM inference is VRAM-bound, not PCIe-bound)
Slot spacing Must match your NVLink bridge size (2-slot or 3-slot gap)
Power supply 650W+ minimum (80 PLUS Gold recommended), 850W+ for headroom
Power connectors 2x 8-pin PCIe power (one per card). Do NOT daisy-chain — use separate cables
CPU platform Any modern platform works. Threadripper/Xeon not required

Recommended motherboards (workstation/server):

  • Any board with 2x PCIe x16 slots spaced 2-3 slots apart
  • Server: Dell R730/R740 with GPU riser (but verify 3-slot bridge clearance in 2U)
  • Workstation: MSI X399 Creation, ASUS WS series, Supermicro X11/X12 boards
  • Desktop: Most ATX boards with 2 full-length x16 slots work

Rack server note: The Quadro RTX 5000's blower cooler exhausts out the bracket — this works well in rack airflow. If using a 2U server, measure clearance for the NVLink bridge sitting on top of the cards. A 4U chassis gives the most room.

Model Arch Quant Weights Context Total VRAM Quality
Qwen3.5-35B-A3B MoE (3B active) Q4_K_M ~12GB 262K ~25GB A-
Qwen3.5-27B Dense Q4_K_M ~17GB 128K+ ~25GB A-
Qwen3-Coder-Next (80B/3B active) MoE Q4 ~20GB 128K ~28GB A
Qwen2.5-Coder-14B Dense Q4_K_M ~10GB 128K ~22GB B+
Qwen2.5-Coder-14B Dense Q8 ~16GB 64K ~28GB A-
Qwen2.5-Coder-32B Dense Q4_K_M ~20GB 16-24K ~28GB A-

The sweet spot: Qwen3.5-35B-A3B at Q4_K_M with 262K context. This is a 35B parameter model with only 3B active at inference (MoE). The Gated Delta Network architecture slashes KV cache memory. The entire model + full 262K context fits in ~25GB — well within 32GB.

What Runs on 16GB (Single RTX 5000 — Phase 1)

Model Quant Context Quality
Qwen3.5-35B-A3B Q4_K_L ~64-128K A-
Qwen3.5-9B Q8 128K+ B+
Qwen3.5-4B Q8 262K B
Qwen2.5-Coder-7B Q8 128K B
Qwen2.5-Coder-14B Q4_K_M 16-32K B+

Even a single card can run the Qwen3.5-35B-A3B MoE model — just with a smaller context window.

Estimated Inference Speed

Model 1x RTX 5000 2x RTX 5000 (NVLink)
Qwen3.5-35B-A3B Q4 (short ctx) ~25-35 tok/s ~25-35 tok/s
Qwen3.5-35B-A3B Q4 (128K ctx) ~10-18 tok/s ~15-25 tok/s
Qwen3.5-35B-A3B Q4 (262K ctx) Won't fit ~10-18 tok/s
Qwen2.5-Coder-14B Q4 ~20-30 tok/s ~25-35 tok/s

NVLink matters most at large context windows where KV cache spans both cards. At short contexts that fit on one card, the second GPU adds less benefit.

Power Consumption & Cost

Config Idle Load Monthly (8hr/day @ $0.09/kWh) Annual
1x Quadro RTX 5000 ~15W ~210W ~$4.50 ~$54
2x Quadro RTX 5000 ~30W ~420W ~$9.00 ~$108

Software Setup

# Build with CUDA support
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j$(nproc)

# Download Qwen3.5-35B-A3B GGUF (Q4_K_M)
# Get from: https://huggingface.co/unsloth/Qwen3.5-35B-A3B-GGUF

# Run on dual GPU with NVLink
./build/bin/llama-server \
  -m Qwen3.5-35B-A3B-Q4_K_M.gguf \
  -ngl 999 \
  -c 262144 \
  --host 0.0.0.0 \
  --port 8080

# llama.cpp auto-detects NVLink and splits layers across both GPUs
# Use -ts 1,1 to manually set equal split if needed

Ollama

# Requires Ollama v0.17+ for Qwen3.5 support
# NOTE: As of March 2026, some Qwen3.5 GGUFs have compatibility issues
# with Ollama due to mmproj vision files. llama.cpp may be more reliable.

# Environment variables for multi-GPU
export OLLAMA_GPU_SPLIT=16,16        # Equal split across both 16GB cards
export OLLAMA_KV_CACHE_TYPE=q8_0     # Halves KV cache VRAM with minimal quality loss
export OLLAMA_KEEP_ALIVE=24h         # Keep model loaded in VRAM
export OLLAMA_FLASH_ATTENTION=1      # Enable flash attention for VRAM savings

# Pull and run
ollama pull qwen3.5:35b-a3b-q4_K_M
ollama run qwen3.5:35b-a3b-q4_K_M
# Check NVLink status
nvidia-smi nvlink --status

# Check NVLink bandwidth
nvidia-smi nvlink -gt d

# Monitor both GPUs during inference
watch -n 0.5 nvidia-smi

Total Cost Summary

Item Cost
1x Quadro RTX 5000 (Phase 1) $400
1x Quadro RTX 5000 (Phase 2) $400
NVLink HB Bridge 3-slot ~$50
Hardware total $850
Monthly power (2 cards, 8hr/day) $9/mo
Claude Pro subscription $20/mo
Monthly operating cost ~$29/mo
3-year total cost of ownership $850 + $1,044 = $1,894

Comparison: This Build vs Alternatives

Setup Cost (3yr) Best Model Max Context Quality
2x RTX 5000 + $20 Pro $1,894 Qwen3.5-35B-A3B + Opus 262K local A- local, A+ cloud
1x RTX 3090 + $20 Pro $1,420 Qwen3.5-35B-A3B + Opus ~128K local A- local, A+ cloud
RTX 8000 (48GB) + $20 Pro $3,220+ Qwen3.5-35B-A3B + Opus 262K+ local A- local, A+ cloud
Claude Max only (no GPU) $3,600 Opus 4.6 200K A+ cloud only
API-only (Opus heavy use) $18,000+ Opus 4.6 200K A+ cloud only

Phased Build Plan

Phase 1 — Start with one card ($400)

  1. Buy Quadro RTX 5000 (VCQRTX5000-PB) — ~$400 on eBay
  2. Install in any PCIe x16 slot
  3. Install llama.cpp or Ollama v0.17+
  4. Run Qwen3.5-35B-A3B at Q4_K_L with 64-128K context
  5. Already A- quality for coding — test if local inference fits your workflow

Phase 2 — Add second card + NVLink ($450)

  1. Buy matching Quadro RTX 5000 — ~$400
  2. Buy NVLink HB Bridge 3-slot (part: P4934 / 1JF3K / 6FY12AA) — ~$50
  3. Install second card in adjacent/nearby x16 slot
  4. Connect NVLink bridge
  5. Verify with nvidia-smi nvlink --status
  6. Now running 32GB unified — Qwen3.5-35B-A3B at Q4_K_M with full 262K context

24GB GPU Options (Previous Research — Still Valid for Tighter Budgets)

GPU VRAM Price Range Best Deals Notes
Tesla P40 24GB $150-320 Newegg refurb $219-270; eBay used $150-200 Best VRAM/$ at 24GB
RTX A2000 12GB 12GB $250-535 eBay used ~$250-350; one listing at $490 Can't run 32B models
Tesla T4 16GB $150-350 eBay used $150-250 Great power efficiency
RTX A4000 16GB $700-750+ eBay used ~$700; new $720+ Too expensive for 16GB

Rack Server Compatibility

Quadro RTX 8000 Passive in R720/R730

  • Physical fit: Full-length, dual-slot — fits in GPU riser slots
  • Power: 260W, requires 8-pin aux power + GPU enablement kit
  • Cooling: Passive — relies on server chassis fans (same as P40)
  • Requirement: Dual CPUs, redundant 1100W PSUs recommended
  • NVLink: Can pair two RTX 8000s for 96GB combined VRAM
  • Very similar physical/power requirements to the Tesla P40

RTX A2000 in R720/R730

  • Physical fit: Yes. Dual-slot, low-profile, 167mm length
  • Power: 70W bus-powered, no aux cable needed. Must use 75W slots (slots 4-7 on R720)
  • Cooling: Blower-style fan exhausts out bracket — ideal for rack airflow
  • Requirement: Dual CPUs needed for GPU PCIe slots
  • Confirmed working in Dell R740XD (similar architecture)

Tesla P40 in R720/R730

  • Physical fit: Yes. Full-length, single-slot, designed for rack servers
  • Power: 250W, requires 8-pin aux power. Needs GPU enablement kit
  • Cooling: Passive — relies on server chassis fans
  • Requirement: Dual CPUs, redundant 1100W PSUs recommended
  • Natively supported in these servers

R720 vs R730

  • R720: PCIe Gen2 (not a bottleneck for LLM inference, which is VRAM-bound)
  • R730: PCIe Gen3, generally preferred
  • Both support up to 2x double-wide or 4x single-wide GPUs

If budget allows ($2,0002,900): RTX 8000 Passive

Single card handles both coding and image gen. 48GB VRAM fits 32B models with massive context windows (32K+). Passive cooling is rack-native. Tensor cores handle FP16 image gen properly. One card, one slot, simple setup. The premium buys you unified VRAM = big context.

If budget allows + dedicated image gen ($2,3003,250): RTX 8000 + A2000

RTX 8000 for coding with full 48GB dedicated to LLM context. A2000 for image gen (3x faster than Turing, 70W, bus-powered, blower cooled). Best separation of concerns — no model swapping needed.

Best value ($400500): Dual P40

Two P40s for 48GB total, but split across cards (can't combine for one model without NVLink, which P40s lack). One for 32B coding (tight fit, ~4K context), one for image gen (slow, needs --force-fp32). 5x cheaper than RTX 8000 but with significant context limitations.

Cheapest entry ($200300): Single P40

Run 32B coding model with very limited context (~4K tokens). Swap to image gen when needed. Good for testing whether local LLM coding works for your workflow before investing more.

Configuration Notes for local-ai stack

For 48GB RTX 8000

# Ollama — take advantage of the full 48GB
OLLAMA_NUM_GPU=999
OLLAMA_NUM_CTX=32768  # Large context window — 48GB can handle it
OLLAMA_KEEP_ALIVE=24h

# Pull best coding models
ollama pull qwen2.5-coder:32b-instruct-q4_K_M  # ~20GB, leaves 28GB for context
ollama pull qwen3.5:27b                          # ~16GB at Q4, even more context room
ollama pull qwen3-coder:30b                      # MoE, very fast inference

# Higher quantization for better quality (48GB allows this)
# Look for Q6_K or Q8_0 variants on Ollama for better output quality

For 32B models on P40 (24GB — tight fit)

OLLAMA_NUM_GPU=999
OLLAMA_NUM_CTX=4096  # Keep context small to fit in remaining VRAM
OLLAMA_KEEP_ALIVE=24h

ollama pull qwen2.5-coder:32b-instruct-q4_K_M

For dual-GPU setup (RTX 8000 + A2000 or P40 + anything)

# Assign GPU 0 to Ollama (coding), GPU 1 to InvokeAI (image gen)
# In docker-compose.yml for Ollama:
CUDA_VISIBLE_DEVICES=0

# In docker-compose.yml for InvokeAI:
CUDA_VISIBLE_DEVICES=1

For image gen on P40 (no tensor cores)

# InvokeAI
INVOKEAI_PRECISION=float32

# ComfyUI launch args
--force-fp32

For image gen on RTX 8000 / A2000 / T4 (has tensor cores)

# InvokeAI — native FP16 works fine
INVOKEAI_PRECISION=float16

# ComfyUI — no special flags needed

Sources