- Corrected all GPU prices to actual eBay/Newegg listings as of March 22, 2026 - Added analysis of 32B coding models (Qwen2.5-Coder 32B, Qwen3.5 27B) on 24GB VRAM - Added honest comparison of local LLM quality vs Claude Code - Revised recommendations: dual P40 ($400-500) or P40+T4 ($400-550) - Added configuration notes for 32B models, dual-GPU, and newer Qwen3 models - RTX A4000 at $700+ is too expensive for this build https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
9.1 KiB
GPU Setup Research: Rack Server AI Workloads
Last updated: March 22, 2026
Goal
Cost-efficient rack-mountable GPU setup for:
- LLM coding inference — Run 32B parameter coding models (Qwen2.5-Coder 32B, Qwen3 32B) for quality closest to Claude/GPT-4o
- Image generation — ComfyUI / InvokeAI with Stable Diffusion SDXL / Flux
Target servers: Dell R720/R730 or HP DL380 equivalent (2U rack)
Why 24GB VRAM Matters for Coding
32B parameter models are the sweet spot for local coding quality:
- Qwen2.5-Coder 32B scores 73.7 on Aider (between GPT-4o at 71% and Claude 3.5 Haiku at 75%)
- Competitive with GPT-4o on EvalPlus, LiveCodeBench, BigCodeBench
- At Q4_K_M quantization (~20GB), fits on 24GB with room for small context
- 14B models are noticeably worse for complex coding tasks
Newer models also fit 24GB:
- Qwen 3.5 27B (dense): ties GPT-5 mini on SWE-bench (72.4%), ~16GB at Q4
- Qwen3-Coder 30B-A3B (MoE): only 3.3B active params, fast inference, fits easily
- Qwen2.5-Coder 32B remains the FIM (autocomplete) king: 92.7% HumanEval
Honest Assessment: Local vs Claude Code
Nothing local approaches Claude Opus 4.6 quality for complex multi-file agentic coding. These 32B models are competitive with GPT-4o — a tier below Claude Sonnet, two tiers below Opus. Best strategy: use local models for routine tasks, save Claude credits for hard problems.
GPU Candidates Compared
| Feature | Tesla P40 | RTX A2000 12GB | Tesla T4 | RTX A4000 |
|---|---|---|---|---|
| Architecture | Pascal (2016) | Ampere (2020) | Turing (2018) | Ampere (2020) |
| VRAM | 24GB GDDR5 | 12GB GDDR6 | 16GB GDDR6 | 16GB GDDR6 |
| Tensor Cores | No | Yes | Yes | Yes |
| TDP | 250W | 70W | 70W | 140W |
| Compute Capability | 6.1 | 8.6 | 7.5 | 8.6 |
| Cooling | Passive (server fans) | Active (blower) | Passive (server fans) | Active (single-slot) |
| Aux Power Required | Yes (8-pin) | No (bus-powered) | No (bus-powered) | Yes (6-pin) |
| PCIe | Gen3 x16 | Gen4 x16 | Gen3 x16 | Gen4 x16 |
| Can run 32B Q4? | Yes (tight) | No (12GB) | No (16GB) | No (16GB) |
Current Prices (March 22, 2026)
| GPU | VRAM | Price Range | Best Deals | Notes |
|---|---|---|---|---|
| Tesla P40 | 24GB | $150-320 | Newegg refurb $219-270; eBay used $150-200 | Best VRAM/$ ratio |
| RTX A2000 12GB | 12GB | $250-535 | eBay used ~$250-350; one listing at $490 | New retail ~$535 |
| Tesla T4 | 16GB | $150-350 | eBay used $150-250 | Great power efficiency |
| RTX A4000 | 16GB | $700-750+ | eBay used ~$700; new $720+ | Too expensive for this build |
Sources: eBay active listings, Newegg, Lowpi.com, Pangoly price history (all checked March 2026)
Performance Benchmarks
LLM Inference (tokens/sec via Ollama)
| Model | A2000 12GB | P40 24GB | RTX 3090 24GB (ref) |
|---|---|---|---|
| qwen2.5:14b | ~21 t/s | ~17 t/s | — |
| qwen2.5-coder:32b Q4_K_M | Won't fit | ~5-12 t/s (est.) | ~37-40 t/s |
| llama3.2:3b-Q4 | ~60 t/s | ~48 t/s | — |
P40 reality check for 32B: The model barely fits (~22GB for weights), leaving only ~2GB for KV cache. Context window will be severely limited. Expect 5-12 tok/s — usable for short prompts, painful for long sessions.
Image Generation (ComfyUI)
| GPU | SDXL 20 steps | Notes |
|---|---|---|
| P40 | ~49 seconds | Requires --force-fp32 flag |
| A2000 | ~16 seconds | ~3x faster than P40 |
| RTX 4090 | ~3 seconds | Reference (16x faster than P40) |
Rack Server Compatibility
RTX A2000 in R720/R730
- Physical fit: Yes. Dual-slot, low-profile, 167mm length
- Power: 70W bus-powered, no aux cable needed. Must use 75W slots (slots 4-7 on R720)
- Cooling: Blower-style fan exhausts out bracket — ideal for rack airflow
- Requirement: Dual CPUs needed for GPU PCIe slots
- Confirmed working in Dell R740XD (similar architecture)
- Third-party single-slot cooler available from n3rdware for tighter fits
Tesla P40 in R720/R730
- Physical fit: Yes. Full-length, single-slot, designed for rack servers
- Power: 250W, requires 8-pin aux power. Needs GPU enablement kit (power cables + riser)
- Cooling: Passive — relies on server chassis fans (which R720/R730 have)
- Requirement: Dual CPUs, redundant 1100W PSUs recommended
- Natively supported in these servers
- Users report power-limiting to 140W with little performance impact
R720 vs R730
- R720: PCIe Gen2 (not a bottleneck for LLM inference, which is VRAM-bound)
- R730: PCIe Gen3, generally preferred
- Both support up to 2x double-wide or 4x single-wide GPUs
Setup Options Analysis
Option A: Dual P40 (~$400-500)
- P40 #1: Qwen2.5-Coder 32B (tight fit, 5-12 tok/s, limited context)
- P40 #2: Image gen with ComfyUI (
--force-fp32, ~49s/image SDXL) - Or: split 32B model across both P40s via tensor parallelism for better speed
- Both are native rack server GPUs (passive, designed for R720/R730)
- Total power: ~500W GPU (can power-limit to ~280W)
- Best value for 32B + image gen
Option B: P40 + A2000 (~$550-810)
- P40 for 32B coding model (24GB, tight but works)
- A2000 for image gen (3x faster than P40, tensor cores, 70W)
- Total power: ~320W GPU
- A2000 at $490 is overpriced — shop for $250-350 range
- Better image gen speed vs Option A
Option C: Single P40 (~$200-300)
- Run 32B coding model OR image gen (not both simultaneously)
- 32B model leaves no room for anything else in VRAM
- Swap between tasks by unloading/loading models
- Cheapest entry point, upgrade later
Option D: P40 (coding) + T4 (image gen) (~$400-550)
- P40: 32B coding model (24GB)
- T4: Image gen with tensor cores, 16GB, 70W, passive
- T4 is faster than P40 for image gen (Turing tensor cores, native FP16)
- Both passive-cooled = true rack-native
- Good balance of price and image gen speed
Recommendations
If 32B coding quality is the priority: Dual P40 ($400-500)
Two P40s give you 48GB total. Run 32B coding model on one, image gen on the other. Or split the model across both for faster inference. Both fit natively in R720/R730.
If image gen speed matters equally: P40 + T4 ($400-550)
P40 for 32B coding, T4 for image gen. Both passive-cooled, both rack-native. T4 has tensor cores + FP16 for much faster image gen than P40.
Cheapest possible: Single P40 ($200-300)
Run 32B coding model with limited context. Swap to image gen when needed. Upgrade to dual-GPU later.
Configuration Notes for local-ai stack
For 32B models on P40
# In Ollama environment
OLLAMA_NUM_GPU=999
OLLAMA_NUM_CTX=4096 # Keep context small to fit in remaining VRAM
OLLAMA_KEEP_ALIVE=24h
# Pull the right quantization
ollama pull qwen2.5-coder:32b-instruct-q4_K_M
For dual-GPU setup
# Assign GPU 0 to Ollama (coding), GPU 1 to InvokeAI (image gen)
# In docker-compose.yml for Ollama:
CUDA_VISIBLE_DEVICES=0
# In docker-compose.yml for InvokeAI:
CUDA_VISIBLE_DEVICES=1
For P40 image gen
# InvokeAI
INVOKEAI_PRECISION=float32
# ComfyUI launch args
--force-fp32
For newer coding models (2026)
# Qwen 3.5 27B — fits easily on 24GB at Q4, better quality than 2.5
ollama pull qwen3.5:27b
# Qwen3-Coder 30B-A3B MoE — fast inference, agentic coding
ollama pull qwen3-coder:30b
Sources
- NVIDIA RTX A2000 Datasheet
- Lenovo ThinkSystem RTX A2000 Product Guide
- n3rdware Single-Slot A2000 Cooler
- Dell R730 Owner's Manual - Expansion Cards
- Dell R720 Owner's Manual - Expansion Cards
- ComfyUI GPU Benchmarks Discussion
- ComfyUI P40 FP32 Issue
- How to Use Tesla P40 Guide
- Build a Local LLM Server Under $1000
- Best Budget GPUs for AI 2026
- Tesla P40 for Local LLMs 2026
- Best Local LLMs for 24GB VRAM 2026
- Best Coding Models 2026
- Ollama VRAM Requirements Guide
- Local LLMs That Can Replace Claude Code
- Qwen2.5-Coder 32B on Ollama
- Qwen2.5-Coder 32B HuggingFace Discussion