Files
local-ai/docs/gpu-setup-research.md
T
Claude 8646dca2bb Fix RTX 8000 pricing: $2,000-2,900 realistic, not $750
The $750 listing was a single outlier. Actual market price for
Quadro RTX 8000 Passive 48GB is $2,000-2,900 on eBay. Updated
all recommendations and cost comparisons accordingly.

https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
2026-03-22 14:32:28 +00:00

236 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GPU Setup Research: Rack Server AI Workloads
*Last updated: March 22, 2026*
## Goal
Cost-efficient rack-mountable GPU setup for:
1. **LLM coding inference** — Run 32B+ parameter coding models with maximum context windows
2. **Image generation** — ComfyUI / InvokeAI with Stable Diffusion SDXL / Flux
Target servers: Dell R720/R730 or HP DL380 equivalent (2U rack)
## Why 48GB VRAM is the Right Target
### The Problem with 24GB
32B coding models at Q4_K_M quantization use ~20GB of weights, leaving only ~4GB for KV cache
on a 24GB card. This severely limits context window size — the key ingredient for complex
coding sessions where the model needs to understand your entire codebase.
### What 48GB Unlocks
- **32B models at higher quantization** (Q6_K/Q8_0) = better output quality
- **28GB+ free for KV cache** = massive context windows (32K+ tokens)
- **70B models** in aggressive quantization (~12 t/s but functional)
- **Simultaneous model loading** — coding model + image gen model at once
- Room for future larger models without hardware changes
## Best Local Coding Models (2026)
| Model | Size at Q4_K_M | Quality | Notes |
|-------|---------------|---------|-------|
| **Qwen2.5-Coder 32B** | ~20GB | 73.7 Aider (≈ GPT-4o) | FIM king, 92.7% HumanEval |
| **Qwen 3.5 27B** | ~16GB | 72.4% SWE-bench (ties GPT-5 mini) | 262K context, multimodal |
| **Qwen3-Coder 30B-A3B** (MoE) | ~18GB | #1 SWE-rebench (64.6%) | Only 3.3B active, very fast |
| **Qwen3-Coder-Next 80B** (MoE) | needs 64GB+ RAM offload | Beats Claude Opus 4.6 on SWE-rebench | Hybrid attention, 256K context |
### Honest Assessment: Local vs Claude Code
Nothing local approaches Claude Opus 4.6 quality for complex multi-file agentic coding.
These 32B models are competitive with **GPT-4o** — a tier below Claude Sonnet, two tiers below Opus.
Best strategy: use local models for routine tasks, save Claude credits for hard problems.
## 48GB GPU Market (March 22, 2026 — Real Prices)
| GPU | Arch | Used Price | TDP | Cooling | Tensor Cores | Mem BW |
|-----|------|------------|-----|---------|--------------|--------|
| **Quadro RTX 8000** | Turing (2018) | **$2,0002,900** | 260W | Passive variant | Yes (576) | 672 GB/s |
| **A40** | Ampere (2020) | **~$5,050+** | 300W | Passive | Yes (336 3rd-gen) | 696 GB/s |
| **RTX A6000** | Ampere (2020) | **~$5,400+** | 300W | Active (blower) | Yes (336 3rd-gen) | 768 GB/s |
| **L40** | Ada (2022) | **~$6,500+** | 300W | Passive | Yes (568 4th-gen) | 864 GB/s |
| **RTX 6000 Ada** | Ada (2022) | **~$6,500+** | 300W | Active | Yes (568 4th-gen) | 960 GB/s |
Sources: eBay active/sold listings, GPUPoet price tracking, Pangoly, CamelCamelCamel (all March 2026)
Note: One outlier RTX 8000 listing at ~$750 exists but is not representative of the market.
### Cheapest 48GB Option: Quadro RTX 8000 Passive ($2,0002,900)
The RTX 8000 is still the cheapest 48GB card — roughly half the price of an A40 and a
third of an A6000. The passive variant is purpose-built for rack servers — no fan, relies
on chassis airflow, designed for 24/7 operation in 2U/4U systems.
Key advantages over the P40:
- **48GB vs 24GB** — room for models + massive context
- **Has Tensor Cores** (576 Turing) — native FP16, no `--force-fp32` hacks for image gen
- **NVLink support** — pair two for 96GB combined (100 GB/s bidirectional)
- 10W idle power draw
### Cost Reality Check
At $2,0002,900 the RTX 8000 is a significant investment. The key question: is unified
48GB VRAM worth 46x the cost of dual P40s ($400500)?
**Yes, if** you need large context windows (32K+) for complex coding — KV cache can't
be split across two GPUs without NVLink (which P40s don't have).
**No, if** you're mostly doing short-prompt coding tasks and image gen — dual P40s give
you 48GB total (split) at a fraction of the cost, and each card can handle its own workload.
## RTX 8000 Performance Benchmarks
### LLM Inference (Exllama, 5.0 bpw quantization)
| Model | Context | Prompt Processing | Generation |
|-------|---------|-------------------|------------|
| Qwen3 30B-A3B (MoE) | 8K | 950 t/s | **34 t/s** |
| Qwen3 30B-A3B (MoE) | 16K | 673 t/s | **21 t/s** |
| Qwen3 30B-A3B (MoE) | 32K | 345 t/s | **11 t/s** |
| Llama 3.3 70B | short | 36 t/s | **13 t/s** |
| Llama 3.1 8B | — | — | **72 t/s** |
### Compared to P40 (24GB)
| Metric | P40 (24GB) | RTX 8000 (48GB) |
|--------|-----------|-----------------|
| **Used price** | **$150320** | **$2,0002,900** |
| 32B model fit | Barely (~2GB free) | Comfortable (~28GB free) |
| 32B generation speed | ~5-12 t/s (est.) | ~20-34 t/s |
| Max practical context | ~4K tokens | **32K+ tokens** |
| Image gen (SDXL) | ~49s (`--force-fp32`) | Faster (native FP16) |
| Rack server ready | Yes (passive) | Yes (passive variant) |
### Image Generation
The RTX 8000 has Turing Tensor Cores with native FP16 support. Unlike the P40, it does NOT
need `--force-fp32` workarounds. Image gen performance is significantly better than the P40,
though still behind Ampere/Ada cards.
## 24GB GPU Options (Previous Research — Still Valid for Tighter Budgets)
| GPU | VRAM | Price Range | Best Deals | Notes |
|-----|------|-------------|------------|-------|
| **Tesla P40** | 24GB | $150-320 | Newegg refurb $219-270; eBay used $150-200 | Best VRAM/$ at 24GB |
| **RTX A2000 12GB** | 12GB | $250-535 | eBay used ~$250-350; one listing at $490 | Can't run 32B models |
| **Tesla T4** | 16GB | $150-350 | eBay used $150-250 | Great power efficiency |
| **RTX A4000** | 16GB | $700-750+ | eBay used ~$700; new $720+ | Too expensive for 16GB |
## Rack Server Compatibility
### Quadro RTX 8000 Passive in R720/R730
- **Physical fit**: Full-length, dual-slot — fits in GPU riser slots
- **Power**: 260W, requires 8-pin aux power + GPU enablement kit
- **Cooling**: Passive — relies on server chassis fans (same as P40)
- **Requirement**: Dual CPUs, redundant 1100W PSUs recommended
- **NVLink**: Can pair two RTX 8000s for 96GB combined VRAM
- Very similar physical/power requirements to the Tesla P40
### RTX A2000 in R720/R730
- **Physical fit**: Yes. Dual-slot, low-profile, 167mm length
- **Power**: 70W bus-powered, no aux cable needed. Must use 75W slots (slots 4-7 on R720)
- **Cooling**: Blower-style fan exhausts out bracket — ideal for rack airflow
- **Requirement**: Dual CPUs needed for GPU PCIe slots
- **Confirmed working** in Dell R740XD (similar architecture)
### Tesla P40 in R720/R730
- **Physical fit**: Yes. Full-length, single-slot, designed for rack servers
- **Power**: 250W, requires 8-pin aux power. Needs GPU enablement kit
- **Cooling**: Passive — relies on server chassis fans
- **Requirement**: Dual CPUs, redundant 1100W PSUs recommended
- **Natively supported** in these servers
### R720 vs R730
- R720: PCIe Gen2 (not a bottleneck for LLM inference, which is VRAM-bound)
- R730: PCIe Gen3, generally preferred
- Both support up to 2x double-wide or 4x single-wide GPUs
## Recommended Setups
### If budget allows ($2,0002,900): RTX 8000 Passive
Single card handles both coding and image gen. 48GB VRAM fits 32B models with massive
context windows (32K+). Passive cooling is rack-native. Tensor cores handle FP16 image gen
properly. One card, one slot, simple setup. The premium buys you unified VRAM = big context.
### If budget allows + dedicated image gen ($2,3003,250): RTX 8000 + A2000
RTX 8000 for coding with full 48GB dedicated to LLM context.
A2000 for image gen (3x faster than Turing, 70W, bus-powered, blower cooled).
Best separation of concerns — no model swapping needed.
### Best value ($400500): Dual P40
Two P40s for 48GB total, but split across cards (can't combine for one model without
NVLink, which P40s lack). One for 32B coding (tight fit, ~4K context), one for image gen
(slow, needs --force-fp32). **5x cheaper than RTX 8000** but with significant context limitations.
### Cheapest entry ($200300): Single P40
Run 32B coding model with very limited context (~4K tokens). Swap to image gen when needed.
Good for testing whether local LLM coding works for your workflow before investing more.
## Configuration Notes for local-ai stack
### For 48GB RTX 8000
```bash
# Ollama — take advantage of the full 48GB
OLLAMA_NUM_GPU=999
OLLAMA_NUM_CTX=32768 # Large context window — 48GB can handle it
OLLAMA_KEEP_ALIVE=24h
# Pull best coding models
ollama pull qwen2.5-coder:32b-instruct-q4_K_M # ~20GB, leaves 28GB for context
ollama pull qwen3.5:27b # ~16GB at Q4, even more context room
ollama pull qwen3-coder:30b # MoE, very fast inference
# Higher quantization for better quality (48GB allows this)
# Look for Q6_K or Q8_0 variants on Ollama for better output quality
```
### For 32B models on P40 (24GB — tight fit)
```bash
OLLAMA_NUM_GPU=999
OLLAMA_NUM_CTX=4096 # Keep context small to fit in remaining VRAM
OLLAMA_KEEP_ALIVE=24h
ollama pull qwen2.5-coder:32b-instruct-q4_K_M
```
### For dual-GPU setup (RTX 8000 + A2000 or P40 + anything)
```bash
# Assign GPU 0 to Ollama (coding), GPU 1 to InvokeAI (image gen)
# In docker-compose.yml for Ollama:
CUDA_VISIBLE_DEVICES=0
# In docker-compose.yml for InvokeAI:
CUDA_VISIBLE_DEVICES=1
```
### For image gen on P40 (no tensor cores)
```bash
# InvokeAI
INVOKEAI_PRECISION=float32
# ComfyUI launch args
--force-fp32
```
### For image gen on RTX 8000 / A2000 / T4 (has tensor cores)
```bash
# InvokeAI — native FP16 works fine
INVOKEAI_PRECISION=float16
# ComfyUI — no special flags needed
```
## Sources
- [Quadro RTX 8000 for Local LLMs — Hardware Corner](https://www.hardware-corner.net/guides/quadro-rtx-8000-for-llm/)
- [RTX 8000 Passive — Network Outlet](https://networkoutlet.com/blogs/articles/nvidia-quadro-rtx-8000-48gb-passive-cooling-powering-ai-rendering-server-workloads)
- [LLM Benchmarks on Turing/Ampere GPUs — Stefandroid](https://blog.stefandroid.com/2025/06/02/benchmark-llm-performance-nvidia-gpus.html)
- [NVIDIA A40 Price Tracking — GPUPoet](https://gpupoet.com/gpu/learn/card/nvidia-a40)
- [NVIDIA L40 Price Tracking — GPUPoet](https://gpupoet.com/gpu/learn/card/nvidia-l40)
- [RTX A6000 Price History — CamelCamelCamel](https://camelcamelcamel.com/product/B09BDH8VZV)
- [RTX A6000 Price History — Pangoly](https://pangoly.com/en/price-history/pny-nvidia-quadro-rtx-a6000)
- [NVIDIA RTX A2000 Datasheet](https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/rtx-a2000/nvidia-rtx-a2000-datasheet-1987439-r5.pdf)
- [Dell R730 Owner's Manual — Expansion Cards](https://www.dell.com/support/manuals/en-us/poweredge-r730/r730_ompublication/expansion-card-installation-guidelines)
- [Dell R720 Owner's Manual — Expansion Cards](https://www.dell.com/support/manuals/en-us/poweredge-r720/720720xdom/expansion-card-installation-guidelines)
- [ComfyUI GPU Benchmarks Discussion](https://github.com/Comfy-Org/ComfyUI/discussions/2970)
- [ComfyUI P40 FP32 Issue](https://github.com/Comfy-Org/ComfyUI/issues/4363)
- [Best Local LLMs for 24GB VRAM 2026](https://localllm.in/blog/best-local-llms-24gb-vram)
- [Best Coding Models 2026](https://localvram.com/en/guides/best-coding-models/)
- [Ollama VRAM Requirements Guide](https://localllm.in/blog/ollama-vram-requirements-for-local-llms)
- [Local LLMs That Can Replace Claude Code](https://agentnativedev.medium.com/local-llms-that-can-replace-claude-code-6f5b6cac93bf)
- [7 Local LLM Families to Replace Claude/Codex](https://agentnativedev.medium.com/7-local-llm-families-to-replace-claude-codex-for-everyday-tasks-25ba74c3635d)
- [Qwen2.5-Coder 32B on Ollama](https://ollama.com/library/qwen2.5-coder:32b-instruct-q4_K_M)
- [Qwen3-Coder — How to Run Locally](https://unsloth.ai/docs/models/qwen3-coder-how-to-run-locally)