Handle old/low-VRAM GPUs and document nvidia-container-toolkit requirement
GPU tier table extended: ultra ≥16 GB → SDXL (unchanged) high 8-16 GB → SDXL (unchanged) medium 4-8 GB → SD 2.x (unchanged) legacy 2-4 GB → SD 1.5 (~1.7 GB fp16) ← new: GTX 970/1060/RX 580 etc. minimal <2 GB → SD 1.5 + sequential CPU offload ← new: very old/integrated GPUs gpu_detect.py: - Detects CUDA compute capability (CC); fp16 disabled for CC < 6.0 (pre-Pascal) - GpuInfo gains compute_capability and warnings fields - _make_warnings() emits human-readable warnings for low VRAM and old CC - model tier fallback updated from 'low' to 'legacy' local_diffusion.py: - minimal/legacy tiers use enable_sequential_cpu_offload() + enable_attention_slicing(1) - target resolution per tier: ultra/high=1024, medium=768, legacy/minimal=512 - .to(device) skipped when sequential CPU offload is active gpu_status.py: - Response now includes compute_capability and warnings docker-compose.gpu.yml: - Full nvidia-container-toolkit install instructions in header comment - nvidia-docker2 (legacy) fallback documented as comment block inline - AMD ROCm swap-in instructions added - GPU tier table documented in header scripts/gpu_setup.py: - Prints compute capability, fp16 status, tier, and model selection at startup - Prints per-tier warnings (old CC, low VRAM) https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
This commit is contained in:
+5
-2
@@ -39,10 +39,13 @@ async def lifespan(app: FastAPI):
|
||||
):
|
||||
from app.services.gpu_detect import get_cached_gpu_info
|
||||
info = get_cached_gpu_info()
|
||||
cc_str = f" | CC={info.compute_capability}" if info.compute_capability else ""
|
||||
print(
|
||||
f"[gpu] {info.device_name} | {info.vram_gb:.1f} GB | tier={info.tier} | "
|
||||
f"backend={info.backend}"
|
||||
f"[gpu] {info.device_name} | {info.vram_gb:.1f} GB{cc_str} | "
|
||||
f"tier={info.tier} | fp16={info.fp16}"
|
||||
)
|
||||
for w in info.warnings:
|
||||
print(f"[gpu] ⚠ {w}")
|
||||
if settings.auto_download_models:
|
||||
# Download model weight files to disk cache in background so first
|
||||
# user request loads from local disk instead of the internet.
|
||||
|
||||
Reference in New Issue
Block a user