Adds AI_PROVIDER=local_gpu — a fully self-contained GPU inference engine
using HuggingFace Diffusers that requires zero InvokeAI/ComfyUI setup.
All existing providers (InvokeAI, ComfyUI, OpenAI, Replicate) remain intact
and can be mixed with local GPU via per-operation overrides.
New features:
- GPU auto-detection (CUDA/NVIDIA, MPS/Apple Silicon, CPU fallback)
- VRAM-tiered model selection:
ultra ≥16 GB → SDXL inpaint + SDXL base
high 8-16 GB → SDXL inpaint + SDXL base
medium 4-8 GB → SD 2.x inpaint + SD 2.1
low <4 GB → SD 2.x (small)
- Auto-download model weights to HuggingFace disk cache at startup
(background task; first request loads from local disk, not internet)
- LRU pipeline cache evicts oldest GPU pipeline when VRAM limit reached
- Per-operation model overrides via HF_MODEL_INPAINT / HF_MODEL_TXT2IMG etc.
- Optional HF_TOKEN for gated/private HuggingFace models
New files:
- backend/app/services/gpu_detect.py — GPU detection + tier/model mapping
- backend/app/services/local_diffusion.py — Diffusers provider + LRU cache
- backend/app/routers/gpu_status.py — GET /api/gpu/status, POST /api/gpu/prefetch
- backend/requirements.gpu.txt — Diffusers ecosystem deps (GPU only)
- docker-compose.gpu.yml — NVIDIA GPU compose (one-command startup)
- Dockerfile.gpu — pytorch/pytorch:2.1.0-cuda12.1 base image
- scripts/gpu_setup.py — Startup GPU info logger
Modified:
- backend/app/config.py — local_gpu settings added
- backend/app/services/remote_provider.py — local_gpu registered as provider
- backend/app/routers/ai_tools.py — /api/config exposes GPU tier + caps
- backend/app/main.py — GPU router + background prefetch task
- backend/entrypoint.sh — runs gpu_setup.py at container start
- .env.example — local_gpu documented as first option
Quick start with GPU:
docker compose -f docker-compose.gpu.yml up --build
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
25 lines
812 B
Plaintext
25 lines
812 B
Plaintext
# =============================================================================
|
|
# GPU / Local Diffusion dependencies
|
|
# Install alongside requirements.txt when running with AI_PROVIDER=local_gpu
|
|
#
|
|
# Usage:
|
|
# pip install -r requirements.txt -r requirements.gpu.txt
|
|
#
|
|
# These are pre-installed in Dockerfile.gpu; optional in the standard image.
|
|
# =============================================================================
|
|
|
|
# HuggingFace Diffusers ecosystem
|
|
diffusers>=0.27.0
|
|
transformers>=4.38.0
|
|
accelerate>=0.27.0
|
|
huggingface-hub>=0.21.0
|
|
safetensors>=0.4.0
|
|
|
|
# Required by SDXL pipelines
|
|
invisible-watermark>=0.2.0
|
|
omegaconf>=2.3.0
|
|
|
|
# xformers — further reduces VRAM usage on CUDA (install separately, version must
|
|
# match your PyTorch/CUDA; leave out if unsure and use attention_slicing instead)
|
|
# xformers
|