Lets SAM, U2Net, BEN2, and BiRefNet-HR (optionally SDXL via --sdxl) be
downloaded outside Docker into ./data/, which is already bind-mounted
into the GPU container — so a blocked container network no longer blocks
first-run setup. Reuses the existing dual-mode download_sam_model.py and
download_u2net_model.py as-is. For the HuggingFace Hub models, sets
HF_HOME (rather than --cache-dir) so the host-side cache layout matches
the container's default ~/.cache/huggingface resolution exactly, avoiding
a path-nesting mismatch between the two.
Wired into install-local-gpu.sh's completion banner and
bring-up-local-gpu.sh's header, and referenced from the relevant README
troubleshooting sections and the hf_cache bind-mount comment in
docker-compose.gpu.yml.
BEN2 becomes the new default local backend (clean cutouts, strong on
hair/fur edges), with BiRefNet-HR available as a high-res/print
alternate and U2Net kept as the lightweight fallback. Both are
MIT-licensed and download weights from HuggingFace on first use
(cached via the existing hf_cache bind mount), unlike U2Net/SAM which
need an explicit download script.
- config: new BG_REMOVAL_MODEL setting (default "ben2")
- tools.py: remove-background-base64 now tries local backends in
order (request.model override > BG_REMOVAL_MODEL > ben2/u2net),
falling back to rembg's birefnet-general session as a last resort
- requirements.gpu.txt / Dockerfile.gpu: add ben2 + transformers deps
needed for the new backends, with a build-time smoke test for ben2
- frontend: model dropdown in the Remove Background dialog, threaded
through api.js to the new request field
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ro4PwQKvSc3CH19LSN21Ht
Commit 5c85987 appended a second copy of the toolkit-install/DNS-fix/
GPU-verify steps instead of inserting the new prerequisite-checks
section once. Removed the duplicate, keeping the more informative
GPU-check failure message from the second copy.
U2Net only ever downloaded lazily on the first Remove Background click,
unlike SAM which retries on every container start. If that one attempt
failed (DNS/firewall) the model was never fetched again, surfacing as
"No background removal method available. Install u2net or rembg."
Mirrors the existing SAM auto-download/AUTO_DOWNLOAD_SAM pattern for
U2Net, and documents manual host-side recovery in the README.
Also deletes backend/app/services/u2net_model.py (hand-written
U2NET/U2NETP PyTorch classes) — unused since tools.py switched to
cv2.dnn.readNetFromONNX for background removal.
- Backend: POST /api/image/replace-subject
Uses rembg to extract the subject from a source photo, scales it to fit
the selection mask bounding box (or canvas centre when no selection is
active), applies a partial LAB color transfer (blend=0.45) so the
subject's lighting matches the background, then composites the result.
Also adds POST /api/image/extract-subject for standalone subject extraction.
- Frontend: "Replace with subject from file" button in the selection panel
(selection_actions.js) — opens a native file picker so no clipboard API
or HTTPS is required. Calls the new endpoint with the current SAM mask.
- Frontend: Image > Replace Subject (AI)... menu entry backed by
modules/image/replace_subject.js — a full dialog with file picker,
thumbnail preview, and "match background lighting" toggle. Works with
or without a prior Smart Select; if a selection exists it confines the
subject to that region.
https://claude.ai/code/session_01UtrvbisMp1yGu6PqeLmrFr
Checks Docker installed + daemon running, NVIDIA driver (nvidia-smi),
and curl. Prints specific install commands for each missing dependency
and exits early with a clear summary rather than failing partway through.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
The iptables DOCKER-USER rule is lost on reboot; the script re-applies
it each run, checks for duplicates, and is silently skipped on macOS/WSL.
Default behaviour (no args): docker compose up -d --build.
All docker compose subcommands can be passed as args (logs, down, etc.).
README Quick Start and Updates sections now reference ./start-gpu.sh.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
iptables fix is simpler (one command, no 13 GB download) and does not
affect container isolation — adds clarifying note so users understand
it only restores Docker's default outbound DNS behaviour.
Docker helper container kept as fallback option.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
Replaces 'pip install huggingface-hub' (breaks on PEP 668 / Debian 12+)
with a docker run --rm python:3.11-slim one-liner that downloads directly
into ./data/hf_cache without touching host Python packages.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
Build / pip layer fixes:
- Add BUILDID ARG to Dockerfile.gpu; pass from docker-compose.gpu.yml build args
so pip layers can be force-busted without --no-cache:
BUILDID=$(date +%s) docker compose -f docker-compose.gpu.yml up --build
Model download (DNS-blocked environments):
- Change HF model cache from named volume to ./data/hf_cache bind mount
so models can be pre-downloaded on the host (no rebuild needed)
- Remove now-unused hf_model_cache named volume
- README: add iptables fix + huggingface-cli offline download instructions
Error handling improvements:
- ai_edit_region: catch ConnectError/Errno-3 → return 503 with exact fix commands
- _require_remote: give actionable message when local_gpu provider fails to load
- _build_provider: catch AttributeError (torch.xpu from wrong diffusers) not just ImportError
- local_diffusion.py: fix docstring to reflect <0.29.0 pin
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
- Pin diffusers to >=0.28.0,<0.29.0 to avoid AttributeError on torch.xpu
(diffusers 0.29+ requires PyTorch 2.4 but base image ships 2.1.2)
- Pin transformers to <4.40.0 to match
- README: add SAM offline download troubleshooting for DNS-blocked containers
- SelectionActions panel: 'Selection ready' title + subtitle makes clear
nothing has fired yet; AI actions show inline description (not just tooltip);
section labels 'AI Actions' / 'Classic Tools'; scale hint updates live;
Enter key submits AI edit prompt; _actionCard hover border for clickability
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
1.1.1.1 uses a different anycast route than 8.8.8.8 so may reach the
container even when Google's servers don't. Falls back to Google if
Cloudflare is also unreachable. If all fail (Errno -3), port 53 UDP is
blocked at the Docker bridge level — comment has the iptables fix.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
ai_tools.py: add 'from io import BytesIO' — scale_selection and paste
endpoints used BytesIO directly but it was only imported locally in one
unrelated function, causing NameError on every scale call.
provider-badge.js: redesign badge to fit the narrow (~40px) left toolbar.
Was rendering 'GPU · sdxl_offload · Quadro RTX 3000' inline which wrapped
into multiple lines covering tool icons. Now shows a status dot + short
label (SDXL / OAI / Rep…) with all details moved to the hover tooltip.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
selection_actions.js uses named exports only (no default export).
load_plugins() was calling new classObj.default(ctx) on every file in tools/,
causing TypeError which aborted render_main_tools() -> render_main_gui() ->
leaving menu and toolbar blank. Added null guard and per-file try/catch,
matching the defensive pattern already added to load_modules().
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
Without this, any constructor throw in any module stops load_modules() mid-loop
and render_main_gui() (which builds the menu) never runs — resulting in a blank
page with no menu. The failing module is now logged to the browser console
instead of silently aborting startup.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
main.py referenced info.vram_gb but the field is info.vram_total_gb.
This crashed the FastAPI lifespan hook on every startup when
AI_PROVIDER=local_gpu, causing a restart loop.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
Old README described the original React/Fabric.js UI and listed
"No local GPU inference" as a non-goal. Updated to reflect:
- miniPaint-based editor with SAM brush selection
- GPU quick-start (nvidia-container-toolkit prereqs, docker-compose.gpu.yml)
- Cloud API quick-start
- GPU tier auto-selection table (FLUX/SDXL/SD by VRAM)
- Full feature list (selection actions, print tools, progress bars)
- Correct clone URL and update commands
- Troubleshooting section
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
Backend:
- local_diffusion.py: add _make_step_cb() that writes step/total_steps/
progress into _states on every diffusers callback_on_step_end; wired into
txt2img, inpaint, img2img with TypeError fallback for older diffusers
- ai_tools.py: GET /api/generate/progress SSE endpoint — streams _states
as JSON array every 200ms so clients get live denoising step counts
Frontend:
- progress_overlay.js: add connectProgressSSE(pipeType, baseUrl) /
disconnectProgressSSE() — opens EventSource, maps step/total_steps
to bar percentage (0→85% during denoising, 85→100 for decode/place)
- text_to_image.js: connect SSE before POST, disconnect on done/error
- selection_actions.js: connect SSE for AI edit / asymmetry operations
Result: for local GPU, progress bar shows "Step 12 / 30" with exact fill;
for remote providers and upscale (no step callbacks), shimmer animates.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
- New progress_overlay.js: animated fullscreen overlay with shimmer bar,
fake progress creep, Esc-to-cancel, used by all slow AI operations
- text_to_image.js: allow local_gpu provider (was incorrectly blocked);
show provider/model/VRAM info in dialog; show estimated generation time;
use progress overlay during generation
- upscale.js: replace alertify.message with progress overlay (90s estimate
for AI upscale, 10s for Lanczos)
- frame_fit.js: progress overlay for extend mode (AI outpaint ~45s)
- print_prepare.js: progress overlay for full upscale+frame chain (~2 min)
- selection_actions.js: progress overlay for all AI region edits
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
- Add 18x24" to FRAME_SIZES in backend and frontend (frame_fit.js)
- Add 200 DPI option to frame_fit dialog (adequate for large-format prints)
- Add 18x24 portrait/landscape at 200 and 300 DPI to Canvas Size presets (size.js)
- New /api/print/prepare endpoint: chains AI upscale to target DPI then frame-fit
in one server-side call (avoids round-tripping a large upscaled image)
- New print_prepare.js module: "Prepare for Print" dialog with per-frame quality
assessment (current effective DPI, needed upscale factor, AI vs Lanczos note)
- Add "Prepare for Print..." to Image menu above "Fit to Frame..."
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
Backend (3 new endpoints under /api/image/):
- POST /api/image/scale-selection — scale selected object by any % in-place;
LaMa/OpenCV fills the exposed gap so the scene looks natural
- POST /api/image/ai-edit-region — AI redraws the masked region via the
configured inpaint provider (local_gpu / InvokeAI / ComfyUI / OpenAI)
- POST /api/image/paste-into-selection — scales clipboard image to fit the
selection bounding box, masks it to the selection shape, composites result
Frontend (selection_actions.js + tool integration):
- New SelectionActions panel: fixed bottom-center HUD that appears
automatically after every SAM selection (click or paint)
- Panel actions: Scale by % (default 3%), Make less symmetrical (AI),
custom AI Edit prompt, Replace with clipboard, Copy/Cut to layer, Erase
- Both smart_select.js and brush_select.js updated to show the panel,
add updateLayerWithResult(), and hide panel on clearSelection/on_leave
- brush_select: offerFloatSelection() replaced with richer action panel
Real-world workflows now supported in one click after painting over object:
"Make this 3% bigger" → scale-selection (LaMa fills gap)
"Make this less symmetrical" → ai-edit-region with asymmetry prompt
"Replace this with what I copied" → paste-into-selection
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
Model selection:
- Add sdxl_offload tier (eff_vram ≥ 4.0 GB) for GTX 1060 6GB and Quadro
6GB cards that were falling through to SD 2.1 despite SDXL fitting with
model_cpu_offload. Cards with 5.3 GB effective VRAM now get SDXL quality.
- Update _tier_label(), _caps(), _build_warnings() for new tier.
Frontend GPU display:
- api.js: add getGpuStatus() fetching /api/gpu/status
- capabilities.js: add getGpuStatus() export with own LRU cache;
refreshCapabilities() now also resets GPU status cache
- provider-badge.js: when AI_PROVIDER=local_gpu show green badge with
GPU name, tier, VRAM, CC, feature flags, and capabilities in tooltip.
Strip "NVIDIA GeForce" prefix so "GTX 1060 6GB" fits in badge.
- ai_provider_settings.js: add local_gpu to all provider dropdowns;
show GPU info panel (device, VRAM, CC, features, tier, model table per
operation) in the settings dialog when a GPU is detected.
https://claude.ai/code/session_01WVDg7amsy1TTtxvpku7bcM
The fitted image was processed and inserted correctly but the canvas
(config.WIDTH/HEIGHT) was never updated, so the result was clipped to
the original canvas size. Now wraps both new-layer and replace-layer
paths in Prepare_canvas_action + Update_config_action so the canvas
expands (or shrinks) to the target frame size automatically.
https://claude.ai/code/session_017wupXfpjuoSdAqVyak4fJf
Resize dialog:
- Units selector (pixels / inches) at the top; switching updates the
Width/Height placeholders and clears any partially entered values
- New "Crop to fill" checkbox: when both Width and Height are given,
scales the image with cover-fit (fills target without letterboxing)
then center-crops — subject looks the same at 5x7, 8x10, 11x14
- resize_layer/resize_gui now honour params.units instead of only
reading the global default_units setting
Global units persistence:
- Switching units in either Resize or Canvas Size saves the choice to
default_units so both dialogs open with the same unit next time
https://claude.ai/code/session_017wupXfpjuoSdAqVyak4fJf
- Units selector (pixels/inches) directly in the dialog; switching live-converts
the width/height fields so you can type e.g. 5 / 7 without mental math
- Six print-size presets at 300 DPI added to the Resolution dropdown:
5x7, 8x10, and 11x14 in both Portrait and Landscape orientations
- New "Resize & crop image" checkbox: when enabled, image layers are scaled
with cover-fit (fills the target canvas, no letterboxing) and center-cropped
so the subject stays proportionally the same across all three print sizes
https://claude.ai/code/session_017wupXfpjuoSdAqVyak4fJf
ai_edit.js:
- Extend Base_tools_class and add load() + default_events() so mouse
events actually wire up (was the root cause of tool not working)
- Use get_mouse_info() for coordinate mapping instead of manual
clientX/Y math — consistent with all other miniPaint tools
- Fix _mouseToImage() to use mouse.x/y (already in image coords)
and correct display-scale for overlay brush rendering
- mousedown/mousemove/mouseup now guard on config.TOOL.name
ai_edit.svg: new icon (brush + sparkle star) for the left toolbar
layout.css: add .ai_edit:after CSS rule for the icon
main.js:
- Collapse right-panel Colors section by default (respects saved cookie
so user preference persists)
- Mount compact foreground/background color swatches at the bottom of
the left toolbar; click either square to toggle the full color picker
open/closed; syncs live with config.COLOR every 250ms
https://claude.ai/code/session_01B58MaJCU1R6KwBDJCp8AfN
sklearn was not installed in the container, causing ModuleNotFoundError on
import of ai_tools.py and preventing the server from starting.
Replaced with a self-contained numpy k-means++ implementation:
- k-means++ seeding for better initial centers
- 20-iteration Lloyd's algorithm
- Same output: hex colors sorted by cluster frequency
No new dependencies required.
https://claude.ai/code/session_01B58MaJCU1R6KwBDJCp8AfN