The #1 friction point: users had to manually build a ComfyUI workflow,
export it in API format, then figure out which node IDs to enter in
Open WebUI. Now:
- comfyui-workflow-default.json: SD 1.5 text-to-image (512x512)
- comfyui-workflow-sdxl.json: SDXL text-to-image (1024x1024)
Upload either file directly into Admin → Settings → Images → Upload.
Then fill in the node ID mapping: Prompt=6, Model=4, Width/Height=5,
Steps/Seed=3.
Also fixes:
- Clear instructions to remove stale sk-1234 API key (ComfyUI doesn't
use API keys)
- Explains the localhost:1111 error (stale AUTOMATIC1111 config in
Open WebUI's database — just re-enter the ComfyUI URL)
- Updated setup steps to use setup-image-models.sh instead of manual
wget/docker cp
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Previously display-only: IMG_TIER/IMG_MODELS were detected but never
used. Now they drive actual behavior:
- Both setup scripts call setup-image-models.sh --auto after model
pulls, installing the right SD/SDXL/Flux model for the detected GPU
- INVOKEAI_PRECISION is now GPU-aware: bfloat16 for Ampere+ (compute
8.0+), float16 for Pascal+, auto fallback for older cards. Was
hardcoded to float16.
- local-ai-setup.sh now includes ComfyUI service + Open WebUI
integration (ENABLE_IMAGE_GENERATION=true, COMFYUI_BASE_URL) — was
completely missing, only laptop_full_setup.sh had it
- Added ComfyUI port 8188 to UFW firewall rules in local-ai-setup.sh
- Added comfyui-data/comfyui-output directories to mkdir loop
- Updated start.sh and final output to show ComfyUI URL
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- Both setup scripts now detect VRAM and determine which image gen
models the GPU can run (SD 1.5 at 4GB, SDXL at 8GB, Flux at 12-20GB)
- New setup-image-models.sh: interactive script that detects GPU,
shows available models with VRAM requirements, and installs into
InvokeAI and/or ComfyUI. Supports --auto for unattended install.
- Scales from 4GB cards through dual RTX 5000s to high-end 48GB cards
- README: added image gen VRAM tier table, expanded inpainting docs
with practical fix recipes (hands, fingers, eyes, backgrounds),
mask tips, and denoising strength guidance
- Setup end messages now show image gen capabilities and point to
setup-image-models.sh instead of manual model install instructions
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Major expansion of the InvokeAI section:
- Add InvokeAI vs ComfyUI comparison table (when to use each)
- Step-by-step: install base model, import RunPod LoRA, text-to-image
with LoRA, img2img iteration with denoising strength guide
- Denoising strength table explaining what 0.2 vs 0.8 actually does
- Example prompts for age up/down, emotion, setting, art style changes
- Unified Canvas / inpainting instructions for selective editing
- 6GB GPU notes (SD 1.5 fits, SDXL is tight)
- Expanded troubleshooting for greyed-out buttons and VRAM issues
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
New script: comfyui-install-ipadapter.sh
- Installs cubiq/ComfyUI_IPAdapter_plus custom nodes
- Downloads CLIP Vision encoders and IP-Adapter models (SDXL/SD1.5)
- Optional --faceid flag for stronger face identity lock
- Skips already-downloaded files, pulls updates on re-run
- Prints wiring diagram and example prompts after install
README:
- Add "IP-Adapter: same face, different settings" section with
task table, install commands, workflow guide, weight tuning
- Clarify IP-Adapter = ComfyUI direct (not from OWUI chat)
- Add both new scripts to the file listing table
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- README: Add "Combining multiple LoRAs" section with wiring diagram,
expand style management table with combo examples, clarify that
LoRAs are baked into workflows with no chat keyword activation
- Script: Add multi-LoRA tip and clarify no-keyword behavior
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- New script: comfyui-import-lora.sh — copies .safetensors into
ComfyUI's Docker volume and prints step-by-step instructions for
wiring it into a workflow and exporting to Open WebUI
- README: Add "Using LoRAs with Open WebUI" section documenting the
workflow-per-style pattern, multi-LoRA management, and architecture
compatibility table
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Setup script changes:
- All prompts now use whiptail dialogs with text fallback
- Q1b (SSH), Q2 (storage), Q3 (Kiwix), Q4b (firewall), Q5 (models),
Q6 (download), final confirm all converted
- Model tier selection uses radiolist with recommended tier pre-selected
- Custom model entry uses inputbox with current defaults pre-filled
- Fix speed_label: now shows actual VRAM needed (file size + 2GB overhead)
instead of misleading "fully in VRAM" for models that don't fit
- qwen3.5-35b-a3b MoE already in tier list (was there, now with accurate
VRAM estimate shown)
README changes:
- Add "Realistic expectations by model size" table
- 35B MoE highlighted as sweet spot for small GPUs
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
The stack already has a code-aware RAG server that auto-indexes repos
and retrieves relevant code chunks during chat - but this wasn't
documented clearly enough. Added:
- Architecture diagram showing RAG server data flow
- Three methods to index repos (manual, API, Gitea webhook)
- What gets indexed (file types, AST parsing, smart chunking)
- Comparison table: RAG server vs Knowledge Collections vs Memories
- Updated Local AI vs Claude Code comparison to reflect code awareness
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- How to create knowledge collections from chat summaries (handoff workflow)
- RAG tuning settings for better retrieval quality
- Community functions for context management (summarization, clipping)
- How Open WebUI handles context overflow (truncation, not summarization)
- Honest comparison table: Local AI vs Claude Code tradeoffs
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- Manual function/action installation without community signup
- Auto Memory setup with Ollama configuration
- Memory vs Knowledge collections comparison (context impact)
- Context window consumption analysis (memories are ~200 tokens fixed,
conversation history is the real context hog)
- Project-scoped memory workarounds (Knowledge collections recommended)
- System prompt fix for models outputting code instead of natural language
- VRAM reality check table (model file size != inference VRAM needed)
- Qwen 9B does NOT fit in 6GB VRAM despite setup script claiming so
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- Add ComfyUI as optional service in setup wizard (alongside InvokeAI)
- Configure Open WebUI env vars for ComfyUI integration when selected
(ENABLE_IMAGE_GENERATION, IMAGE_GENERATION_ENGINE, COMFYUI_BASE_URL)
- Add ComfyUI docker service (ai-dock/comfyui, port 8188, GPU access)
- Add comprehensive image generation documentation to README:
- ComfyUI + Open WebUI setup steps (model install, workflow export, node mapping)
- AUTOMATIC1111 alternative setup
- Environment variables reference table
- VRAM considerations for simultaneous LLM + image gen
- Update all touchpoints: UFW rules, Caddyfile, start.sh URLs, volumes,
compose services, summary output, directory creation
- Note: InvokeAI does NOT integrate with Open WebUI natively (no compatible API)
ComfyUI is the recommended path for chat-integrated image generation
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Setup wizard ZIM checklist now includes:
- Stack Overflow, Ask Ubuntu, Super User, Unix & Linux SE, Server Fault
(was only Stack Overflow before)
- DevDocs and FreeCodeCamp moved up near the top
- List grouped: Dev & Sysadmin → Reference → Other
- Whiptail dialog height increased to fit all 22 items
Both whiptail (GUI) and plain terminal (fallback) lists match.
Download handlers added for all new SE sites.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- Option 2 description: "Qwen2.5 · Qwen2.5-Coder" → "Qwen 3.5 (Feb 2026)"
- Default choice: 1 (Western) → 2 (Performance/Qwen 3.5)
- Recommended tier logic: use TOTAL_VRAM and match actual tier names
(4B/9B/35B/27B for Qwen 3.5, not old 7B/14B/22B/32B)
- Users with 6GB GPUs now see the 4B option marked "fast — fully in VRAM"
instead of having to pick Custom and type model names manually
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Context management for local models with limited windows:
- Context Tracker (community function): shows tokens used vs available,
progress bar, percentage remaining — so you see the cliff coming
- Checkpoint Summarization Filter: auto-summarizes old messages when
context fills up, like Claude's auto-compaction
- Both added as recommended post-install links in setup output
(Open WebUI Functions install with one click from the UI)
6GB GPU tier (Quadro P3300, GTX 1060, etc.):
- Qwen 3.5 4B at Q4_K_M = ~2.5GB weights, leaves 3.5GB for KV cache
- With Q8 KV cache: ~32K usable context on 6GB
- Better than squeezing 9B into nothing — more context > slightly smarter
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Main added a basic kiwix_search tool. Our branch already has a unified
search() that does everything kiwix_search did plus: DDG fallback,
freshness detection, structured results, and read_doc() for full articles.
Keep our version, keep the sync tool from our branch.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Replaced three separate tools (search_docs, web_search, search) with
a smart unified search() that:
- Always hits Kiwix first (instant, offline, no rate limit)
- Checks query for freshness keywords (latest, 2026, release, CVE, etc)
- If time-sensitive: also hits DDG, flags "prefer live results"
- If timeless (algorithms, docs, concepts): Kiwix only, skips web hit
- If Kiwix returns nothing: falls back to DDG automatically
Model sees both result sets with clear guidance on which to trust.
Keeps read_doc() for reading full Kiwix articles and web_search()
for explicit live-only queries when verifying offline currency.
130GB of ZIMs earn their keep on timeless topics (no rate limit,
instant, complete articles). DDG covers everything else.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
The model can now search before it generates:
MCP tools (available via Open WebUI + Claude Code):
- search_docs(query): searches Kiwix ZIM files (Wikipedia, Stack Overflow,
DevDocs, Arch Wiki) — instant, offline, no rate limits
- read_doc(path): reads full article content from Kiwix results
- web_search(query): DuckDuckGo search, no API key needed
Open WebUI native search:
- ENABLE_RAG_WEB_SEARCH=true + RAG_WEB_SEARCH_ENGINE=duckduckgo
- Switched from SearXNG (not in stack) to DDG (zero config)
Search priority: Kiwix first (offline, fast) → DDG fallback (live web)
All 3 setup scripts updated:
- duckduckgo-search added to mcp_requirements.txt
- KIWIX_URL=http://kiwix:80 added to MCP container env
- curl added to MCP container deps (for sync script)
- Open WebUI DDG search enabled by default
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Six stackable techniques that compound:
1. KV cache quantization (Q8_0 = 2x context, asymmetric K=Q8/V=Q4 = 2.6x)
2. Flash attention (free VRAM + speed, zero quality loss)
3. Host-memory prompt caching (--cram, use the server's 128-384GB RAM)
4. KV to system RAM (-nkvo, last resort, 5-20x slower)
5. Architecture selection (GQA + MoE = tiny KV footprint)
6. NVMe mmap for model loading (fast cold starts, not inference)
Stacked result: single card goes from ~50K to ~130K usable context
with the MoE model. Updated llama.cpp config with all flags.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
VRAM must hold weights AND KV cache. What's left after weights is
your real context window. For a 10K-line project (~100-150K tokens):
- 1 card: 32-100K usable, file-by-file workflow
- 2 cards: 80-180K usable, whole-project-in-one-shot workflow
- This Opus session: 1M tokens (neither setup comes close)
Comparison table vs Claude Opus 1M session for perspective.
Keep Pro for the hard stuff, use local for the daily 80%.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Single card table was showing A- for 35B-A3B which undercut the case
for a second card. Restructured to lead with 16GB (start here) at
honest B+ ratings, then show what 32GB unlocks: dense 27B/32B models
that physically can't fit on 16GB and are genuinely A- quality.
The upgrade isn't about 262K context - it's about accessing better
models (27B dense > 35B MoE with 3B active params).
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
The 35B-A3B MoE activates only 3B params per token - quality tracks
active params, making it more like a smart 7B than a true 35B. Revised
rating to B+ to A-. The 27B dense model is the real A- but needs 17GB
weights. Also notes: 262K is a VRAM ceiling not a quality guarantee,
Q4 quantization costs something, and context quality degrades at edges.
Practical high-quality context is more like 64-128K.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Both GPUs must be dedicated to the LLM for 262K context window.
Image gen and LLM run one at a time (swap takes seconds).
Simultaneous only possible with smaller context (~64K on one card).
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
The real use case: drop Claude Max ($100/mo), keep Pro ($20/mo), offload
bulk codebase work to local. Local handles the 80% (reading 10K-line
projects, routine fixes, boilerplate) with no rate limits. Pro handles
the hard 20% where Opus quality matters. GPU pays for itself in 13 months,
saves $1,596 over 3 years vs Max.
Also adds: NVLink speed reality check, image gen capabilities (SDXL/Flux
included at no extra cost), hardware longevity estimate (3-5 years),
and updated cost comparison tables.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Complete shopping list with Dell-specific part numbers:
- GPU power cables: 9H6FV / N08NH (~$10-15 each, need 2)
- GPU Riser 3 required for second GPU slot
- Low-profile heatsinks needed on R720 (usually pre-installed on R730)
- 2x 1100W PSUs mandatory, non-redundant mode for full wattage
Documents riser layout, NVLink bridge clearance in 2U, potential issues
(CPU TDP limits, "unsupported" GPU warning, blower noise), and
R720 vs R730 comparison. Total build cost ~$940-960 with Dell parts.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Research and document the $850 budget build: 2x Quadro RTX 5000 (16GB each)
connected via NVLink for 32GB unified VRAM. With Qwen 3.5's Gated Delta
Network architecture (Feb 2026), this setup runs the 35B-A3B MoE model
with full 262K context in ~25GB — the best price-to-capability ratio
for local AI coding available.
Includes:
- Exact NVLink bridge part numbers (RTX 5000 uses unique smaller connector)
- Motherboard/PSU requirements and slot spacing guidance
- llama.cpp and Ollama multi-GPU configuration
- VRAM budget calculations for all Qwen 3.5 model sizes
- Phased build plan (start with 1 card at $400, add second later)
- Updated model table with full Qwen 3.5 family specs
- Cost comparison vs RTX 8000, RTX 3090, Claude Max, and API pricing
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
The $750 listing was a single outlier. Actual market price for
Quadro RTX 8000 Passive 48GB is $2,000-2,900 on eBay. Updated
all recommendations and cost comparisons accordingly.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- Added full 48GB GPU market comparison (RTX 8000, A40, A6000, L40, RTX 6000 Ada)
- Quadro RTX 8000 Passive at $750-1,400 is 4-5x cheaper than alternatives
- Added RTX 8000 LLM benchmarks (34 t/s on 30B models at 8K context)
- Explained why 48GB >> 24GB for coding: context window is the bottleneck
- Added 2026 coding model landscape (Qwen3.5 27B, Qwen3-Coder, etc.)
- Revised recommendations: RTX 8000 as primary, dual P40 as budget alt
- Updated config notes for 48GB (32K context, higher quantization options)
- All prices verified from real listings as of March 22, 2026
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- Corrected all GPU prices to actual eBay/Newegg listings as of March 22, 2026
- Added analysis of 32B coding models (Qwen2.5-Coder 32B, Qwen3.5 27B) on 24GB VRAM
- Added honest comparison of local LLM quality vs Claude Code
- Revised recommendations: dual P40 ($400-500) or P40+T4 ($400-550)
- Added configuration notes for 32B models, dual-GPU, and newer Qwen3 models
- RTX A4000 at $700+ is too expensive for this build
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Research comparing P40, A2000, T4, and M40 GPUs for LLM inference
and image generation in Dell R720/R730 rack servers. Includes
benchmarks, compatibility notes, pricing, and recommendations.
https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
- Add aider Docker service (paulgauthier/aider) with GUI on port 8080
pointing at local ollama code model (auto-detected by VRAM tier)
- Add aider.sh CLI wrapper: cd into any git repo and run files through
aider terminal mode against the same local model
- Add port 8080 to UFW firewall rules
- Surface Aider UI URL in start.sh output and final install summary
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
ollama pull compares digests server-side so no download happens if
a model is already current. Timer runs monthly offset 2h from kiwix.
install.sh updated to handle all units generically.
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
User units in systemd/ with an install.sh that symlinks and enables.
Persistent=true means missed runs (machine off) catch up on next boot.
Logs via journald: journalctl --user -u kiwix-update.service
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
- kiwix_download.sh: add Ask Ubuntu, Super User, Unix & Linux SE, Server Fault
- mcp_server.py: add kiwix_search tool — models can now query all offline ZIMs
- local-ai-setup.sh: pass KIWIX_URL env to MCP container, add kiwix dependency
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
Replace subshell ls expansion (which produces no args when /data is empty)
with a conditional: serve ZIM files if they exist, otherwise sleep infinity
so the container stays up gracefully until ZIM files are added.
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
SearXNG:
- Dropped from both setup scripts and docker-compose (Google/Startpage
block self-hosted instances by IP — not reliable enough to include)
- Removed all interactive safe-search and engine-selection prompts
- Removed Open WebUI RAG web-search env vars (ENABLE_RAG_WEB_SEARCH,
SEARXNG_QUERY_URL, etc.)
- Removed port 8888 from UFW rules, start.sh URLs, done output, and
the Caddyfile template
- Removed searxng/ from .gitignore (directory no longer created)
- configure-searxng-safesearch.sh kept in repo for optional manual use
Kiwix:
- Replace the blocking wait-loop (`until ls *.zim`) with a one-liner
that passes whatever ZIM files exist (or none) directly to kiwix-serve,
so the container starts immediately and shows an empty library page
rather than hanging until ZIMs are downloaded
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
All scripts now default BASE to their own directory (SCRIPT_DIR) instead
of ~/docker/ai-stack, so after a git clone the user never needs to change
folders — docker-compose.yml, .env, settings, and data dirs all live next
to the scripts.
configure-searxng-safesearch.sh and configure-storage.sh do the same for
standalone use (still honoured when called with BASE= from a parent script).
docker-compose.yml (generated by both setup scripts) gains a comment block
at the top listing the everyday docker compose commands:
up -d / down / restart / stop / logs -f / pull / ps
so the file itself is the reference for managing containers.
.gitignore added to exclude generated files (docker-compose.yml, .env,
requirements.txt, helper scripts) and data directories (workspace/, kiwix/,
searxng/, gitea/, etc.) from git tracking.
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
mcp_server.py / local-ai-setup.sh:
FastMCP.get_asgi_app() was removed in mcp 1.x — replace with
sse_app(), which returns the same Starlette ASGI app for SSE
transport. Fixes the crash loop the container was stuck in.
laptop_full_setup.sh Caddyfile:
$(hostname).local requires mDNS (Avahi/Bonjour) to resolve from a
proxy machine, which is often not available on all LAN clients.
Replace every occurrence with $LOCAL_IP (the machine's LAN IP)
so the generated Caddyfile.example works reliably regardless of
mDNS support.
https://claude.ai/code/session_012gDnantBmFTWZGCiKyjazx
Adds invokeai-import-lora.sh to copy LoRA .safetensors files directly
into InvokeAI's Docker model volume, bypassing the greyed-out UI upload
buttons. Also adds README documentation for LoRA usage and troubleshooting.
https://claude.ai/code/session_01RU7NQuTbA8S8NRoWojvhR5