# Local AI Stack A fully offline, self-hosted AI environment for Ubuntu 24.04. Runs on any NVIDIA GPU (or CPU-only). **Services:** Ollama · Open WebUI · RAG · MCP · ChromaDB · SearXNG · Kiwix · Gitea · InvokeAI · Portainer --- ## Quick Start ```bash git clone cd local-ai ./laptop_full_setup.sh ``` That's it. The script installs Docker, NVIDIA drivers (if needed), generates all config, starts the stack, and optionally pulls models. --- ## Service URLs After setup, all services are available on your LAN: | Service | URL | Purpose | |------------|----------------------------|--------------------------------| | Open WebUI | `http://:3000` | Chat interface (Ollama + RAG) | | InvokeAI | `http://:9090` | Image generation | | SearXNG | `http://:8888` | Private web search | | Kiwix | `http://:8181` | Offline Wikipedia / docs | | Gitea | `http://:3001` | Self-hosted Git | | RAG Health | `http://:8001/health` | RAG server status | | MCP SSE | `http://:8002/sse` | MCP endpoint for Claude Code | | Portainer | `https://:9443` | Docker management UI | --- ## Day-to-Day Commands All generated into `~/docker/ai-stack/` by the setup script: ```bash bash ~/docker/ai-stack/start.sh # pull latest images + docker compose up -d bash ~/docker/ai-stack/stop.sh # docker compose down bash ~/docker/ai-stack/status.sh # GPU / container / RAG health bash ~/docker/ai-stack/pull-models.sh # pull Ollama models (run once after first install) ``` The stack also registers as a **systemd service** that starts on boot: ```bash sudo systemctl start local-ai sudo systemctl stop local-ai sudo systemctl status local-ai ``` --- ## Script Reference | Script | Lines | What it does | |---|---|---| | `laptop_full_setup.sh` | 620 | **Main setup.** Installs Docker + NVIDIA toolkit, creates `~/docker/ai-stack/`, writes `docker-compose.yml`, starts stack, registers systemd service. | | `local-ai-setup.sh` | 837 | Alternative setup script. Same as above but also **auto-detects VRAM** and selects models accordingly (14B for ≥14GB VRAM, 7B for CPU). Use this instead of `laptop_full_setup.sh` if you want VRAM-aware model selection. | | `ubuntu-post-install.sh` | 8,889 | Full Ubuntu 24.04 post-install (dev tools, fonts, apps, tweaks). Run once on a fresh OS install. Independent of the AI stack. | | `configure-storage.sh` | 239 | Storage/mount configuration helper. Run separately if you have a secondary drive for AI data. | | `kiwix_download.sh` | 198 | Downloads ZIM files (Wikipedia, Stack Overflow, etc.) for offline use. Run separately — files are large. | | `invokeai-import-lora.sh` | 85 | Copies a LoRA `.safetensors` file into InvokeAI's Docker model volume. | ### Which setup script should I use? - `laptop_full_setup.sh` — fixed model selection (`qwen2.5:14b` / `qwen2.5-coder:7b`), simpler - `local-ai-setup.sh` — detects your VRAM at runtime and picks appropriate models, also embeds `server.py` and `mcp_server.py` directly (doesn't need repo files copied separately) Both scripts are **idempotent** — safe to re-run for updates. Config files are kept on re-run unless you pass `--force`. --- ## Generated File Layout ``` ~/docker/ai-stack/ ├── docker-compose.yml # generated by setup script ├── .env # API tokens — edit this, never overwritten ├── server.py # RAG server (copied from repo) ├── mcp_server.py # MCP server (copied from repo) ├── requirements.txt # RAG Python deps ├── mcp_requirements.txt # MCP Python deps ├── start.sh # start the stack ├── stop.sh # stop the stack ├── status.sh # GPU + container + RAG health ├── pull-models.sh # pull Ollama models ├── Caddyfile.example # reverse proxy config template ├── papers/ # drop PDFs here for RAG indexing ├── repos/ # git repos indexed by RAG ├── workspace/ # MCP working directory ├── index/ # ChromaDB vector store (persistent) ├── kiwix/ # ZIM files for Kiwix ├── gitea/ # Gitea data ├── invokeai-outputs/ # InvokeAI generated images └── logs/ ``` --- ## First Run Checklist 1. **Run setup:** ```bash ./laptop_full_setup.sh ``` 2. **Pull models** (prompted at end of setup, or run manually): ```bash bash ~/docker/ai-stack/pull-models.sh ``` Downloads ~15-30GB. Takes 10-40 min depending on connection. 3. **Add API tokens** (optional — for Gitea/GitHub MCP tools): ```bash nano ~/docker/ai-stack/.env ``` 4. **Connect Claude Code to MCP:** ```bash claude mcp add local http://:8002/sse ``` 5. **Download ZIMs** for offline docs (optional, large): ```bash ./kiwix_download.sh ``` --- ## Using LoRA Models in InvokeAI LoRA (Low-Rank Adaptation) files let you customize image generation with fine-tuned styles or characters. If you trained a LoRA on RunPod or elsewhere, here's how to use it. ### Import a LoRA file ```bash # Copy your LoRA into the InvokeAI Docker volume: ./invokeai-import-lora.sh ~/Downloads/my-lora.safetensors # Optionally give it a display name: ./invokeai-import-lora.sh ~/Downloads/my-lora.safetensors "My Custom Style" ``` ### Use the LoRA in InvokeAI 1. Open InvokeAI at `http://:9090` 2. Go to **Model Manager** (cube icon, left sidebar) and click **Scan for Models** / **Sync Models** 3. Your LoRA should appear in the model list 4. Switch to **Text to Image** tab 5. In the left panel, find the **LoRA** section (below the model selector) 6. Click **+** to add your LoRA, then adjust the **weight** slider (start at 0.7–0.85) ### Troubleshooting greyed-out upload buttons - **No base model installed:** You need a fully downloaded base model (e.g., SD 1.5) before InvokeAI enables LoRA uploads. Use Model Manager to install one first. - **Model not synced:** After copying files, click **Scan for Models** in Model Manager. - **Architecture mismatch:** A LoRA trained on SD 1.5 only works with SD 1.5 base models — not SDXL or SD 2.x. - **Use the import script instead:** The greyed-out UI upload can be bypassed entirely by using `invokeai-import-lora.sh` to copy files directly into the model volume. --- ## Updating Re-run the setup script — it detects an existing install and skips prereqs: ```bash ./laptop_full_setup.sh # or ./laptop_full_setup.sh --force # also overwrites config files ``` --- ## GPU / Model Tiers (`local-ai-setup.sh`) | VRAM | Chat model | Code model | Context | |---------|-----------------|-----------------------|---------| | ≥ 14 GB | qwen2.5:14b | qwen2.5-coder:14b | 32k | | 8–14 GB | qwen2.5:14b | qwen2.5-coder:7b | 16k | | 4–8 GB | qwen2.5:7b | qwen2.5-coder:7b | 8k | | CPU | qwen2.5:7b | qwen2.5-coder:7b | 4k | Embed model is always `nomic-embed-text` (required for RAG).