Adds invokeai-import-lora.sh to copy LoRA .safetensors files directly into InvokeAI's Docker model volume, bypassing the greyed-out UI upload buttons. Also adds README documentation for LoRA usage and troubleshooting. https://claude.ai/code/session_01RU7NQuTbA8S8NRoWojvhR5
189 lines
7.1 KiB
Markdown
189 lines
7.1 KiB
Markdown
# Local AI Stack
|
||
|
||
A fully offline, self-hosted AI environment for Ubuntu 24.04. Runs on any NVIDIA GPU (or CPU-only).
|
||
|
||
**Services:** Ollama · Open WebUI · RAG · MCP · ChromaDB · SearXNG · Kiwix · Gitea · InvokeAI · Portainer
|
||
|
||
---
|
||
|
||
## Quick Start
|
||
|
||
```bash
|
||
git clone <this-repo>
|
||
cd local-ai
|
||
./laptop_full_setup.sh
|
||
```
|
||
|
||
That's it. The script installs Docker, NVIDIA drivers (if needed), generates all config, starts the stack, and optionally pulls models.
|
||
|
||
---
|
||
|
||
## Service URLs
|
||
|
||
After setup, all services are available on your LAN:
|
||
|
||
| Service | URL | Purpose |
|
||
|------------|----------------------------|--------------------------------|
|
||
| Open WebUI | `http://<ip>:3000` | Chat interface (Ollama + RAG) |
|
||
| InvokeAI | `http://<ip>:9090` | Image generation |
|
||
| SearXNG | `http://<ip>:8888` | Private web search |
|
||
| Kiwix | `http://<ip>:8181` | Offline Wikipedia / docs |
|
||
| Gitea | `http://<ip>:3001` | Self-hosted Git |
|
||
| RAG Health | `http://<ip>:8001/health` | RAG server status |
|
||
| MCP SSE | `http://<ip>:8002/sse` | MCP endpoint for Claude Code |
|
||
| Portainer | `https://<ip>:9443` | Docker management UI |
|
||
|
||
---
|
||
|
||
## Day-to-Day Commands
|
||
|
||
All generated into `~/docker/ai-stack/` by the setup script:
|
||
|
||
```bash
|
||
bash ~/docker/ai-stack/start.sh # pull latest images + docker compose up -d
|
||
bash ~/docker/ai-stack/stop.sh # docker compose down
|
||
bash ~/docker/ai-stack/status.sh # GPU / container / RAG health
|
||
bash ~/docker/ai-stack/pull-models.sh # pull Ollama models (run once after first install)
|
||
```
|
||
|
||
The stack also registers as a **systemd service** that starts on boot:
|
||
|
||
```bash
|
||
sudo systemctl start local-ai
|
||
sudo systemctl stop local-ai
|
||
sudo systemctl status local-ai
|
||
```
|
||
|
||
---
|
||
|
||
## Script Reference
|
||
|
||
| Script | Lines | What it does |
|
||
|---|---|---|
|
||
| `laptop_full_setup.sh` | 620 | **Main setup.** Installs Docker + NVIDIA toolkit, creates `~/docker/ai-stack/`, writes `docker-compose.yml`, starts stack, registers systemd service. |
|
||
| `local-ai-setup.sh` | 837 | Alternative setup script. Same as above but also **auto-detects VRAM** and selects models accordingly (14B for ≥14GB VRAM, 7B for CPU). Use this instead of `laptop_full_setup.sh` if you want VRAM-aware model selection. |
|
||
| `ubuntu-post-install.sh` | 8,889 | Full Ubuntu 24.04 post-install (dev tools, fonts, apps, tweaks). Run once on a fresh OS install. Independent of the AI stack. |
|
||
| `configure-storage.sh` | 239 | Storage/mount configuration helper. Run separately if you have a secondary drive for AI data. |
|
||
| `kiwix_download.sh` | 198 | Downloads ZIM files (Wikipedia, Stack Overflow, etc.) for offline use. Run separately — files are large. |
|
||
| `invokeai-import-lora.sh` | 85 | Copies a LoRA `.safetensors` file into InvokeAI's Docker model volume. |
|
||
|
||
### Which setup script should I use?
|
||
|
||
- `laptop_full_setup.sh` — fixed model selection (`qwen2.5:14b` / `qwen2.5-coder:7b`), simpler
|
||
- `local-ai-setup.sh` — detects your VRAM at runtime and picks appropriate models, also embeds `server.py` and `mcp_server.py` directly (doesn't need repo files copied separately)
|
||
|
||
Both scripts are **idempotent** — safe to re-run for updates. Config files are kept on re-run unless you pass `--force`.
|
||
|
||
---
|
||
|
||
## Generated File Layout
|
||
|
||
```
|
||
~/docker/ai-stack/
|
||
├── docker-compose.yml # generated by setup script
|
||
├── .env # API tokens — edit this, never overwritten
|
||
├── server.py # RAG server (copied from repo)
|
||
├── mcp_server.py # MCP server (copied from repo)
|
||
├── requirements.txt # RAG Python deps
|
||
├── mcp_requirements.txt # MCP Python deps
|
||
├── start.sh # start the stack
|
||
├── stop.sh # stop the stack
|
||
├── status.sh # GPU + container + RAG health
|
||
├── pull-models.sh # pull Ollama models
|
||
├── Caddyfile.example # reverse proxy config template
|
||
├── papers/ # drop PDFs here for RAG indexing
|
||
├── repos/ # git repos indexed by RAG
|
||
├── workspace/ # MCP working directory
|
||
├── index/ # ChromaDB vector store (persistent)
|
||
├── kiwix/ # ZIM files for Kiwix
|
||
├── gitea/ # Gitea data
|
||
├── invokeai-outputs/ # InvokeAI generated images
|
||
└── logs/
|
||
```
|
||
|
||
---
|
||
|
||
## First Run Checklist
|
||
|
||
1. **Run setup:**
|
||
```bash
|
||
./laptop_full_setup.sh
|
||
```
|
||
2. **Pull models** (prompted at end of setup, or run manually):
|
||
```bash
|
||
bash ~/docker/ai-stack/pull-models.sh
|
||
```
|
||
Downloads ~15-30GB. Takes 10-40 min depending on connection.
|
||
|
||
3. **Add API tokens** (optional — for Gitea/GitHub MCP tools):
|
||
```bash
|
||
nano ~/docker/ai-stack/.env
|
||
```
|
||
|
||
4. **Connect Claude Code to MCP:**
|
||
```bash
|
||
claude mcp add local http://<your-ip>:8002/sse
|
||
```
|
||
|
||
5. **Download ZIMs** for offline docs (optional, large):
|
||
```bash
|
||
./kiwix_download.sh
|
||
```
|
||
|
||
---
|
||
|
||
## Using LoRA Models in InvokeAI
|
||
|
||
LoRA (Low-Rank Adaptation) files let you customize image generation with fine-tuned styles or characters. If you trained a LoRA on RunPod or elsewhere, here's how to use it.
|
||
|
||
### Import a LoRA file
|
||
|
||
```bash
|
||
# Copy your LoRA into the InvokeAI Docker volume:
|
||
./invokeai-import-lora.sh ~/Downloads/my-lora.safetensors
|
||
|
||
# Optionally give it a display name:
|
||
./invokeai-import-lora.sh ~/Downloads/my-lora.safetensors "My Custom Style"
|
||
```
|
||
|
||
### Use the LoRA in InvokeAI
|
||
|
||
1. Open InvokeAI at `http://<ip>:9090`
|
||
2. Go to **Model Manager** (cube icon, left sidebar) and click **Scan for Models** / **Sync Models**
|
||
3. Your LoRA should appear in the model list
|
||
4. Switch to **Text to Image** tab
|
||
5. In the left panel, find the **LoRA** section (below the model selector)
|
||
6. Click **+** to add your LoRA, then adjust the **weight** slider (start at 0.7–0.85)
|
||
|
||
### Troubleshooting greyed-out upload buttons
|
||
|
||
- **No base model installed:** You need a fully downloaded base model (e.g., SD 1.5) before InvokeAI enables LoRA uploads. Use Model Manager to install one first.
|
||
- **Model not synced:** After copying files, click **Scan for Models** in Model Manager.
|
||
- **Architecture mismatch:** A LoRA trained on SD 1.5 only works with SD 1.5 base models — not SDXL or SD 2.x.
|
||
- **Use the import script instead:** The greyed-out UI upload can be bypassed entirely by using `invokeai-import-lora.sh` to copy files directly into the model volume.
|
||
|
||
---
|
||
|
||
## Updating
|
||
|
||
Re-run the setup script — it detects an existing install and skips prereqs:
|
||
|
||
```bash
|
||
./laptop_full_setup.sh
|
||
# or
|
||
./laptop_full_setup.sh --force # also overwrites config files
|
||
```
|
||
|
||
---
|
||
|
||
## GPU / Model Tiers (`local-ai-setup.sh`)
|
||
|
||
| VRAM | Chat model | Code model | Context |
|
||
|---------|-----------------|-----------------------|---------|
|
||
| ≥ 14 GB | qwen2.5:14b | qwen2.5-coder:14b | 32k |
|
||
| 8–14 GB | qwen2.5:14b | qwen2.5-coder:7b | 16k |
|
||
| 4–8 GB | qwen2.5:7b | qwen2.5-coder:7b | 8k |
|
||
| CPU | qwen2.5:7b | qwen2.5-coder:7b | 4k |
|
||
|
||
Embed model is always `nomic-embed-text` (required for RAG).
|