Commit Graph
100 Commits
Author SHA1 Message Date
Claude cd003dbaf3 Add Gatus auto-sync from Caddyfile, promote ensure_yq to lib/common.sh
Adds one Gatus endpoint per Caddy site block automatically, tagged
group: caddy-sync — the sync only ever adds/removes entries in that
exact group, so anything added by hand (the default external checks,
a custom endpoint) is never touched regardless of what the Caddyfile
looks like. Offered at install time (syncs once immediately) and, if
systemd is available, scheduled via a timer every 15 minutes so a site
added or removed later gets picked up without re-running the installer
— matches the "schedule that checks the Caddyfile" shape asked for.

Domain extraction tracks actual brace depth (reusing the same approach
as remove_service's Caddy block removal) rather than a naive
line-by-line scan, so it correctly skips the global options block and
parenthesized snippet definitions like (authelia) without needing to
special-case them by name.

Verified end-to-end against a real Caddyfile/config.yaml fixture with
the actual mikefarah/yq binary: initial sync adds the right entries
and leaves the default "external" group alone, a second run with no
Caddyfile changes is a true no-op (0 added, 0 removed), and changing
the Caddyfile (removing one site, adding another) correctly adds the
new endpoint and removes only the stale one.

Also fixes a real gap surfaced while building this: ensure_yq (used by
both gatus.sh now and onlyoffice.sh already) checked `command -v yq`
alone, which a box can satisfy with a completely different, incompatible
yq — confirmed live in this environment, Debian/Ubuntu's own `yq`
apt package is kislyuk/yq (a Python jq-wrapper) which silently errors
on mikefarah/yq's `e '.path' file` syntax every caller here depends on.
Now checks the version string actually identifies as mikefarah's
before trusting it, installing to /usr/local/bin (which precedes
/usr/bin on Ubuntu's default PATH) if not. Promoted ensure_yq itself
from onlyoffice.sh (its only previous user) to lib/common.sh now that
gatus.sh needs the same thing, so both share one implementation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 04:26:08 +00:00
Claude 7c85f326c8 Accept Beszel's own "copy for docker compose" snippet directly
Confirmed live: Beszel's Settings -> Tokens & Fingerprints page surfaces
a "copy for docker compose" shortcut as the prominent way to grab the
key/token — not a bare string — so the previous two-prompt flow (paste
plain key, paste plain token) didn't match what people actually have
in their clipboard. The user pasted that YAML snippet into .env by
hand afterward, using the container's raw KEY/TOKEN names and YAML
`NAME: 'value'` syntax instead of what the compose file's own
${AGENT_KEY:-}/${AGENT_TOKEN:-} substitution actually reads — the
agent then failed with "no key provided" since AGENT_KEY was never
actually set.

_beszel_extract_field pulls KEY/TOKEN out of whatever shape the paste
arrives in — YAML mapping (`KEY: 'value'`), compose list style
(`- KEY=value`), or plain `KEY=value` — regardless of quoting. The
agent prompt now accepts a multi-line paste (the whole snippet, or
just the two lines) instead of asking for two separately pre-extracted
values; if no labeled KEY/TOKEN line is found at all, it falls back to
treating the paste as a bare key and asks for the token separately, so
a Beszel version that really does just show plain strings still works.

Verified all three paths directly: the exact mixed KEY:/TOKEN= paste
the user had, the skip path (blank first line leaves .env untouched),
and the bare-value fallback with no labels at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 04:11:56 +00:00
Claude e10e1e90b4 Fix invalid docker-compose.yml when Beszel is installed with local Caddy
Confirmed live: "networks.beszel-agent additional properties ... not
allowed" — the root-level networks: block (_CADDY_NET_SECTION) was
placed between the two services instead of after both. Since it sits
at 0 indentation, YAML parsed the following beszel-agent: line as a
continuation of the networks: mapping instead of a new services: entry,
so the whole beszel-agent service definition got swallowed as if it
were a (invalid) child of networks.caddy_net. gatus.sh's identical
_CADDY_NET_BLOCK/_CADDY_NET_SECTION pattern never hit this because it
only ever has one service, so the same placement is always the last
content in the file there.

Moved _CADDY_NET_SECTION (the root-level networks: definition) to
after both services; _CADDY_NET_BLOCK (the per-service "join
caddy_net" snippet) stays right after the hub's own volumes, where it
correctly nests under the beszel: service only.

Verified by regenerating the compose file with local Caddy present and
parsing it with PyYAML: services.beszel and services.beszel-agent are
now proper siblings, beszel-agent keeps its image/environment/volumes
keys, beszel's own networks: is scoped to just that service, and the
root networks: definition is separate and correctly placed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:51:02 +00:00
Claude fc8f57ab51 Add bash tab-completion for setup.sh
./setup.sh mat<TAB> now completes to ./setup.sh mattermost, same for
flags. Service names are read fresh from services/*.sh on every
completion — never a hardcoded list, which would go stale the moment
a new service file gets added (matches this repo's own "adding a
service = adding one file, nothing generated" rule from CLAUDE.md).

Verified live: sourced the script and confirmed completions for "mat"
and "--li", and specifically confirmed "bes" resolves to "beszel" —
the service added earlier this same session — with zero changes
needed to the completion script itself, proving the list is genuinely
dynamic rather than something that looked right once and then rotted.

Self-locating via its own BASH_SOURCE path rather than a hardcoded
install directory, so it keeps working regardless of where the repo
is cloned. Works through a leading `sudo` via bash-completion's
standard sudo pass-through (enabled by default on Ubuntu).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:46:09 +00:00
Claude 7938399e99 Add Beszel for lightweight server + Docker monitoring
Answers "what's the best way to see CPU/RAM/disk usage on this box"
(IONOS's own dashboard doesn't expose it) and "does Gatus cover this" —
it doesn't, Gatus is a black-box HTTP check (is the site responding
from the outside), Beszel is white-box host/process monitoring (is the
box under memory/disk pressure, is a container actually running vs.
crash-looping). Complements Gatus rather than replacing it.

Mirrors the hub+agent same-system layout from beszel's own
supplemental/docker/same-system/docker-compose.yml (fetched from the
actual upstream repo, not reconstructed from memory) — hub is the web
dashboard, agent reads /var/run/docker.sock (read-only) to report every
currently-running container automatically, no per-service config
needed as containers get added or removed.

Genuinely a two-phase install: the hub's SSH keypair and universal
token only exist after logging into its web UI once, so this starts
the hub, walks through where to find both values, and finishes wiring
the agent once provided — skipping is fine, a rerun in "update" mode
detects the agent was never connected and offers to finish it.

Verified the generated docker-compose.yml/.env by running the actual
file-writing code path with docker/configure_caddy_for_service mocked
out — confirmed TOKEN/KEY are correctly left as literal
${AGENT_TOKEN:-}/${AGENT_KEY:-} for Docker Compose's own substitution
at "up" time, not prematurely expanded by the heredoc itself.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:43:58 +00:00
Claude eeaa2e6c64 Accept plain "remove"/"uninstall" as aliases for --remove
./setup.sh filebrowser remove (no dashes) fell through to the normal
install dispatch instead of removing anything, since only the --remove
flag form was recognized. Accept the bare words too — order-independent
either way (./setup.sh remove filebrowser works the same).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:34:57 +00:00
Claude efcad48755 Add a generic service removal command: ./setup.sh <name> --remove
No removal path existed anywhere in this repo — manually removing a
service meant hand-editing docker-compose.yml, the Caddyfile, and UFW
rules yourself, or just leaving orphaned config behind.

remove_service (lib/common.sh) handles the common case: stop/remove
the service's containers (with an explicit y/n on whether to also wipe
data volumes, default no), find and remove its Caddy site block if one
exists, remove any UFW rule tagged with its name, and optionally
delete its ~/docker/<name> directory (default no — keep data as a
safety net unless explicitly confirmed).

The Caddy site block removal (_remove_caddy_site_block) tracks actual
brace depth rather than scanning to the next blank line or EOF — the
same class of bug this repo already hit once with a naive Samba
config edit. Verified against a multi-block test Caddyfile with nested
log{}/header{} blocks: removes exactly the targeted block, leaves
every other block (including ones with their own nested braces)
byte-for-byte intact, and is a safe no-op when nothing matches.

Also fixes the UFW rule-number extraction: ufw status numbered pads
single-digit rule numbers with a leading space ("[ 3]" vs "[10]") to
align columns, which the regex didn't account for — every single-digit
rule would have silently never matched and never gotten deleted.

Wired into setup.sh as a new --remove flag, resolving SERVICE_ALIAS
and validating the name the same way run_service already does.
Scoped to the common case (a Docker service at $DOCKER_DIR/<name> with
a standard configure_caddy_for_service site block); a hand-built Caddy
block or non-standard layout may need manual cleanup for the parts
this can't find. Non-Docker services (base, ssh-key-import, etc.)
report cleanly that they're not handled rather than erroring
confusingly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:29:10 +00:00
Claude ad011b2db8 Add check_container_health helper, wired into mattermost.sh as reference
A service's "Started" message after docker compose up -d doesn't mean
the app is actually working — it can still crash-loop (bad DB
password, missing required env var, etc.) with no visible sign until
someone separately runs docker ps -a much later, exactly what happened
repeatedly this session (mattermost, koha-db, homebox, vaultwarden,
filebrowser all showed a clean "Started" message while crash-looping).

check_container_health (lib/common.sh) waits briefly, checks the
container's actual status and restart count via docker inspect, and
prints recent logs automatically if it's not running or has already
restarted — instead of a misleading one-line success message.

Wired into mattermost.sh's own start step as the reference
implementation, guarded by declare -F so standalone runs (no
lib/common.sh sourced) degrade gracefully. Not retrofitted across
every other service in one pass — this establishes the shared helper
so other services can adopt it incrementally.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:24:50 +00:00
Claude 74628f3736 Fix vaultwarden SMTP false-positive and mattermost DB password mismatch
vaultwarden: SMTP_PORT defaulted to "587" and SMTP_SECURITY was a
hardcoded "starttls" literal in the .env template, written
unconditionally regardless of whether SMTP_HOST was ever provided.
Confirmed live: skipping SMTP entirely (blank SMTP_HOST) still wrote
real values for those two, and Vaultwarden reads that as "some SMTP
config is present," refusing to start ("Both SMTP_HOST and SMTP_FROM
need to be set") even with host/from genuinely blank. Both now stay
empty unless SMTP_HOST is actually set.

mattermost: DB_PASS/MM_SECRET were only reused from the existing .env
when MODE=update — a "fresh" reinstall always generated a new
POSTGRES_PASSWORD. Confirmed live: choosing fresh after removing only
the mattermost app container (not the whole directory) regenerates the
password in .env while db/'s existing Postgres data still enforces the
OLD one from its first init (the entrypoint skips re-init on existing
data), causing "password authentication failed for user mattermost" on
every start. Whether db/ already has real data is what actually
determines whether the old password is still live, not which reinstall
mode was chosen — reuse the existing secrets whenever db/ is non-empty,
regardless of MODE.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:22:53 +00:00
Claude c177312947 Fix three crash-looping services: koha-db, homebox, vaultwarden
koha-db: compose used MYSQL_ROOT_PASSWORD/MYSQL_DATABASE/MYSQL_USER/
MYSQL_PASSWORD, but this mariadb:11 image version's entrypoint doesn't
recognize MYSQL_ROOT_PASSWORD as any of its accepted root-password
options at all. Confirmed live: "Database is uninitialized and password
option is not specified" on every start, even though DB_ROOT_PASS was
correctly generated and present in .env the whole time. Switched all
four to their MARIADB_* equivalents.

homebox: a newer homebox release requires HBOX_AUTH_API_KEY_PEPPER (at
least 32 bytes) or the container panics on startup — this installer
never set it. Generate one with generate_password 48 and wire it
through .env + the compose environment block.

vaultwarden: the SMTP setup prompts let you enter a host but leave
"SMTP from address" blank (no default), writing a half-configured state
Vaultwarden refuses to start with ("Both SMTP_HOST and SMTP_FROM need
to be set"). Validate after prompting — if SMTP_HOST is set but
SMTP_FROM came back empty, disable SMTP entirely instead of writing a
config known to crash the container.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 03:09:33 +00:00
Claude 77440c8f75 Live-scan Traccar's device-protocol range for collisions, not just Asterisk's ports
The 5000-5150 range was only ever checked against Asterisk's hardcoded
fixed ports (5038/5060/5061) for a first instance — no live scan of the
rest of the range, because the directory-count-based offset mechanism
only triggers for an explicit additional instance.

Confirmed live: this range sat unclaimed at the OS level while this
Traccar instance's container had never actually started, so an
unrelated service's own find_free_port scan found port 5007 genuinely
free (nothing was listening there yet) and took it — invisible to any
check until Traccar itself tried to bind its declared range for the
first time, failing with "port is already allocated".

Add a live scan across the whole intended range (skipping Asterisk's
expected carve-outs at the base 5000-5150 range) and shift by 1000,
same step the multi-instance path already uses, until genuinely clear.
Also fixed the compose-block and README generation, which keyed off
INSTANCE_SUFFIX being empty to decide whether Asterisk's exclusions
were needed — now keyed off whether the range is still the unshifted
default (PROTO_MIN -eq 5000), since a first instance can now end up
shifted too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 02:43:37 +00:00
Claude 8bdf4c0a07 Add explicit UFW rules for Caddy's 80/443 — Docker was silently bypassing UFW
services/caddy.sh never called ufw allow for any of its published ports.
Confirmed live: ufw status showed no rule for 80 or 443 on a box with
UFW active (default deny incoming), yet HTTPS sites were reachable
fine — Docker manipulates iptables directly for published container
ports (the ports: mapping in Caddy's own compose file), which bypasses
UFW's filtering entirely regardless of what ufw status reports.

This wasn't an actual exposure gap — 80/443 are supposed to be open to
everyone, that's the point of a reverse proxy — but it means ufw status
was actively misrepresenting this box's real firewall state on its two
most externally-facing ports, which is exactly the kind of thing that
looks like a problem (and did, when investigating an unrelated
Let's-Encrypt failure) even though nothing was actually unprotected.

Add explicit ufw allow rules for 80/tcp, 443/tcp, and 443/udp (HTTP/3)
so ufw status reflects reality, matching every other service in this
repo managing its own firewall rules instead of relying on undocumented
Docker/iptables interaction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 01:51:04 +00:00
Claude 060f288107 Fix authelia.sh skipping the real Caddy snippet due to a commented example
grep -q "(authelia)" matches caddy.sh's starter Caddyfile's own commented-
out example block ("# (authelia) {", included as documentation), so
authelia.sh believed the real snippet already existed and never wrote
it. Any later service adding `import authelia` to its own site block
then references a snippet that only exists as a comment.

Confirmed live: this takes Caddy down completely, not just the
Authelia-protected site — "Error: adapting config using caddyfile:
File to import not found: authelia" is a load-time failure, so Caddy
restart-loops and every site it fronts goes with it.

Anchor the check to an actual uncommented snippet definition
(^\(authelia\)\s*\{) instead of a bare substring match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 00:21:13 +00:00
Claude 6588e3abf1 Add a way to remove an existing vpn-data-mount without re-adding it
Removal only existed as a side effect of picking the same share again
in the "fully redo this mount" path — there was no direct way to just
remove a mount you no longer want, without walking back through host/
share selection first.

Adds a top-level "Remove any existing VPN data mounts?" prompt that
lists every configured mount by number (via the new
_vdm_list_all_mounts) and lets you remove one or more, reusing the
existing _vdm_remove_mount teardown (decrypt-layer unit, unmount,
credentials file, tagged /etc/fstab entry).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-11 00:15:30 +00:00
Claude 39e7b2ae6e Fix Mattermost crash-looping with permission denied on config.json
The official mattermost/mattermost-team-edition image runs as a fixed
UID/GID 2000 baked into the image — it does not read PUID/PGID env vars,
that's a LinuxServer.io s6-overlay convention this image doesn't use.
This file set them anyway (computed from ACTUAL_USER's uid/gid), which
did nothing, while the actual host directories (./data, ./logs,
./config, ./plugins) got chowned to ACTUAL_USER instead of 2000:2000.

Confirmed live: the container fails on its very first start with
"could not create config file: open /mattermost/config/config.json:
permission denied" and crash-loops — which then presents as a 502 from
Caddy, an easy trail to follow to the wrong place since Caddy itself
was fine.

Removed the dead PUID/PGID mechanism and chown the app's own volumes to
2000:2000 after the existing ACTUAL_USER chown. db (postgres:15-alpine)
isn't affected — its entrypoint fixes its own volume ownership on
startup. Runs on both fresh installs and "update" reruns, so re-running
the installer on an already-broken instance self-heals it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-10 23:41:41 +00:00
Claude cfd3b04b7b Add a full-redo option for an already-mounted vpn-data-mount share
The previous fix only let an already-mounted share reconfigure its
decrypt layer — there was still no way to change the mount point or
re-enter credentials for a share that's already set up, since the label
prompt was skipped entirely in that path. Add a real choice when an
existing mount is found: reconfigure the decrypt layer in place (as
before), fully redo the mount (tears down the old one via the new
_vdm_remove_mount and falls through to the normal fresh-mount flow,
label pre-filled from the old one), or skip.

_vdm_remove_mount stops/removes any decrypt-layer systemd unit first
(it sits on top of the CIFS mount), then unmounts, removes the
credentials file, and removes the /etc/fstab tag+entry via a fixed
",+1d" range — the tag line plus exactly the one mount line that always
immediately follows it, not an open-ended range to the next blank line
or EOF (the class of bug fixed earlier in this file's history for the
now-removed remote smb.conf-writing code).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-10 23:34:01 +00:00
Claude 304e4b644c Let an already-mounted share be reconfigured instead of blocking on label reuse
Re-running vpn-data-mount for a share that's already mounted hit the
label-uniqueness check with no way through it — picking the same share
always re-prompted for a label, and the existing label was always
already taken by definition, so it just looped rejecting every input.
Confirmed live: reported as an infinite "Label 'data1' is already
used" loop right after this share had already been mounted in an
earlier run.

Detect the existing fstab tag for the same host+share up front and
reconfigure it in place — currently the one thing safe to redo without
touching a working plain mount: the gocryptfs decrypt layer added
previously. Skips the label prompt and remount entirely for a share
that's already set up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-10 23:08:05 +00:00
Claude a23d5d6fc7 Add optional client-side encryption layer for vpn-data-mount
The VPS side of a plain SMB mount necessarily sees plaintext while it's
mounted and in use — that's unavoidable for data a VPS service actually
needs to read. What's avoidable is everything else: a disk image,
backup, or provider-side look at the VPS while the mount isn't actively
in use showing your actual files instead of ciphertext.

tools/gocryptfs-setup-home.sh (new): standalone tool for the home box.
Creates a gocryptfs-encrypted directory and passphrase file; the user
points their existing Samba share's `path =` at the cipherdir (manual
step — same read-only stance on remote Samba config vpn-data-mount.sh
already takes, this tool doesn't touch smb.conf either).

services/vpn-data-mount.sh: after mounting a share over CIFS as before,
optionally offers a gocryptfs decrypt layer on top. Fetches the
passphrase fresh over the same SSH trust already used for share
discovery, pipes it straight into gocryptfs, and never writes it to the
VPS's own disk. A generated systemd unit (via a wrapper script, not one
long quoted ExecStart= one-liner — avoids stacking systemd's own
word-splitting on top of bash -c's) keeps the decrypted view coming back
on boot, re-fetching the passphrase each time rather than caching it.

Fully opt-in and per-share — a plain unencrypted mount works exactly as
before if declined.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-10 23:06:13 +00:00
Claude 9e886ff9a7 Fix CIFS mount error(79) caused by missing nls_utf8 kernel module
The keyutils fix alone didn't resolve it — confirmed live with keyutils
already installed, the same error persisted. Root cause: the hardcoded
iocharset=utf8 mount option requires the kernel's nls_utf8 module, which
some kernels don't ship at all (confirmed live: `modprobe nls_utf8` on a
stock Ubuntu 6.8.0-137-generic VPS kernel returns "FATAL: Module
nls_utf8 not found" — not loadable, not built in). Every such mount
fails with errno 79 (ELIBACC) regardless of credentials, which is why
this recurred identically after the keyutils fix.

Both vpn-data-mount.sh and mount-network-drive.sh now probe with a
harmless `modprobe nls_utf8` before adding the option, and mount without
it (falling back to the kernel's build-time nls_default) with a clear
warning if the module isn't available, instead of hard-failing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-10 20:11:05 +00:00
Claude 2d9501a56c Fix CIFS mount error(79) caused by missing keyutils package
Errno 79 is ELIBACC ("Can not access a needed shared library"), not
ENOKEY as previously assumed — mount.cifs prints glibc's literal
strerror() text for it. It recurred with valid, correctly-captured
credentials because the real cause was never authentication: cifs-utils
hard-depends on the libkeyutils1 library but only Recommends the
keyutils package itself, which ships /sbin/request-key and the
/etc/request-key.d/*.conf handlers the kernel's upcall path invokes.
Minimal cloud VPS images commonly disable install-recommends, so
`apt-get install cifs-utils` alone silently skips it and every mount —
guest or fully credentialed — fails identically.

Install keyutils explicitly wherever cifs-utils is installed:
services/base.sh's unconditional package list, vpn-data-mount.sh's
lazy install-on-mount path, and tools/mount-network-drive.sh's SMB
branch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
2026-08-10 20:00:24 +00:00
Claude 1a259e0895 Fix password prompt silently stripping leading/trailing whitespace
Reported live: mount error(79) again despite already switching to real
credentials + sec=ntlmssp — this time with a Samba password containing
special characters. Root cause confirmed directly: `read -r -s pw1`
without `IFS=` silently strips leading/trailing whitespace even when
reading into a single variable (verified: " P@ss word! " -> "P@ss word!",
10 chars instead of 12). A password with a leading/trailing space —
common from a password manager's copy-paste, or a stray keystroke — got
quietly trimmed on the way into the credentials file, so it no longer
matched what was actually set on the Samba account. That mismatch
surfaces as this same cryptic ENOKEY mount error, not an obvious "wrong
password".

Fixed with IFS= on both reads. Also echo the captured length (never the
password itself) right after entry, so a silently-stripped character is
something you can catch and cross-check yourself before the mount even
attempts, instead of only after it fails.
2026-08-10 19:35:57 +00:00
Claude dd6f1d0a5d Make vpn-data-mount strictly read-only on the remote Samba config
Per direct request: never write to the home box's smb.conf at all, not
even carefully — just discover what's already shared there and mount it.
Removes all remote provisioning (installing Samba, creating/removing
share blocks, resetting smbpasswd accounts) entirely, which also removes
the whole class of bug the previous two fixes were patching around
(destructive section-removal, clobbering another mount's saved password) —
a tool that can't write can't repeat that kind of damage.

New flow: resolve/name the host and bootstrap SSH trust as before, then
read-only list every real share already in the home box's smb.conf
(skipping [global]/[homes]/[printers]/[print$]) via a plain SSH `cat`,
falling back to a sudo'd read only if that comes back empty — still only
ever reading. Presents them as a numbered list and accepts a flexible
selection ('1', '1,3', '1-3', '1 3 5', or combinations), asks once for the
Samba username/password to connect with (reusing a previously-saved
password for the same user+host if one exists), then mounts each picked
share locally over CIFS with its own /etc/fstab entry — same as before.

Verified the selection parser against all the documented formats plus a
mixed comma+range case and garbage/empty input.
2026-08-10 19:24:59 +00:00
Claude a2d3b0a651 Fix smb.conf section removal deleting everything after the target share
Reported live: Samba broke on the home box after this ran. Root cause
confirmed by reproducing it directly: the old removal step used
`sed -i "/^\[share\]$/,/^$/d"` — a range delete from the share's header
through the next BLANK line. A home box whose smb.conf has no blank line
separating sections (common — nothing requires one) means that range
never finds a terminator and sed deletes straight through to end of file,
taking every share defined after the target one down with it. Reproduced
against a 4-section smb.conf with no blank lines: the old approach left
only [global] standing, silently destroying two unrelated, pre-existing
shares that had nothing to do with this tool.

Replaced with an awk pass that removes lines from the target share's own
[header] up to the next `[section]` header or EOF — the actual boundary
of an INI-style section, independent of blank-line formatting. Also now
builds the new config in a scratch file and validates it with `testparm`
before it's ever copied over the live smb.conf; on validation failure it
leaves the existing file untouched and exits instead of restarting smbd
against a config that might not even parse. The existing
smb.conf.backup.<timestamp> step (already present before this fix) is
what the user is recovering the home box with in the meantime.

Verified the fix against the exact reproduction: the same 4-section,
no-blank-line smb.conf now retains all three untouched sections after
removing only the target one.
2026-08-10 19:18:55 +00:00
Claude 9e5a84f4f5 Don't blindly overwrite existing Samba config; stop clobbering shared account passwords
Reported live: the tool unconditionally reconfigured Samba even though
"Samba already installed on the home box" was already correctly detected —
that check only ever covered whether the smbd package exists, never
whether a share for the requested path (or the Samba account itself) was
already set up. Two real problems, not just a UX one:

1. Every run appended/replaced a [share] block and reset the target
   account's password unconditionally, even against a share the user had
   already configured by hand.
2. Since the Samba account is the SSH username (shared across every mount
   from the same home box), setting up a SECOND mount from the same box
   would silently reset the account's password — breaking the FIRST
   mount's already-saved credentials file with no warning.

Now: checks the remote smb.conf for an existing share exporting the exact
requested path first (via a plain SSH+awk query) and offers to reuse it
as-is (prompting for its real credentials, since a Samba password is
stored hashed and can't be read back) instead of overwriting it. If
creating a new share, checks whether this tool already set a password for
the same user+host pair (from another mount) and reuses it instead of
resetting the account; if the account exists with an unknown password
(set up some other way), asks rather than silently clobbering it.

_vdm_find_remote_share/_vdm_find_existing_smb_password/_vdm_prompt_password
are all called via command substitution by their caller, so none of them
call log_info/log_warning/etc. internally — those all write to stdout in
this codebase, which would corrupt the captured value. Verified the awk
share-lookup and the fstab-tag password lookup against sample data.
2026-08-10 19:15:11 +00:00
Claude abd1bcd35d Fix wordpress.sh picking an already-occupied port
Reported live: a fresh site's container failed to start with "address
already in use" on its assigned port. wordpress.sh scanned for a free
port by grepping `docker ps -a`'s port list — that only reflects ports
Docker itself currently has bound, so it's blind to ports held by
non-Docker processes or anything Docker isn't reporting cleanly at that
instant. Every other service in this repo scans with find_free_port
(checks actual OS-level listening sockets via ss) per CLAUDE.md's "Port
collision avoidance" section; wordpress.sh was the one holdout still using
its own weaker check. Switched to the shared helper, already available in
this file's own standalone stub and via lib/common.sh — no new dependency,
just using what was already sitting there unused.
2026-08-10 18:40:06 +00:00
Claude 28252b97d3 Fix vpn-data-mount's guest-mount errno 79 bug, add per-share SMB accounts,
host naming, and chain-in from filebrowser/audiobookshelf/emby

Reported live: "mount error(79): Can not access a needed shared library"
on the local CIFS mount step. That message is misleadingly worded — errno
79 is ENOKEY, not a real missing-library problem, and a plain `guest`
mount with no explicit `sec=` hitting it against a real Samba server is a
known cifs-utils/kernel-cifs rough edge in the anonymous-session keyring
path. Fixed as a side effect of switching away from guest access per
direct request (real per-share Samba accounts, not root/guest, matching
"user accounts for data directories"): each mount now gets a dedicated
Samba account (reusing the SSH username — that Unix account already
exists on the home box) with a generated password, remotely provisioned
via smbpasswd over the same SSH trust, and mounted locally via a
root-only credentials file (same convention tools/mount-network-drive.sh
already uses) plus an explicit sec=ntlmssp instead of guest.

Host naming: entering a raw IP now offers to name it in /etc/hosts, then
uses that name for everything from then on (SSH commands, the CIFS mount
address, and re-runs against the same IP). Deliberately /etc/hosts, not
~/.ssh/config — an SSH Host alias only helps the `ssh` command resolve a
name, mount.cifs never consults ~/.ssh/config at all, so an alias alone
wouldn't get the actual mount using a name. Still offers to also add a
matching SSH Host alias on top (pure convenience — skips typing the
username for interactive ssh use) when services/ssh-config.sh's helpers
are available.

Chain-in: filebrowser/audiobookshelf/emby now offer to run
vpn-data-mount first if their data is on a home box that isn't mounted
yet, and default their own directory prompt to whatever was just mounted
(VDM_LAST_MOUNT_POINT, explicitly unset before each chain call so an
unrelated earlier vpn-data-mount run in the same setup.sh session can't
leak its mount point in as a stale default).
2026-08-10 18:37:59 +00:00
Claude ad1a955096 Document ssh-key-import in the README
New section covering what it does, the public-vs-private-key security
model (only public keys are ever fetched, no outbound capability like
private-repo access is granted), and how to run it standalone via
sudo ./setup.sh ssh-key-import. Placed ahead of the existing SSH Host
aliases section since that section already references key import as
prior context ("after SSH key import, the wizard offers to add...").
2026-08-10 18:16:15 +00:00
Claude 8c5be53950 Extract SSH key import out of base.sh into a standalone, re-runnable service
Was only ever runnable once, buried inside base.sh's required-setup flow —
no way to re-run just this step for a box that already went through base
setup but needs another admin's key added later, or (the immediate case)
a home box for services/vpn-data-mount.sh that only needs this one step.

services/ssh-key-import.sh holds the real logic now (GitHub/Launchpad
import via ssh-import-id, optional password-auth lockdown); base.sh's
_base_setup_ssh chains into it the same way services/asterisk.sh chains
into security-dashboard/pstn-trunk, with a degraded (no import, just
ensures the SSH server itself is running) fallback for a pure standalone
`sudo bash base.sh` run with no sibling files sourced. Independently
runnable via `sudo ./setup.sh ssh-key-import` or `sudo bash
services/ssh-key-import.sh`, and shows up in the whiptail menu under
extras alongside ssh-config. Marked as never showing [installed] in
is_installed()/install_count(), same as ssh-config — it's a repeatable
management action, not a thing with an install state.
2026-08-10 18:10:27 +00:00
Claude 0e42de1cda Add vpn-data-mount: SMB mount from a NetBird-connected home box
Offered right after NetBird setup during required/base setup, matching
the requested flow (base packages -> NetBird -> data mount). Repeatable
by design rather than a one-shot step, since different services can have
data on different home boxes — asks for a home box IP every time and can
be run again for additional boxes/shares.

Flow: test for existing passwordless SSH first (covers "both boxes already
share a key via GitHub import, or any other means" for free — if it
already works, nothing else runs). If not, generate an SSH keypair and
offer ssh-copy-id or a manual/GitHub-import fallback (ssh-import-id, the
same mechanism base.sh's own SSH setup already uses) — needed because a
home box that took base.sh's "disable password login" option won't accept
ssh-copy-id at all. Once passwordless SSH works, use it to remotely
install and configure Samba on the home box for a chosen path, then mount
it locally over CIFS with a tagged /etc/fstab entry.

SMB over NFS/SSHFS per this session's direction: not a "huge" speed gap
for normal use, and SSHFS's own encryption is redundant overhead once the
VPN tunnel already encrypts everything. Guest-accessible (no separate
Samba credentials) since the VPN is the real access control — only
NetBird-connected peers can reach the home box's NetBird IP at all.

Also:
- cifs-utils added to base.sh's always-installed packages, same reasoning
  as Docker/Compose being unconditional there instead of installed lazily
  on first mount.
- is_installed()/install_count() in setup.sh gained a vpn-data-mount case
  (state lives in tagged /etc/fstab entries, not $DOCKER_DIR, since this
  isn't a Docker service) — mirrors wordpress's "count real instances"
  handling rather than a flat 0/1.
- Every SSH call in the new service explicitly runs as $ACTUAL_USER
  (sudo -u), not root — the script itself runs as root throughout, but the
  SSH key lives in $ACTUAL_HOME/.ssh, so a bare `ssh` call would silently
  use root's own ~/.ssh instead and never find it. Caught by review before
  this shipped, not after.
- UNATTENDED mode skips outright with a message instead of spinning
  forever on prompt_text's always-blank default under --unattended, since
  none of this flow's prompts (home box IP, remote path, ...) have a
  sane non-interactive default.
2026-08-10 17:56:50 +00:00
Claude b4a402e399 Drop the header row, tighten name-to-count spacing
Confirmed (again, by rendering into a captured pty and inspecting the
character grid) that whiptail always renders a blank line between the
instructional text and the checklist box itself, with no parameter to
remove it — so a header "directly above the purple box" isn't achievable
no matter how it's built. Per this session's direction: drop the header
line entirely and just tighten the gap between the count and the service
name (was up to 15 chars of mostly blank space from the wide count field
sized to match the now-removed header label; down to ~5).

Verified end-to-end in the same pty harness: rendered the real dialog,
sent actual keystrokes to toggle two items (one plain, one with a
double-digit count), captured the raw whiptail selection output, and
confirmed the existing "extract text after the last space" logic still
pulls the correct plain service names back out.
2026-08-10 16:47:48 +00:00
Claude 3afd7226f2 Pixel-align the header labels with their data columns
Previous commit's leading-space count for the header was an estimate and
visibly off in the follow-up screenshot. Rather than guess again, actually
rendered the dialog into a captured pty (whiptail installed locally,
output fed through pyte to reconstruct the real character grid) and
measured exact column offsets instead of eyeballing.

Root fix: "installed" (9 chars) and "# of installs" (13 chars) are wider
than the underlying data (an "x"-or-blank mark, a 1-2 digit count) — a
narrow data column can never align under a wide label and stay readable,
so it's the data fields that got widened to match the label widths, not
the other way around. Verified alignment holds across installed/
not-installed/double-digit-count rows and at the narrow 78-column width
floor (where the description truncates first now, not the install status —
correct priority, since status is the more critical of the two).
2026-08-10 16:40:18 +00:00
Claude 9bc2e6c510 Move the column header out of the checklist into the non-selectable instruction text
Requested: no checkbox on the header row at all, not just a harmless one.
The previous fake-row header still drew a real [ ] like every other row —
whiptail has no way to suppress that per-row, there's no such thing as a
non-selectable list item in a --checklist.

The instructional text above the list has no checkbox rendering at all
though, since it isn't a list item — moved the header there instead:
"installed" / "# of installs" / "service", spelled out per this session's
request instead of the terse "x"/"#". Spelled-out words can't line up
character-for-character under the 1-2-char data columns below and stay
readable, so the leading spaces are a best-effort approximation, not exact
alignment.

Adjusted the box-height overhead constant (+8 -> +9) since the
instruction text is now two lines instead of one, and dropped the
now-unnecessary sentinel-row filtering from the selection-handling code.
2026-08-10 16:33:32 +00:00
Claude d4a7b5a60e Add a fake header row and size the checklist width to the terminal
Header row: a first, non-functional checklist entry using the exact same
printf field widths as the real rows ("x #  NAME" / "x 1  caddy" / ...),
so it visually reads as column headers for the x/# prefix even though
whiptail has no real header concept. Its sentinel tag ("NAME") is filtered
back out of the selection after the dialog closes, so it's harmless even
if someone checks it and hits <Ok>.

Width: was a flat 78 regardless of the actual terminal, so descriptions
got cut off mid-sentence on anything wider with no way to read the rest
(confirmed from a screenshot — "TURN via the shared coturn s..." trailing
off). Scale with tput cols instead, floored at the old 78 (safe on a plain
80-column terminal) and capped at 160 so a very wide terminal doesn't get
an absurdly wide dialog.
2026-08-10 16:23:35 +00:00
Claude 2a11993c4e Fake dedicated "installed"/"#" columns in the checklist via a fixed-width tag prefix
Requested: separate, non-interactive "installed" (x) and "#" (instance
count) columns ahead of the actual selectable checkbox, with the
description no longer carrying any install-status text at all.

whiptail's checklist only has one interactive element per row — the
checkbox — so there's no such thing as a real extra column, tabbable or
not; the tag and item fields are always just inert display text regardless
of what's in them. The closest real equivalent: bake a fixed-width "x"
(installed) + count prefix into the tag field itself. whiptail pads every
row's tag field to the same width, so it lines up visually like columns
even though it's one string underneath. Extract the plain name back out
before dispatch by taking the last whitespace-separated token, since
service names never contain spaces — robust regardless of the exact
prefix width.

Description field is back to plain SERVICE_DESC text now that install
status lives in the tag prefix instead.
2026-08-10 16:18:34 +00:00
Claude 7f69d1dbee Replace "[installed]" text with an install count "[N]"
"[installed]" was 11 characters of an already-tight 78-column checklist
row, most of the reason the marker had so little room to spare before
whiptail's width truncation silently dropped it (previous commit). "[N]"
says the same thing in 3 characters — and for services that support
CLAUDE.md's multi-instance pattern (a base install plus any number of
"<name>-<suffix>" siblings, e.g. two separate mattermost instances), it's
more informative than a flat "installed": N > 1 means several instances
exist, not just one.

Add install_count() alongside is_installed() in setup.sh: the default case
counts $DOCKER_DIR/<name> plus any $DOCKER_DIR/<name>-* siblings; the
specially-cased services (asterisk, wordpress, etc.) either already count
sites directly (wordpress) or aren't part of the multi-instance pattern, so
they just mirror is_installed() as 0 or 1. Wired into the whiptail
checklist, the non-whiptail plain-text fallback, and --status.
2026-08-10 16:10:31 +00:00
Claude c7f9e5caa1 Add a * marker next to the checkbox for already-installed services
The [installed] text label (previous commit) confirmed working from a
screenshot, but the checkbox itself stays unchecked for installed items by
design — checking it means "install/reinstall this on <Ok>", so
pre-checking every already-installed service would risk a mass reinstall
from just hitting Ok without manually unchecking each one.

Add a second, more immediate cue right next to the checkbox instead:
prefix the item's own tag with "*" when installed (whiptail's checklist
tag is the first column, directly after the checkbox). The "*" is
display-only — stripped back off the selected values before they reach
run_service, so dispatch is unaffected.
2026-08-10 16:03:19 +00:00
Claude 3e75c51d18 Fix "local: can only be used in a function" crash in the category menu
The dynamic checklist-sizing code added in the previous commit used
`local` for its variables, but the category menu loop it lives in is
top-level script code, not inside a function — `local` only works inside
one. Confirmed live: this broke the whiptail menu outright on first
`sudo ./setup.sh` run after pulling ("only be used in a function", then an
unbound-variable error under set -u since the assignment before it never
ran). Drop `local`; these are the same kind of plain loop-scoped variables
every other var in this loop (CHOSEN_CAT, SVCS, CHOICE, SELECTED) already
is.
2026-08-10 14:53:44 +00:00
Claude 3a3833596c Fix filebrowser crash-looping on permission denied opening its database
Confirmed from gtstef/filebrowser's own Dockerfile (_docker/Dockerfile):
the image runs as a fixed non-root user (adduser -u 1000 filebrowser;
USER filebrowser), not root and not remappable via PUID/PGID. The
installer's broad `chown -R $ACTUAL_USER:$ACTUAL_USER "$FB_DIR"` left the
bind-mounted ./data owned by $ACTUAL_USER (root, on a box where the
installer itself runs as root) — UID 1000 inside the container then had no
write access to it, so every start failed with "could not open database:
open /home/filebrowser/data/database.db: permission denied" and the
container crash-looped indefinitely (restart: unless-stopped kept retrying
every ~60s, matching the log timestamps this was diagnosed from).

Re-chown ./data to 1000:1000 specifically, after the broad chown so it
isn't clobbered back to $ACTUAL_USER.
2026-08-10 14:52:48 +00:00
Claude 0cf859704f Fix whiptail checklist silently dropping [installed] on long descriptions
Root cause of the "installed services not shown as installed" report,
confirmed from a screenshot: the [installed] marker was appended AFTER the
service description, and whiptail hard-truncates each checklist row to the
dialog's fixed width (78) with no ellipsis or other sign it happened.
fmd's description alone is 68 characters — adding "  [installed]" pushes
it to 81, past the width, so the marker silently fell off the end. fmd was
actually installed the whole time (confirmed via setup.sh's own pre-wizard
summary and the new --status flag); the checklist just never showed it.

Move the marker to the front of the tag instead, where a long description
can still lose its own tail to truncation but the install status — the
part that actually matters — always survives. Mirrored the same fix into
the non-whiptail plain-text fallback path for consistency.

Also size the checklist's listheight/height to the category instead of a
flat 14 rows: utilities alone has 35+ services, so anything past row 14
was only reachable by scrolling with no on-screen hint more rows existed.
Now scales with the category size, capped to what the actual terminal can
show (tput lines) so it can't request a dialog taller than the screen.
2026-08-10 14:47:35 +00:00
Claude 11a4e249b6 Add missing cancel option to 15 more multi-instance services; add setup.sh --status
Same bug as the previous filebrowser/fmd fix: vaultwarden, immich,
audiobookshelf, homebox, rustdesk, emby, meshcentral, traccar, lyrion,
actualbudget, mealie, joplin, jellyfin, unifi, and ntfy all showed "Manage
that install (update / full reinstall / cancel)" when re-run against an
existing install, but choosing "1) Manage" fell straight through into the
same unconditional fresh-install flow every time regardless of choice —
no way to actually cancel or update in place. Wired all 15 up to
prompt_reinstall_mode, matching the reference pattern in
services/mattermost.sh: update pulls + restarts the existing container
without touching config, cancel leaves the install untouched, fresh falls
through to the existing full-install flow unchanged.

Also add `setup.sh --status`: a plain-text listing of every service with
its install state, using the exact same is_installed() calls the whiptail
checklist's [installed] marker uses. Exists so "is X actually installed"
can be answered by reading terminal output directly, without depending on
a whiptail checklist screen where a narrow/resized terminal can truncate
the "[installed]" suffix off-screen with no visible sign that happened.
2026-08-10 14:39:34 +00:00
Claude 4123662571 Fix fmd's broken Docker image and add missing cancel option to two installers
fmd.sh pointed at nulide/findmydevice, which no longer exists on Docker
Hub — the project has moved twice (nulide/findmydevice ->
gitlab.com/Nulide/findmydeviceserver -> gitlab.com/fmd-foss/fmd-server) and
was rewritten from Node.js to Go+React along the way, confirmed against the
current upstream repo and its GitLab container registry. This means the
service never actually started for anyone who installed it before this fix
("pull access denied for nulide/findmydevice, repository does not exist").

Switch to registry.gitlab.com/fmd-foss/fmd-server:0 (GitLab's own registry
has no "latest" tag; ":0" tracks the current major release the same way
this repo's other services use a floating tag). The old FMD_ADMIN_PASSWORD
model is gone from the app too — replaced with FMD_REGISTRATIONTOKEN
(self-registration gated by a token instead of one shared admin login), and
the database path moved from /fmd/data to /var/lib/fmd-server/db.

Also: filebrowser.sh and fmd.sh both showed "Manage that install (update /
full reinstall / cancel)" when re-run against an existing install, but
choosing "1) Manage" fell straight through into the same unconditional
fresh-install flow every time — no way to actually cancel or update in
place, contradicting both the banner text and the documented
prompt_reinstall_mode contract (CLAUDE.md's "Update vs. fresh reinstall on
rerun"). Wired both up to prompt_reinstall_mode, matching the reference
pattern in services/mattermost.sh. The same gap exists in 15 other
multi-instance services (vaultwarden, immich, audiobookshelf, homebox,
rustdesk, emby, meshcentral, traccar, lyrion, actualbudget, mealie, joplin,
jellyfin, unifi, ntfy) — not fixed here, flagged for a follow-up pass.
2026-08-10 14:25:34 +00:00
Claude 32240e18f9 Fix ensure_coturn_user leaving callers in the wrong directory
install_coturn (services/coturn.sh) cd's into $DOCKER_DIR/coturn and never
restores the caller's original working directory. A consumer that chain-
installs coturn mid-flow (e.g. asterisk.sh, already cd'd into its own
install directory) returned from ensure_coturn_user still sitting in
coturn's directory, then went on to write its own docker-compose.yml/.env
there instead of its own directory — clobbering coturn's compose file and
leaving the consumer's directory without one. The consumer's later
`docker compose up --build` then failed with "Dockerfile: no such file or
directory", since the Dockerfile was correctly in the consumer's directory
but the misplaced compose file (and the build) were not.

ensure_coturn_user now saves/restores the caller's cwd around the
install_coturn call, fixing this for every consumer (asterisk, mattermost).

Also fold cloud-init.sh's contents into a collapsible README section so
it's copy-pasteable straight from the repo instead of requiring a separate
file download.
2026-08-10 03:10:26 +00:00
Claude 613625da09 Fix stale usage comment in cloud-init.sh
The header still told readers to paste the raw GitHub URL, left over from
before we confirmed provider user-data fields run pasted/imported content
directly rather than fetching a URL.
2026-08-10 02:35:34 +00:00
Claude 5cab7fe9c0 Correct cloud-init.sh usage instructions for real provider UIs
IONOS's User Data field takes a Script Type choice (Cloud Config vs Shell
Script) and runs the pasted/imported content directly rather than fetching
a URL. Update the README to say so, add DEBIAN_FRONTEND=noninteractive for
genuine unattended cloud-init execution.
2026-08-10 02:33:26 +00:00
Claude b2b4b6dd19 Add cloud-init.sh for provider install-script/user-data fields
IONOS Cloud Server, DigitalOcean, and Hetzner all offer an "install
script"/user-data field that runs as root with no TTY while the image is
still provisioning, so bootstrap.sh's interactive tail can't run there.

cloud-init.sh clones the repo unattended and drops a one-shot
/etc/profile.d hook that launches the normal whiptail setup.sh wizard on
the first interactive login, then removes itself.
2026-08-10 02:29:02 +00:00
Claude 666d280179 Document IONOS Object Storage pricing
Storage cost matches what was already known from IONOS chat support
(~$0.49/100GB/month). Found the two unknowns from IONOS's own published
price list rather than pricing-comparison sites, which had conflicting
numbers for the API-cost line: API requests (PUT/COPY/POST/LIST/GET/
DELETE) are free with no per-request charge, and outbound transfer is
free up to 2TB/month (shared across the whole IONOS contract, not scoped
to Object Storage alone) before tiered per-GB rates kick in. Relevant to
services/immich.sh's S3 storage engine.
2026-08-10 00:15:33 +00:00
Claude 2fb2a2980d Document the 6vCPU/8GB Tier 3 sizing plan
Settled stack: 3x Mattermost, 1x Traccar (down from 2x to buy back RAM),
2x each of ntfy/mealie/wordpress/actualbudget/audiobookshelf/emby
(music-only + everything)/filebrowser/fmd/homebox/joplin/rustdesk/
vaultwarden, 1x each of asterisk/security-dashboard/sms-inbound (all
three are singleton-by-design, no multi-instance support exists for
them). changedetection and magicmirror x6 dropped — the former never got
the full multi-instance retrofit, the latter's existing pattern caps at
3 instances.

Comes out to ~6.0GB of 8GB (~25% headroom) with RustDesk's relay for
screen sharing, or ~6.4GB (~20% headroom) with MeshCentral instead.
2026-08-10 00:09:49 +00:00
Claude 9e06ed4b83 Bake cross-service port collision avoidance into every service script
With 70+ services sharing a handful of common default ports (emby and
jellyfin both default to 8096, changedetection and frigate both default
to 5000, arm and nextcloud both default to 8080...), nothing previously
checked whether a service's default port was actually free on the host.
Whichever service installed second would silently write a compose file
claiming an already-held port, only failing at `docker compose up` time.

Adds two shared helpers to lib/common.sh:
- port_in_use PORT [PROTO] — true if something's already listening
- find_free_port VARNAME START [PROTO] — scans upward, writes back the
  first free port

Every service that publishes a fixed host port now scans before writing
docker-compose.yml, on every install (not just when adding an explicit
additional instance). On a normal single-install host this is a silent
no-op; it only changes behavior when something else already holds the
port.

- The 19 services already given multi-instance support this session had
  their port scan moved out of the "add instance" branch to run
  unconditionally, since the same collision risk exists on a plain first
  install.
- 20 more services with previously-hardcoded ports gained scanning for
  the first time: archivebox, arm, calibre-web, changedetection,
  drum-rhythm-game, gatus, n8n, nextcloud, onlyoffice, stirling-pdf,
  uptimekuma, portainer, iopaint (both GPU/CPU compose branches), koha
  (paired), syncthing (paired), wg-easy (paired, plus WG_PORT env so
  generated peer configs keep the right Endpoint), homeassistant
  (bridge-mode only — host mode can only warn), frigate and
  frigate-audio (multi-port stacks, moved together).
- caddy.sh is the deliberate exception: 80/443 stay fixed and only warn
  on collision, since silently moving Caddy itself would leave nothing
  listening where any client actually looks.
- authelia.sh needs no change — it has no published host port at all.
- Every service's standalone bootstrap fallback (sudo bash services/x.sh
  with no sibling files) got the same two helpers duplicated into its
  stub block, matching how every other shared helper is already handled
  there.

Documents the full pattern in CLAUDE.md's new "Port collision avoidance"
section, including the quoted-heredoc/backtick-escaping gotcha and the
network_mode:host limitation (can only scan ports the app takes as a
configurable env var).

Verified via bash -n on every changed file, plus functional runs seeding
occupied ports for each collision shape used here (single, paired,
multi-port stacks) and confirming the scan/shift and generated
compose/README output are correct — including the emby/jellyfin,
nextcloud/arm, and frigate/changedetection collision scenarios that
originally motivated this.
2026-08-09 23:55:49 +00:00
Claude b860a8b174 Add multi-instance support to 13 more services
Retrofits the standard multi-instance pattern (documented in CLAUDE.md)
onto actualbudget, filebrowser, fmd, homebox, immich, jellyfin, joplin,
lyrion, meshcentral, ntfy, rustdesk, unifi, and vaultwarden. First
instance of each keeps its original name/paths/ports unchanged; adding a
second instance prompts for a short name and auto-scans for free ports.

Service-specific handling beyond the base pattern:
- joplin, immich, unifi: dedicated Postgres/Mongo container per instance
  (not shared), matching the backup-isolation reasoning in CLAUDE.md.
- meshcentral, unifi: multiple fixed ports scanned/shifted together so
  they stay paired per instance.
- rustdesk: 6-port block shifted by a fixed offset per instance, since
  the image hardcodes its internal ports with no per-port env override.
- jellyfin: DLNA/discovery UDP ports only published for the first
  instance to avoid a host-wide fixed-port conflict.
- lyrion: first instance keeps network_mode: host (required for
  Chromecast/Squeezebox broadcast discovery); additional instances fall
  back to bridge networking with auto-scanned ports, trading away
  zero-config discovery since a second container can't also bind host
  networking's fixed ports.
- magicmirror.sh already had its own working multi-instance pattern
  (upfront instance count, numbered subdirs) and was left as-is.

Verified via bash -n on every changed file, plus scripted functional
runs (fake docker/ss) exercising first + second instance installs for
every port-scanning shape used here (single, dual-paired, quad-paired,
block-offset) and confirming dedicated per-instance DB naming and the
lyrion host->bridge compose output.
2026-08-09 22:53:40 +00:00
Claude 3fc20238af docs: document the multi-instance service pattern in CLAUDE.md
Establishes multi-instance as the default expectation for any service
that stores its own data and isn't inherently single-tenant, not an
opt-in special case -- matching the direction taken this session
(audiobookshelf, emby, mealie, traccar all just got it; mattermost and
wordpress already had it).

Documents the reusable pattern with a code skeleton (first instance
stays plain-named, adding a second introduces suffixed naming with no
further branching downstream), plus the three sharp edges found while
actually building it into four more services rather than just
theorizing about it:
- dedicated-per-instance databases over shared, and why (Kopia's
  generic backup stops a container to snapshot it, so a shared
  instance backs up and restores as one unit covering every instance
  at once -- this is the same reasoning already applied to
  wordpress.sh, now generalized)
- large port ranges shift by an offset instead of being scanned
  port-by-port, including the find -mindepth 1 gotcha discovered
  while building this into traccar.sh
- sidecar tooling that watches Docker labels host-wide (autoheal)
  needs the label itself scoped per instance, not just container names

Also states plainly: verify this kind of port/count logic by actually
running it, not by reading it -- both real bugs it references were
things code review alone missed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 21:58:25 +00:00
Claude d5d979ac31 Add multi-instance support to audiobookshelf, emby, mealie, traccar
Same pattern already established by services/mattermost.sh and
services/wordpress.sh: first instance keeps the plain name/paths/
ports exactly as before (zero behavior change for anyone with a
single instance already installed), and only choosing to add a second
introduces suffixed naming with its own directory, containers, and
ports.

- audiobookshelf.sh, emby.sh, mealie.sh: straightforward -- suffixed
  dir/container name, auto-scanned free host port(s) via `ss`, Caddy
  subdomain default suffixed to avoid collision. emby.sh's existing
  music-only mode is untouched, just correctly parameterized.
- traccar.sh: the harder one -- has its own dedicated Postgres
  container, an autoheal container, and a 150-port device-protocol
  range that can't be scanned port-by-port. Additional instances shift
  the whole range by 1000 (6000-6150, 7000-7150, ...) based on how
  many traccar/traccar-* directories already exist, which never lands
  on Asterisk's fixed ports the way the first instance's range does,
  so no exclusions are needed there. Also scoped the autoheal label
  per-instance (autoheal-traccar-<suffix>) -- autoheal watches by
  Docker label host-wide, not scoped to a compose project, so two
  instances sharing the generic "autoheal" label would each try to
  manage the other's container too.

Found and fixed two real bugs via testing before committing, not just
code review:
- The device-protocol range offset counted existing instances via
  `find $DOCKER_DIR -maxdepth 1 -name 'traccar*'`, which also matches
  $DOCKER_DIR itself if its own basename happens to start with
  "traccar" (true in my test harness, structurally possible in real
  use too) -- fixed with -mindepth 1.
- Verified port auto-scanning actually detects a simulated in-use
  port and increments past it, using a stateful fake `ss` rather than
  trusting the logic by inspection alone.

Verified end-to-end for all four: first instance unchanged from prior
behavior, second instance gets fully distinct dir/containers/ports,
and (traccar specifically) correct DB container, correctly-scoped
autoheal label, and correct shifted port range in the generated
compose file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 21:57:23 +00:00
Claude 5f36b14f93 mattermost: add PikaPods migration helper (DB dump + files import)
New opt-in prompt on fresh/new installs (skipped on "update" reruns,
where an existing instance is already in real use and importing over
it would be destructive): "Migrating from an existing Mattermost
instance (e.g. PikaPods)?" -- if yes, generates
migrate-from-pikapods.sh in the instance's own directory, same
generated-helper pattern as Immich's import-photos.sh.

Checked PikaPods' own docs before writing this rather than guessing
at their export mechanics: they expose per-pod SFTP (file access) and
a Database-access toggle that hands you an Adminer link for a full
SQL dump -- their own documented backup/migration flow is stop the
pod, SFTP the files, export the DB via Adminer. The generated script
assumes that shape (plain-text SQL dump + a files directory) and says
so in its header, including that PikaPods' exact SFTP layout wasn't
verified against a live pod so the files-argument path needs the
user's own confirmation.

What the script does: stops the mattermost container (leaves the DB
container running), drops and recreates the database owned by the
same existing role -- so .env's credentials are never touched or
regenerated, avoiding the "restored data, mismatched password" bug
class fixed elsewhere in this repo -- imports the dump via psql,
rsyncs the files directory into ./data, restarts. Requires typing
"YES" to proceed since it's destructive to whatever's currently in
the fresh instance's database.

Correctly parameterized per-instance: pulled from install_mattermost's
own MM_CONTAINER/DB_CONTAINER variables, so it's already correct for
either the first instance or an additional named one.

Verified end-to-end: prompt fires correctly at the right point in the
flow, generated script is syntactically valid, and the container
names/paths it's parameterized with match the actual instance being
installed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 21:45:26 +00:00
Claude 344bdf4a0f docs: settle on 2 WordPress sites, drop actualbudget
Final decision: actualbudget dropped and WordPress site count settled
at 2 (not 4) specifically to restore real headroom after dedicated-
per-site MariaDB made the 4-site case tight. ~1.09GB headroom (~27%)
now, back in the ideal 25-30% range.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 21:37:39 +00:00
Claude a5d57050b3 wordpress: switch to dedicated MariaDB per site (was shared)
Reconsidered after the shared-MariaDB design's real cost became clear:
Kopia's generic backup (services/backup.sh) stops a service's
container to snapshot it, so a shared MariaDB instance would back up
-- and would have to be restored -- as one unit covering every site at
once. Restoring just one site's database to an earlier point meant
restoring the whole shared snapshot to a temporary location first and
manually extracting that site's data back out, not a direct restore.

Each site now gets its own dedicated MariaDB container embedded in its
own docker-compose.yml (same pattern as services/nextcloud.sh) instead
of registering a database on a shared instance:
- Removed _wordpress_ensure_shared_db() and the wordpress-db/
  wordpress_net shared resources entirely.
- Each site's compose file gets a `db` service (container
  <site>-db) on an explicitly-named per-site default network
  (<site>_net), so wp-cli's one-off container reliably joins the
  right network without depending on Docker Compose's implicit
  naming convention.
- DB creation goes through the mariadb image's own MYSQL_DATABASE/
  MYSQL_USER/MYSQL_PASSWORD env vars on first boot (same as
  nextcloud.sh) instead of an imperative `docker exec mysql -e
  "CREATE DATABASE..."` against a shared container.
- Root and site DB passwords are both reused across reruns (read from
  the existing .env), verified via a real update-mode rerun.

Tradeoff, stated in both the script's header comment and the generated
per-site README: more RAM per site (~100-150MB for a full MariaDB
container instead of a slice of one shared instance) in exchange for
independent backup/restore. Data was already fully isolated either way
(separate database + user, always required since WordPress's schema
uses generic table names) -- the shared-vs-dedicated choice was only
ever about the container/process, not the data.

Re-verified end-to-end against the fake docker shim: distinct ports,
distinct dedicated DB containers/networks per site, correct compose/
.env structure, credentials preserved across an update-mode rerun.

docs/vps-sizing-recommendations.md: updated to match -- WordPress
capacity recomputed for dedicated-per-site MariaDB (~580MB headroom at
4 sites, ~976MB at 2, vs. the shared design's ~700MB/~950MB).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 21:23:24 +00:00
Claude e96a257d8f docs: record WordPress decision (2-4 sites, ecommerce-capable, Emby dropped)
Emby traded off for WordPress capacity rather than run alongside it —
still fully built and ready in services/emby.sh, just not part of the
current baseline. Updates the final RAM budget table to swap Emby for
the shared MariaDB + WordPress sites, and notes wg-easy/homebox/
audiobookshelf aren't included in that specific table since they
weren't part of the baseline as most recently stated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 21:16:00 +00:00
Claude a26d1831ee Add services/wordpress.sh — multi-site WordPress with shared MariaDB
New service: self-hosted WordPress, sized for running several
independent sites the way a hosting company would, not just one blog.

- Multi-site from the start: every site requires a name (no unnamed
  "first instance" special case like mattermost's — there's no
  backward-compat reason to special-case one here) and gets its own
  directory/container/port, but all sites share ONE MariaDB container
  (chain-installed on first site, reused by every other one) instead of
  a dedicated database container per site — same resource-sharing idea
  as services/coturn.sh, just scoped to WordPress's own sites rather
  than shared across different services. Each site gets its own
  database + user within that shared instance.
- E-commerce is just WooCommerce, a normal WordPress plugin — no
  separate infrastructure. PHP memory_limit/upload_max_filesize/
  post_max_size are pre-tuned (256M/64M/64M) so a product-catalog
  import doesn't hit default-image limits on the first try.
- wp-cli (official wordpress:cli image, run as a one-off container
  sharing the site's html volume) does the initial WordPress core
  install non-interactively — title, admin account — so there's no
  browser setup wizard to remember per site. Falls back to printing
  the exact manual command if the site wasn't ready in time.
- Auto-scans for a free host port per site (multiple sites can't all
  bind 8090), matching the "auto-scanned free ports for extras" idea
  already used by mattermost's multi-instance support.
- DB and admin passwords are reused across reruns (checked against the
  DB-password-regeneration bug class already fixed elsewhere in this
  repo, e.g. PR #265) — verified via a real update-mode rerun that the
  credential doesn't change.
- setup.sh: is_installed() gets a wordpress case — every site is named
  from the first one on, so there's never a plain $DOCKER_DIR/wordpress
  directory the default case could match against.
- README.md: added to the utilities services table + copiable list per
  CLAUDE.md's three-step rule for new services. Also fixed `coturn`
  being in the homelab row's prose but missing from the copiable list
  block below it — a pre-existing gap from when coturn.sh was merged.

Verified end-to-end via non-interactive dry runs against a fake docker
shim (no live daemon in this environment): 3 sites installed in
sequence get 3 distinct databases, 3 distinct auto-scanned ports, the
shared DB is only set up once, and an update-mode rerun preserves the
existing DB password rather than regenerating it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 20:57:18 +00:00
Claude 22a6366258 coturn: fix unescaped backticks corrupting generated README + stray error
The "Adding a new service that needs TURN" example in coturn.sh's
write_readme heredoc had one unescaped backtick pair (`sudo ./setup.sh
coturn`) while every other backtick in the same heredoc was correctly
escaped. Since write_readme's heredoc is unquoted (intentionally, so
$DIR-style interpolation works elsewhere in the file), bash treated it
as a command substitution: it actually tried to execute `sudo
./setup.sh coturn` at install time, printed "sudo: ./setup.sh: command
not found" to the terminal on every coturn install, and silently
dropped the intended text from the generated README.

Found while verifying the shared-coturn multi-consumer flow end-to-end
(coturn install -> asterisk + 2 mattermost instances all registering
concurrently) — confirmed working correctly otherwise: three distinct
credential files, no collisions, all three referencing the same host/
port, and reruns correctly reuse the cached credential instead of
regenerating.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 20:49:22 +00:00
Claude d2848cecc3 immich: add native S3 storage engine support for thumbnails/uploads
Adds an opt-in prompt to store Immich-managed data (thumbnails, encoded
video, new uploads) in S3-compatible object storage instead of local
disk, using Immich's native IMMICH_STORAGE_ENGINE=s3 — deliberately NOT
a FUSE-mounted bucket. Checked this against real reported issues before
implementing: Immich uses symlinks internally that S3 doesn't support
under FUSE (ENOSYS errors), and its startup does thousands of stat()/
read() calls that FUSE-over-network handles badly enough to crash the
mount under latency spikes as small as 100ms. Native S3 mode talks to
the bucket over the S3 API directly, sidestepping both problems.

Independent of the existing external-library strategy — an external
library (existing photos indexed read-only, e.g. over a VPN mount) is
a separate mount either way and works the same regardless of where
Immich's own managed data lives, since S3 mode only replaces
UPLOAD_LOCATION.

- New prompts: bucket, region, endpoint (for non-AWS S3-compatible
  providers — auto-sets S3_FORCE_PATH_STYLE when given), prefix, access
  key ID, and secret key (read via `read -rs` so it doesn't echo; left
  blank with a warning under UNATTENDED, since there's no sane default).
- Refactored the docker-compose.yml generation from two near-duplicate
  heredocs (with/without external library) into one with composable
  volume-line variables, to avoid quadrupling the duplication once S3
  was added as a second axis.
- Skips creating local upload-location subdirectories entirely in S3
  mode (thumbs/upload/backups/library/profile/encoded-video) — Immich
  manages that structure inside the bucket itself.
- .env now gets chmod 600 (previously ungated) — more pointed now that
  it can hold an S3 secret key, not just the DB password.
- Generated README documents the S3 setup and carries the FUSE-mount
  warning forward so a future reader doesn't try that route instead.

Verified both the non-S3 baseline (unchanged output) and S3 mode
end-to-end via non-interactive dry runs — correct .env, correct
compose volumes, no local upload dirs created, 0600 permissions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 20:06:50 +00:00
Claude 033ffeee48 docs: drop lyrion from the VPS plan, add emby music-only, update swap notes
lyrion was ruled out for two protocol-level reasons Authelia can't work
around (single shared server password, and SlimProto has no auth of its
own for Authelia's HTTP-only forward_auth to gate) — emby covers music
instead, with real per-user library access. Also updates the swapfile
rule of thumb to reflect it now being a default for every install
rather than an Asterisk-droplet-specific behavior, and adds a final
RAM budget table/verdict for the full confirmed service list.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 18:14:32 +00:00
Claude ad38b96cfe Make swapfile a default for every install, not just Asterisk droplets
Extracts the swapfile logic out of services/asterisk.sh (previously
DigitalOcean-droplet-gated) into lib/common.sh's ensure_swapfile() —
provider detection was never really the point, the actual condition
that matters is "modest RAM, no swap yet," which applies just as much
to a non-DO VPS running several Docker services at once as it did to a
single-purpose droplet.

- lib/common.sh: new ensure_swapfile(), same fallocate/mkswap/fstab/
  swappiness logic as before, threshold raised from 2048MB to 4096MB
  (a 4GB box running a full service stack is exactly the case that
  motivated this change — the old threshold would have skipped it).
- services/base.sh: calls it unconditionally so every install gets the
  same check regardless of which other services get chosen.
- services/asterisk.sh: swapfile call is no longer gated behind
  IS_DO — calls the shared helper directly. Kept a standalone-mode
  stub (same pattern as this file's other stubbed helpers) so
  `sudo bash asterisk.sh` with no base.sh in the picture still gets
  it. Idempotent either way: a box that already has swap, or already
  got it from base.sh earlier in the same run, no-ops immediately.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 18:13:37 +00:00
Claude 7a66ef2e2f emby: add music-only setup mode with per-user library access guidance
Prompts whether this install is music-only (changes the default folder
to ~/music and the prompt wording — Emby has no compose/env flag for
"music-only", library types are chosen in its own web setup wizard, so
this is guidance plus a sane default, not a functional restriction).
Generated README walks through adding only a Music library and, the
actual reason to pick Emby for this role over Lyrion, per-user library
access under Dashboard → Users → Access — LMS/Lyrion has no equivalent,
just one shared server-wide password.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 18:07:09 +00:00
Claude 1b54fdbeb3 docs: correct VPS service recap — drop portainer/syncthing, confirm audiobookshelf
portainer and syncthing were never actually agreed to, and
audiobookshelf-with-remote-home-library access (over NetBird or wg-easy)
was confirmed, not just floated — lyrion remains the only still-open item
from that same idea.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 16:57:11 +00:00
Claude 0fde1a2542 docs: add VPS sizing reference and recap of the IONOS box's planned services
Records the sizing methodology worked out for this repo's services (RAM
as the binding constraint, per-service budget ranges, when a swapfile
matters) against two real VPS plans, plus a recap of the full service
list planned for the 4 vCPU/4GB/120GB IONOS box: the core stack, the
utility adds, the NetBird-for-remote-access vs wg-easy-for-local-testing
split, and what was deliberately left out and why.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 16:53:24 +00:00
Claude 4a0ec63a04 wg-easy: hub-routed peer mesh by default, plus SSH-alias sync script
Pins WG_DEFAULT_ADDRESS/WG_ALLOWED_IPS explicitly instead of relying on
wg-easy's own internal default, so every client is created with
0.0.0.0/0 Allowed IPs — client-to-client traffic already routes through
this VPS automatically with no per-pair config, confirmed against
wg-easy's own docs/issue tracker as the documented way to get this
behavior. Also adds an opt-in, additive-only UFW rule to reach SSH over
the VPN subnet (never touches the existing public SSH rule — narrowing
that is left as a manual step so a misconfigured VPN can't lock anyone
out), and a self-contained sync-ssh-aliases.sh companion script that
reads connected peers straight off the live WireGuard interface (`wg
show`, not wg-easy's own undocumented/unstable HTTP API or its
internal storage format) to generate ~/.ssh/config Host aliases.

This is a hub-and-spoke design, not true peer-to-peer mesh: the VPS is
a single point of failure for inter-peer connectivity specifically
(not just VPS access), and unlike NetBird/Tailscale there's no direct
P2P fallback or centralized identity for multi-person key management.
Documented in the generated README as a real tradeoff, not a full
substitute for NetBird once more than one box or one person is
involved.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TBtExJcqxnokyZZKmphdug
2026-08-09 12:44:46 +00:00
Claude d4ca8277a6 security-dashboard: add Calls & Texts tab for PSTN calls and SIP/SMS texts
Surfaces pstn-trunk.sh's existing pstn-trunk-calls.log (already recording
every PSTN call's numbers, just never shown on the dashboard) plus two new
metadata-only logs: sip-messages.log for internal SIP MESSAGE
deliveries/denials (asterisk.sh) and pstn-sms.log for SMS-over-SIP arrivals
(pstn-trunk.sh). No message bodies are ever logged. The dashboard reads all
three via a new /api/pstn-calls and /api/comms-texts pair, sharing a common
bounded tail helper with the Security Log parser.
2026-08-08 02:32:42 +00:00
Claude e04b133e87 Add cross-platform duplicate file finder (tools/dedupe-finder.py)
Content-hash based dedup (size -> partial SHA-256 -> full SHA-256), not
filename matching, so identically-named files with different content are
never confused for duplicates and differently-named files with identical
bytes always are. Stdlib-only Python so it runs unmodified on Windows
10/11 and Debian-flavored Linux (Ubuntu, Mint). Supports photo/music/
video/docs extension categories or custom extensions, JSON/CSV reports,
and optional delete/move/hardlink cleanup actions that default to a dry
run and require --yes to actually touch files.
2026-08-05 12:34:26 +00:00
Claude f731efa2fa Fix DB/admin password regeneration on rerun in 5 services
Same bug class just fixed in mattermost.sh: immich, joplin, koha,
mail-archiver, and nextcloud all generated a fresh random
DB/admin password on every single run with no check for an
existing one. Each backs its database with a persistent volume, so
Postgres/MariaDB keeps the password from its first init while the
freshly overwritten .env (or config-main.env for koha) no longer
matches it — any rerun would have locked the app out of its own
database. koha, mail-archiver, and nextcloud also regenerated an
app-level admin login password the same way.

Found by cross-referencing every service with a DB password against
which ones actually guard reuse on rerun (only traccar.sh did,
already correctly) rather than waiting to be told about each one
individually.

Fix mirrors traccar.sh's existing pattern: read the password back out
of the existing .env/config file if present, only generate fresh when
there's genuinely nothing there yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NQkdAn3iG5A4WoqU9FHMaN
2026-08-04 17:08:55 +00:00
Claude cf3d4bf6de Extract coturn into a shared service; add Mattermost multi-instance support
Asterisk and Mattermost each used to embed their own dedicated coturn
container (network_mode: host), and their default relay port ranges
overlapped by ~100 UDP ports — running both on one box meant a
coin-flip over which service's active call lost its media relay.

- services/coturn.sh: new shared TURN/STUN relay, one instance for
  every consumer instead of one each. Runs --lt-cred-mech with a
  SQLite user database (not --use-auth-secret — coturn doesn't
  support both auth mechanisms on one instance at once, confirmed via
  coturn's own upstream docs/issues) so each consumer gets its own
  dedicated username/password without stepping on any other's.

- lib/common.sh: ensure_coturn_user() — chain-installs coturn.sh on
  first need (same declare -F guard pattern as the existing
  asterisk -> security-dashboard chaining) and registers/reuses a
  per-consumer credential, mirroring configure_caddy_for_service's
  out-param convention.

- services/asterisk.sh: _asterisk_write_compose gains a
  USE_EMBEDDED_COTURN flag. New installs use the shared service;
  existing installs keep their dedicated coturn exactly as-is on
  every "update" (detected from the existing compose file before
  regenerating it, so a rebuild can never silently drop the container
  its own .env TURN_PASSWORD still points at) and only switch on an
  explicit "fresh" reinstall, with a warning first.

- services/mattermost.sh: same embedded/shared coturn handling, plus
  genuine multi-instance support (separate dir/containers/DB/ports per
  instance, auto-scanned free ports for extras) for real isolation
  between groups, as opposed to Team Edition's built-in Teams feature.
  Calls plugin TURN config switched from the HMAC "TURN Static Auth
  Secret" field to the verified "ICE Servers Configurations" JSON
  field, which accepts the same fixed username/credential shared
  coturn issues. Also fixes a latent bug found while adding proper
  update-mode detection: DB_PASS/MM_SECRET were regenerated on every
  single rerun with no existing-install check at all, silently
  breaking Postgres auth on any reinstall.

- CLAUDE.md: documents the ensure_coturn_user pattern (including the
  auth-mechanism constraint and the embedded-coturn migration-safety
  rule) for any future service that needs TURN.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NQkdAn3iG5A4WoqU9FHMaN
2026-08-04 16:32:08 +00:00
Claude 883f0d2557 Document restricting DMs to teammates in Mattermost's README
Answers a real gap: Teams alone don't limit who can Direct Message
whom — that's a separate System Console setting
(TeamSettings.RestrictDirectMessage), free in Team Edition. Documented
it in the Teams section along with the caveat that it only filters the
DM picker UI, not a hard boundary (existing DMs unaffected, multi-team
users can still DM across all their teams).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NQkdAn3iG5A4WoqU9FHMaN
2026-08-04 13:57:51 +00:00
Claude 2fc49fa43e Document Teams setup in Mattermost's generated README
Team Edition includes multiple Teams natively (no Enterprise license
needed), but nothing in the generated README said so or explained how
to create one. Added a Teams section covering creation, adding
members, and multi-team membership, plus a one-line pointer in the
install summary.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NQkdAn3iG5A4WoqU9FHMaN
2026-08-04 13:48:49 +00:00
Claude b7f69090a6 Sync backup.conf/README to a spare box; make DR bring-up fault-tolerant
- dr_bringup.sh: bound every kopia call and docker compose up with a
  timeout so one stuck service can't stall the rest of the batch, and
  only exit non-zero if literally nothing came up — a partial recovery
  is a partial success, not a failed run.
- backup_kopia.sh: optional DR_SYNC_HOST/DR_SYNC_PATH in backup.conf
  scp's backup.conf + README.md to a spare box over SSH after every
  successful backup, so dr_bringup.sh is ready there with no manual
  copy step.
- backup.sh: prompts for the spare's SSH destination, verifies
  connectivity at install time instead of failing silently at 2am, and
  writes ~/docker/backup/README.md (this service never had one) so the
  synced copy documents every command listed above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NQkdAn3iG5A4WoqU9FHMaN
2026-08-04 13:40:53 +00:00
Claude ed7270dcf2 Add unattended DR bring-up script for the backup service
restore_kopia.sh is interactive and one-service-at-a-time, which doesn't
scale to standing up a cold spare box quickly during a real outage.
dr_bringup.sh restores every service's latest snapshot (or one named
service) and runs docker compose up -d with no prompts, so a full-stack
recovery is one command instead of N interactive restores.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NQkdAn3iG5A4WoqU9FHMaN
2026-08-04 13:31:42 +00:00
Claude b2d064573f Move ai-stack and paintplus vendored source under vendor/
Matches the existing vendor/easy-asterisk convention (used by
services/asterisk.sh) instead of two one-off top-level directories that
cluttered the repo root and didn't look like anything else next to
setup.sh, lib/, services/, extras/. Only the two services' own SRC_DIR
path resolution and header comments needed updating — nothing else in
the repo referenced the old ./ai-stack / ./paintplus paths.

Also documents vendor/ in README.md's Layout section.
2026-08-04 12:55:09 +00:00
Claude d6092dde4c Move static service docs into services/<name>.md companion files
Applies the new write_readme companion-doc convention to ai-stack,
paintplus, and kyber-launcher: install-time-invariant content (usage
walkthroughs, service tables, troubleshooting) moves out of the .sh
heredocs into sibling services/<name>.md files, leaving only genuinely
install-specific content inline.

- services/paintplus.md: config/cloud/GPU/ai-stack-backend/update/Caddy
  sections, picked up automatically via write_readme's companion-doc
  support.
- services/ai-stack.md: roles, GPU switcher, service URLs, cloud LLM
  provider setup, update, Caddy. ai-stack.sh can't use write_readme
  directly (its POST-INSTALL-NOTES.md filename deliberately avoids
  colliding with the vendored app's own README.md in the same
  directory), so it appends the companion file manually.
- services/kyber-launcher.md: the full SWBF2 (2017) + Kyber walkthrough,
  moved out of the root README's "Gaming scripts" section (which now
  just points here). kyber-launcher.sh now calls write_readme to deploy
  it to ~/.local/share/kyber/README.md, fixing a stale in-script pointer
  to a README section that no longer exists.
2026-08-04 12:34:37 +00:00
Claude 52b4606ab5 Add optional per-service companion doc files to write_readme
Any services/<name>.md next to services/<name>.sh gets appended to the
generated ~/docker/<name>/README.md automatically, with no changes needed
to the calling install_<name>() function. Keeps install-time-invariant
walkthroughs (third-party UI linking steps, multi-account setup) out of
the compose/README heredocs, which should stay focused on values chosen
during install.

Add services/traccar.md as the reference example: documents the
per-user Connections-tab linking needed for ntfy notifications to reach
non-admin accounts, and why TEST CHANNELS can return success without
actually sending anything (Traccar's SMS notificator silently no-ops
when the logged-in user's Phone field is empty).
2026-08-04 12:14:32 +00:00
Claude 2d7b9f8d56 traccar: scripted, idempotent ntfy push-notification setup
Traccar has no native ntfy integration, but its "SMS" notification
channel is a generic HTTP webhook (sms.http.* config) under the hood —
pointing it at ntfy's JSON publish API instead of a real SMS gateway is
a well-known trick. Adds an opt-in prompt during install (or reruns)
for an ntfy server URL — self-hosted on this box, a different server
entirely, or public ntfy.sh — a topic name, and optional username/
password if that topic needs auth.

Everything lands in .env only: NOTIFICATOR_TYPES, SMS_HTTP_URL,
SMS_HTTP_TEMPLATE, and (if given) SMS_HTTP_USER/SMS_HTTP_PASSWORD.
traccar's env_file: .env already passes these straight through to the
container with no docker-compose.yml changes needed, since Traccar's
env-var config naming convention already matches these key names
exactly. The topic name itself isn't consumed by Traccar (it's set
per-user via the web UI's Phone field), so it's stored as a `#
NTFY_TOPIC=` comment line purely so reruns can recall it and so it's
discoverable on disk.

Reruns default every prompt to whatever was configured last time
(reading it back out of the existing .env), so accepting the defaults
is a no-op — the same non-destructive-by-default pattern already used
for the DB password.

Fixed a heredoc bug caught while testing this: `${_NTFY_ENV_BLOCK}
TRACCAR_ENV` on one line never matches the heredoc terminator, because
bash matches heredoc delimiters against the literal source line before
variable expansion, not after — regardless of what the variable
expands to. Moved the terminator to its own line, matching the
${_CADDY_NET_SECTION} / TRACCAR_COMPOSE pattern already used a few
lines above for the same reason.

Verified: bash -n, and four full install runs (ntfy disabled, enabled
without auth, enabled with auth, and a rerun accepting defaults) with
`docker compose config` confirming SMS_HTTP_TEMPLATE's JSON and all
other values resolve correctly through env_file with no compose-level
escaping needed.
2026-08-04 10:42:09 +00:00
Claude e1129b7f27 traccar: never recursively chown the Postgres data directory
Confirmed live: after re-running the installer to pick up the
caddy_net/port fixes, Traccar crash-looped with
"FATAL: could not open file \"global/pg_filenode.map\": Permission
denied" — a Postgres-side error, not a Traccar or Caddy problem.

install_traccar() had two `chown -R $ACTUAL_USER "$TRACCAR_DIR"` calls
(one via ensure_docker_dir_ownership at the top, one explicit near the
end) inherited from the original H2-only script, where that was safe —
everything under the directory (logs/, data/, config/) was meant to be
host-user-owned. Once db/ started holding Postgres's own data files
(owned internally by whatever uid the postgres container runs as, not
$ACTUAL_USER), both of those recursive chowns reassign db/'s contents
to $ACTUAL_USER on every rerun, and Postgres can no longer read its own
files afterward.

Replaced both with non-recursive/scoped chowns that never touch db/:
the top-level directory itself, docker-compose.yml, and .env directly,
plus a separate `chown -R` limited to logs/ and data/ (which are
Traccar's own app-writable directories and always safe to reassign).
2026-08-03 21:34:15 +00:00
Claude 2ced57db55 traccar: also exclude Asterisk's SIP ports (5060 tcp+udp, 5061 tcp)
The previous fix only excluded 5038 (AMI). Confirmed live: after
attaching caddy_net, `docker compose up -d` still failed —
"failed to bind host port 0.0.0.0:5060/tcp: address already in use" —
because Asterisk (network_mode: host) also owns 5060 (SIP, tcp+udp)
and 5061 (SIP TLS, tcp), both inside Traccar's 5000-5150 range.

Asterisk gets priority on all three ports; Traccar's range just skips
them. Audited every other network_mode: host service in the repo
(caddy, homeassistant, kyber-server, lyrion, mattermost, watchyourlan,
wolf-pair, wolf) — none of them land in 5000-5150, so Asterisk is the
only conflict to account for.
2026-08-03 21:01:22 +00:00
Claude 78692c5999 traccar: exclude Asterisk's AMI port (5038/tcp) from the published range
Confirmed live: on a box also running Asterisk from this repo,
`docker network connect caddy_net traccar` failed with "failed to bind
host port 0.0.0.0:5038/tcp: address already in use". Asterisk runs with
network_mode: host (services/asterisk.sh), so its AMI (port 5038,
hardcoded in services/sms-inbound.sh) binds directly on the host's
network stack — no Docker NAT involved. Traccar's docker-compose.yml
published the entire 5000-5150 range for device protocols, which
needs Docker to also bind host port 5038 for its own port-forwarding,
directly colliding with Asterisk's AMI on any box running both
services from this repo.

Split the TCP range into 5000-5037 and 5039-5150 to skip that one
port; left UDP as a single 5000-5150 range since AMI is TCP-only.
2026-08-03 20:56:41 +00:00
Claude 737bc873d1 traccar: fix stale admin@admin.com/admin default-login messaging
Current Traccar images ship with no built-in account at all — the login
screen's Register flow creates the first user, and that user is
automatically made admin. The admin@admin.com/admin default our messages
still quoted belongs to older Traccar versions and no longer exists,
so anyone following our own output would try that login and fail.

Updated the dry-run summary, README, and final on-screen message to
describe the real flow, and to flag that self-registration stays open
to anyone who reaches the server until it's turned off (Settings →
Server → Permissions), since that's a real exposure window on a
freshly-installed instance with no way to lock it down at config time.
2026-08-03 18:56:02 +00:00
Claude 1fc0a6edfe Apply the local/remote Caddy mode resolution to every service, not just traccar
traccar.sh's caddy_net wiring was fixed to mirror configure_caddy_for_service's
own mode resolution (CADDY_MODE from site config, then a local ~/docker/caddy,
then the legacy CADDY_REMOTE_HOST var) instead of only checking for the local
directory. That same bare directory check was copy-pasted into the caddy_net
wiring of every other Docker service in the repo, so a site with Caddy on a
different box would silently fail to join any of their containers to caddy_net
during setup (or, for homeassistant/koha, only get half the wiring right).

Applied the same fix mechanically across all 37 services using the standard
_CADDY_NET_BLOCK/_CADDY_NET_SECTION pattern (verified identical text via
scripted diff before touching any of them), plus by hand for:
- homeassistant.sh and koha.sh, which use their own differently-shaped
  variables (HA_CADDY_NET_LINES / _CADDY_NET_ENTRY) for the same decision
- paintplus.sh and ai-stack.sh, which do a live `docker network connect`
  instead of a compose network block
- watchyourlan.sh, whose Caddy note was worded for local-only setups

sms-inbound.sh got more than a mode swap: its Caddy wiring was hand-rolled
(not routed through configure_caddy_for_service) and had no remote-Caddy
path at all — a remote Caddy box would get a misleading "Caddy isn't
installed here" message instead of a snippet. Added
_sms_write_caddy_snippet(), mirroring the snippet-file pattern
configure_caddy_for_service uses everywhere else, and pointed the firewall
gate at the same three-way mode instead of a two-way dir check.

Verified: bash -n across all of services/*.sh, a scripted check that every
touched file has exactly one _CADDY_MODE resolution and no leftover bare
`[ -d "$DOCKER_DIR/caddy" ]` feeding a caddy_net decision, and spot-checked
docker compose config renders (traccar, mattermost) confirming the ${VAR}
interpolation and multi-service usage sites still resolve correctly.
2026-08-03 18:54:02 +00:00
Claude 17201b647f Sync drum-rhythm-game service with upstream's Docker deployment
The game repo grew from a single genre/song count to 18 genres (119
synth-orchestra songs + 120 drum patterns), gamepad remap, and multiplayer,
and added its own Dockerfile/nginx.conf/.dockerignore specifically so only
index.html gets served. This service had fallen behind on both fronts: it
bind-mounted the whole cloned repo into nginx:alpine, publicly serving
README.md, CLAUDE.md, DEPLOYMENT.md, the Dockerfile itself, and old
versions_to_compare/*.html snapshots alongside the game. Build from the
repo's own Dockerfile instead (matching its .dockerignore) so only the game
is served, with gzip and /healthz along for free, and refresh the written
README's stale feature description.
2026-08-03 18:50:04 +00:00
Claude 6554256e8e traccar: drive database config from .env instead of a static XML file
The previous fix still baked database.user/database.password directly
into config/traccar.xml, duplicating the secret that .env already held
and leaving a second, unmanaged copy of it on disk.

Traccar supports reading its config from environment variables
(CONFIG_USE_ENVIRONMENT_VARIABLES=true, confirmed against the official
traccar/traccar docker/compose/traccar-mysql.yaml reference). Use that:
DATABASE_DRIVER/URL/USER/PASSWORD are now set in the compose file via
${POSTGRES_*} interpolation from .env, so .env is the only place the
credentials live — drop config/traccar.xml and its volume mount
entirely, matching the official reference example.

Also switched to the official reference healthcheck (wget against
/api/health, 1h start_period) — a real endpoint on real hardware rather
than a guessed /dev/tcp probe against an unverified image's toolset —
and added the interval/start-period env vars to the autoheal container
to match, while keeping the container scoped to just Traccar via the
autoheal=true label instead of the reference's host-wide "all".

Verified with `docker compose config` (both with and without a local
Caddy directory present) that the ${POSTGRES_DB}/${POSTGRES_USER}/
${POSTGRES_PASSWORD} references resolve correctly from .env with no
warnings.
2026-08-03 18:46:23 +00:00
Claude 5a329cdd48 traccar: resolve Caddy local/remote mode like configure_caddy_for_service does
The caddy_net wiring for the new db/traccar containers was gated only on
whether ~/docker/caddy exists locally, so it didn't account for a site
where Caddy runs on a different box (CADDY_MODE=remote or the legacy
CADDY_REMOTE_HOST var, set with no local Caddy directory). Resolve the
mode the same way configure_caddy_for_service does — explicit CADDY_MODE
first, then the local directory, then CADDY_REMOTE_HOST — so caddy_net is
only joined when Caddy is actually local. A remote Caddy reaches Traccar
via this host's published 8082 port regardless, so no other change is
needed for that path.
2026-08-03 18:43:02 +00:00
Claude 09eff1e397 Fix Traccar install: add PostgreSQL database and autoheal
Traccar's docker image no longer bundles the H2 driver, so the
generated traccar.xml (org.h2.Driver / jdbc:h2:...) failed at startup
with no working database. Add a postgres:15-alpine db container with
a healthcheck, point traccar.xml/POSTGRES_* at it via .env, and gate
traccar's startup on db being healthy.

Also add a willfarrell/autoheal container scoped to just the traccar
container (via the autoheal=true label) that restarts it if its own
TCP healthcheck on 8082 fails.

Reruns reuse the existing DB_PASS from .env instead of generating a
new one, since the postgres volume keeps the original password from
its first init.
2026-08-03 18:02:08 +00:00
Claude 9390382bf4 Add Sipnetic QR-scan provisioning and a kiosk client installer download
Two concrete asks: "a QR code generator for sipnetic... displayed on the
security website" and a real download for the baresip kiosk client
instead of just pointing at docs.

Sipnetic QR: ea_device_sipnetic_string() builds Sipnetic's own documented
account-string format (n=/u=/d=/p=/dt=, semicolon-separated, from
https://www.sipnetic.com/qr-codes) from the exact same data the existing
device-details panel already reads back out of pjsip.conf -- unlike
ea_device_provisioning()'s deliberately generic plain-text file (written
because no vendor XML format could be verified), this one has real
documentation to build against. Rendered client-side with
davidshimjs/qrcodejs (MIT, wraps Kazuhiko Arase's original QRCode for
JavaScript) embedded verbatim with its license header intact -- no CDN
call, no new Python dependency, same self-contained approach as the rest
of this page. New "Sipnetic QR code" button on each extension's detail
panel; verified the embedded library actually renders (not just parses)
via a real jsdom run producing real QR cell output.

Kiosk installer download: _secdash_copy_kiosk_installer copies
vendor/easy-asterisk/easy-asterisk-v0.10.0.sh alongside the deployed app
at install time (both update and fresh-install paths) and a new
/download/kiosk-client-installer.sh route serves it -- linked directly
from the Ring Groups card's help text instead of only being reachable by
manually finding the file in the repo. Copied at install time rather
than read live, since a standalone run of this one file (no full repo
clone -- explicitly supported here) has no vendor/ directory to read
from; missing source degrades to a clean 404 on that one link, not an
install failure.

Verified: bash -n, py_compile on the extracted embedded app.py, node
--check on the extracted embedded JS (including the newly-embedded QR
library), and the account-string output checked directly against
Sipnetic's own documented example format.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-29 03:30:13 +00:00
Claude 646b7b6fde Let an existing Ring Group's type/timeout be edited, not just set at creation
Rename, member add/remove, and DID assignment could all already be
changed after a room existed -- type and timeout could only be set once,
at creation, with no way to flip an existing Ring group to Page (or back)
without deleting and recreating it, losing its members/DID assignment in
the process. Reported directly: "I can't edit the ring group to change
it to a page group."

Turns the Timeout/Type columns into inline-editable controls (matching
the same select/input the creation form already uses) with a Save button
per row, backed by a new ea_update_room_settings() that rewrites just
those two fields in rooms.conf, leaving name/members/DID untouched --
same read-modify-write pattern ea_rename_room() already uses.

Verified: bash -n, py_compile on the extracted embedded app.py, node
--check on the extracted embedded JS.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-29 03:03:15 +00:00
Claude ca750f5432 Document Ring vs Page and auto-answer kiosks, in the dashboard itself
The user asked for the security webpage to actually explain how to set
these up, not just this conversation. Adds a "How this works" disclosure
to the Ring Groups card (matching the existing help-block pattern already
used on Extensions and Personal numbers) covering:

- Ring vs Page's actual difference (first-answer-wins hunt group vs
  Asterisk signaling auto-answer via SIP headers).
- Mixing an auto-answering device with normally-ringing phones needs no
  new group type or per-member setting -- auto-answer lives in the
  device's own SIP client config, so a plain Ring group already dials
  everyone at once and an auto-answer-configured device just picks up
  faster than a human can.

New docs/kiosk-paging-setup.md walks through the concrete path for a
dedicated always-on auto-answer device: Easy Asterisk's own baresip-based
"kiosk" client, installed on a separate small Linux machine (not the
Asterisk server's Docker container -- the vendor script explicitly skips
baresip install when it detects Docker locally). Also notes plainly that
a browser/WebRTC auto-answer option doesn't exist on this server today
(no WSS transport configured) and would be new infrastructure work, not
a quick addition.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-28 23:57:42 +00:00
Claude eaca45c83c Read domain/TURN from Easy Asterisk's own config, and offer a settings download
The panel kept showing "no domain set" and "TURN not configured" on a box
that has both. The cause was reading the wrong file: I had it reading the
host-side .env, a 600 file needing an ACL grant, a ProtectSystem exception
and a systemd Environment= line to even locate — three things that each had
to be right, and weren't.

The vendored web admin gets this right because it reads
/etc/easy-asterisk/config, written by the container's own entrypoint. That is
mounted on the host inside ASTERISK_EA_CONFIG_DIR, a directory this dashboard
is already granted read access to for rooms.conf. So it now reads that first,
falls back to .env, and finally to `docker exec cat` of the same file — each
source only filling gaps the previous left. No new permissions, and the value
shown is what Asterisk is actually running with rather than what the
installer asked for.

Also adds the requested provisioning file: a "Download settings" link per
extension serving a plain text file with server, username, password, display
name, transport, port, SRTP, a ready-made SIP URI, and TURN server/user/
password. Deliberately not a vendor-specific format — Sipnetic, Linphone,
Zoiper and Groundwire each want a different one and none could be verified
from here, and a confidently-wrong .xml is worse than a file you can read. It
warns in-file when the extension is UDP-only, when the cert is self-signed,
and that it contains a password.

The panel gains the SIP URI and says when the address shown is the host IP
rather than a configured domain.

Fixes a JS syntax error introduced with that note: an apostrophe escaped for
Python's benefit left a bare quote inside a single-quoted JS string, which
broke the whole page script — the table rendered empty. Now checked properly
by extracting the script and running `node --check` over it, which catches
this class of fault directly instead of inferring it from missing elements.

Verified with only Easy Asterisk's config present and .env absent entirely —
the state that failed before: domain, TURN server, user and password all
resolve, the download serves with the right filename, and the page has no
errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NAddJGE1G6eGaPzmScG5Vh
2026-07-28 20:19:46 +00:00
Claude 1b3e7a6110 Fix extensions created from the dashboard, and show details by their row
Three separate faults, reported together as "the old web admin made
extensions correctly and this doesn't".

1. Every write through ea_docker_write left pjsip.conf owned by root. `tee`
   runs as root inside the container while Asterisk runs as `asterisk` and
   expects to own its own config — the vendored admin's add_device() chowns
   it back immediately after writing, and services/asterisk.sh's device
   migration does too. This was the one writer in the project that didn't,
   and is the most likely reason an extension created here behaves
   differently from one created in the vendored admin. Now chowned after
   every write, with a scoped sudoers entry for it, and a warning logged if
   the chown itself fails rather than passing silently.

2. The dashboard's update path rewrote the systemd unit and then restarted
   the service without daemon-reload, so systemd kept running the cached
   unit. Any Environment= line added since the last FRESH install was written
   to disk and ignored — which is exactly how a box with DOMAIN_NAME and TURN
   both set in .env still reported "no domain set" and "TURN not configured".

3. The connection panel rendered at the top of the card, so on any table long
   enough to scroll, the answer appeared off-screen above the row that was
   clicked. It is now a table row injected directly beneath its own
   extension, toggled by the same button, and it survives the transport and
   password actions by reopening after the reload they trigger.

Also: LAN devices now get ice_support=yes when a TURN server is configured,
matching the vendored admin (a device given TURN credentials but no ICE can't
use the relay); the panel reports the device type, so a Mobile extension can
be confirmed as such; and an unreadable .env now says which of "path not
set", "file missing", "permission denied" or "TURN_SERVER empty" applies
instead of the flat "not configured" that covered all four.

Verified: the three .env failure states each produce their own message; the
detail row lands directly after its anchor, only one is ever open, it clears
on re-render and reopens rather than sticking closed; no duplicate or missing
element IDs and no JS errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NAddJGE1G6eGaPzmScG5Vh
2026-07-28 20:08:49 +00:00
Claude 4f216fc8bb Correct two docs left describing SMS as ntfy push, not AMI delivery
services/sms-inbound.sh was rebuilt to deliver texts into Asterisk over AMI,
landing in the softphone's own thread, and ntfy was dropped from that path
entirely. Two pieces of prose written against the earlier design survived and
now contradict the working implementation:

- docs/anveo-direct-setup-guide.md still ended its "native Messages app"
  section with "Codes arrive as ntfy push notifications instead". The point
  of that section — no SIP client can write into Android Messages or iOS
  Messages — is unchanged, but the place texts actually land is Sipnetic's
  message thread.
- services/pstn-trunk.sh's generated README claimed inbound SMS "doesn't
  touch Asterisk at all" and set up ntfy notifications. It is now the exact
  opposite: sms-inbound reads this service's pstn-personal-dids.conf and
  pstn-groups.conf to resolve DID ownership, the same files the inbound-voice
  ring logic uses, which also makes install order matter — noted there now.

Documentation only; no behaviour change, and services/sms-inbound.sh is not
touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NAddJGE1G6eGaPzmScG5Vh
2026-07-28 19:55:34 +00:00
Claude 1f6c7cf349 Default new extensions to TLS, with transport behind an Advanced disclosure
Making the transport a visible, equally-weighted choice meant new users had
to know the answer before they could get one right — and the wrong answer
fails in the least diagnosable way available: a UDP-only endpoint doesn't
refuse a TLS registration, it ignores it, so the phone times out and nothing
is logged anywhere.

The add form now creates Remote/FQDN (TLS 5061) extensions without asking.
TLS is listed first in the markup so it is the default before any script
runs, and the JS no longer switches it to LAN on a box with no domain — that
box is not better served by UDP, it just needs the phone told to trust a
self-signed certificate, which the disclosure now says at the point of
choosing.

Transport and auto-answer moved behind an "Advanced…" toggle, leaving name,
extension and category as the whole form. The toggle resets on cancel and
after a successful add so it doesn't stay open across uses.

Verified in the browser on two fixtures — with and without DOMAIN_NAME set —
that the panel starts hidden, the transport reads fqdn untouched in both
cases, and each shows the right hint when opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NAddJGE1G6eGaPzmScG5Vh
2026-07-28 19:43:21 +00:00
Claude 11f48ed970 Show connection details per extension, and let them be edited
Adding an extension from the dashboard gave back a password and nothing else
— no server, no TURN credentials, and no indication of which transport the
endpoint had actually been written with. A correctly-created extension and
one that could never register looked identical.

Each row gets an info button opening the full set: SIP server, username,
password, transport and port, and the TURN server/user/password. The password
is read back from pjsip.conf rather than regenerated, so re-pairing a handset
no longer means deleting and recreating the extension. The same panel can
switch an extension between LAN (UDP 5060) and Remote/FQDN (TLS 5061) and
reset its password in place, keeping its category, room membership and PSTN
permissions.

That transport choice is the likeliest cause of a phone that looks right and
never registers: an endpoint written transport=transport-udp will not answer
a TLS registration and the phone just times out. The add form now defaults to
Remote/FQDN whenever DOMAIN_NAME is set, rather than always LAN.

Reading TURN details needs Asterisk's .env, which is chmod 600 — the
installer now grants the service user read on that one file via ACL and lists
it in the unit's ReadOnlyPaths, since ProtectSystem=strict would otherwise
hide it. The UI says so plainly if the file still isn't readable.

Also fixes a real gap: ea_add_device is a third, independent writer of
endpoint blocks alongside the two vendor paths that services/asterisk.sh
patches, and it was not emitting message_context=sip-messaging — so an
extension created from this dashboard silently had no internal SIP messaging
while one created from the vendor admin did.

Verified against a fixture: transport round-trips TLS↔LAN without
accumulating duplicate transport/media_encryption/ice_support keys, password
reset rewrites only the target extension's auth section, delete still works
after edits, and the browser shows the populated panel with the add form
defaulting to fqdn when a domain is configured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NAddJGE1G6eGaPzmScG5Vh
2026-07-28 19:42:13 +00:00
Claude a74da52760 Confirm inbound SMS→AMI working live; rule out outbound over SIP
Inbound: a real text through the full path (Anveo webhook -> relay ->
AMI MessageSend -> Sipnetic) landed with a SIP 200 OK, confirmed live.
Removes the last "UNVERIFIED"/"believed correct" hedges from
sms-inbound.sh now that the AMI permission class and the
Destination-not-To fix are both proven, not just plausible.

Outbound over SIP: tested directly by sending a MESSAGE toward Anveo's
trunk (the mirror image of the inbound webhook). Anveo's SBC responded
501 Not Implemented -- a real, unambiguous rejection of the method
itself. Closes off this avenue for good, symmetric with inbound
SMS-over-SIP already being confirmed unavailable on this DID: neither
direction is offered on this account via SIP. Sending still has no
working path here until Anveo activates HTTP SMS-API access.

Also updates docs/anveo-direct-setup-guide.md's SMS section, which still
described the ntfy-based mechanism this session fully replaced with
AMI/Sipnetic delivery, and corrects its stale "not available on Direct"
sending claim with what's actually been confirmed: SIP MESSAGE outbound
is a dead end (501), the HTTP API is real but pending Anveo activation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-27 13:41:31 +00:00
Claude 0135b2ed27 Use Destination, not To, for MessageSend's endpoint resolution
`manager show command MessageSend` (this box's own Asterisk, requested
live) documents Destination as the field that actually resolves an
outgoing message's endpoint/technology; To is documented as a
backward-compatible fallback for the destination when Destination is
omitted, and separately as just the outgoing SIP MESSAGE's To: header
content when Destination IS provided. Two live attempts using only To
(bare "pjsip:212", then domain-qualified "pjsip:212@domain") both
produced zero SIP wire traffic -- confirmed via `pjsip set logger on`
during a real delivery attempt against an actively-registered contact --
meaning that documented fallback path isn't actually wired up on this
Asterisk version regardless of what the docs promise.

Switched to Destination using the docs' own "endpoint" form: bare
"pjsip:<ext>", no domain, which resolves via the endpoint's default
aor/contact -- the same live, registered contact `pjsip show contacts`
already confirmed exists for this extension.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-27 12:45:03 +00:00
Claude 5b339fdc0d Try qualifying MessageSend's To with a domain -- To/From asymmetry
Live test: AMI MessageSend to a bare "pjsip:212" produced zero SIP wire
traffic (confirmed via `pjsip set logger on` during a real delivery
attempt with the target extension actively registered) -- Asterisk
never even tried reaching the registered contact, meaning the failure
was in URI resolution before anything got sent, not a rejection from
the softphone. From already carried a domain (SMS_DOMAIN); To didn't.
Testing whether that asymmetry was the actual cause.

Explicitly a live experiment, not a confirmed fix -- next test will
show whether this produces real SIP MESSAGE traffic in the logger.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-27 04:01:52 +00:00
Claude e48460d6f2 Fix the actual cause of every "file does not exist": ProtectHome=true
Chased this as an ACL-ordering problem for the last several commits, and
those fixes were real and worth keeping, but none of them could ever
have fixed this: ProtectHome=true in the systemd unit doesn't just
restrict permissions, it mounts an empty, invisible filesystem over
/home, /root, and /run/user for the whole unit. ASTERISK_CONFIG_DIR
lives under /root/docker/... (or /home/<user>/docker/... on a non-root
install), so the relay process could never see it regardless of any ACL
grant on the real filesystem underneath -- from inside the sandboxed
unit it genuinely doesn't exist, while a plain unsandboxed shell
(confirmed live: `sudo -u smsrelay cat pstn-personal-dids.conf` outside
systemd) reads the exact same path fine.

Fix: ProtectHome=read-only instead of true. Still stops this service
from writing into /home or /root -- all it should ever need is read --
it just stops hiding them outright.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-27 03:43:38 +00:00
Claude 733a583c6a Grant smsrelay's ACL access after the Asterisk restart, not before
Reproduced live: pstn-personal-dids.conf was confirmed readable by
smsrelay right after a manual ACL grant, then unreadable again
("does not exist" in the relay's log -- os.path.isfile() swallows the
PermissionError and just returns False) immediately after the very next
fresh install. The only thing that ran in between was this same
install's own Asterisk container restart (needed to pick up the new AMI
secret).

The ACL grant was sequenced BEFORE that restart. CLAUDE.md documents the
container's entrypoint re-chowning its mounted config directory on every
restart and says chown alone can't touch ACL entries -- true, but
apparently this image's entrypoint also chmods, and chmod recomputes a
directory's ACL mask entry, which can silently weaken a named-user grant
made before it even though the grant's ACL entry itself is untouched.

Fix: do the grant last, after the restart-or-not branch, so nothing left
in this install run can undo it. ensure_docker_dir_ownership() (chown
only, confirmed in lib/common.sh, no chmod) stays where it was --
chow doesn't need this ordering fix, only the ACL grant does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-27 03:27:14 +00:00
Claude e59f16556a Log rejected sms-inbound webhook requests instead of dropping silently
A token mismatch returned a bare 404 with no journal line at all, so
"nothing is happening" was indistinguishable from "no request ever
arrived" -- exactly what a stale provider URL looks like after a
RELAY_TOKEN rotation (every full reinstall generates a new one, which
invalidates whatever's still pasted into the DID's SMS tab until it's
updated). Now logs the request's source IP and path length -- never the
attempted path itself, since that's unauthenticated input from whoever
hit the port, no reason to trust or echo it into the journal.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-27 03:04:37 +00:00
Claude 9207daa433 Stop the old sms-inbound instance before a fresh reinstall, not after
The port-taken message on the last run ("Port 8093 was taken — the relay
will use 8094") was this service colliding with itself, not a real
conflict. Fresh-install never stopped the previous run before scanning
for a free port, so it always found its own earlier process still bound
to 8093, silently moved to 8094, and then `systemctl enable --now` was a
no-op against an already-active unit -- meaning the OLD process (holding
the OLD AMI secret and OLD relay token, from before this session's
settings.env fix) kept serving traffic while the freshly-written config
and port sat unused underneath it.

Fix: `systemctl stop sms-inbound` right before the port scan. Now the
scan only reports a real conflict from something else, and the box
should settle back on 8093 (or whatever's actually free) with the
process that's really running matching what was just configured.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JDyKC6Kdg7tofmYSmRtgww
2026-07-27 03:00:30 +00:00