9045f8f030441cab153ad7badd5cd9fd66d9e7ff
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9045f8f030 |
Fix DEVICE_MARKER_RE field order so list_extensions() actually finds devices
DEVICE_MARKER_RE expected "; === Device: NAME [AA:marker] (category) ===",
but device_config's own template (further down this file) generates
"; === Device: NAME (category) [AA:marker] ===" -- category parens before
the AA tag, not after. The regex never matched a real device comment, so
list_extensions() silently returned [] for every device on every install,
and /api/pstn-permissions served {"extensions": []} regardless of what was
actually in pstn-permissions.conf. That's why the Extensions tab's
Messaging/Voicemail checkboxes always rendered unchecked after a save +
reload even though the file itself had messaging=yes/voicemail=yes written
correctly -- the JS falls back to an all-default row when the endpoint
returns nothing. ea_list_devices() and the rename-device code parse the
same comment via string-splitting/a differently-shaped regex and were
already correct; this was the one broken parser.
|
||
|
|
de5cfe26bd |
Fix app_voicemail module collision at its real source (entrypoint.sh, not host modules.conf)
The prior fix (
|
||
|
|
930233cc7f |
Fix app_voicemail module-load collision via modules.conf
Live-discovered bug, present on every install using this script, not
specific to any one box or extension: the easy-asterisk image ships
app_voicemail.so, app_voicemail_imap.so, and app_voicemail_odbc.so all
autoloading by default -- three alternative storage backends for the
SAME application (VoiceMail, VoiceMailMain, VMAuthenticate,
VoiceMailPlayMsg, VMSayName, the VM_INFO function, several AMI
actions), which collide registering those names against each other on
every single Asterisk start. This box's own container log showed the
exact signature on every restart: "Already have an application
'VoiceMail'" (and every sibling) followed by "app_voicemail.c:15897
load_module: Failure registering applications, functions or tests" --
app_voicemail never actually finished loading. Confirmed against
Asterisk's own documentation this session rather than assumed: this
is a known multi-backend conflict ("administrators should enable only
one module at a time"), not something specific to this repo's config.
Fix: new _asterisk_write_modules_conf, called from both the fresh-
install and update paths (matching voicemail-dialplan.conf's own
call-site pattern) alongside the other config/asterisk files, all
sharing the already-bind-mounted ./config/asterisk:/etc/asterisk
volume -- no new mount needed. noloads the two backends this repo
never configures (no IMAP/ODBC settings are ever written anywhere in
this script), leaving only the plain file-based app_voicemail.so
(the one voicemail.conf's [default] mailboxes actually target) to
load cleanly. Regenerated on every install/update, unlike
voicemail.conf, since modules.conf carries no per-install state of
its own -- consistent with how messaging-dialplan.conf/voicemail-
dialplan.conf are already handled, and added to their same chmod 644
line.
Requires a container restart to take effect on an existing install
(re-run `sudo ./setup.sh asterisk` -> Update, which regenerates this
file, then restart the container once).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
be684e114d |
asterisk-standalone-backup.sh: fix cross-layout restore (DO -> non-DO)
Live blocker: user's actual migration is DigitalOcean droplet (asterisk-digital-ocean/, container easy-asterisk-do) -> IONOS (plain asterisk/, container easy-asterisk) -- exactly the case this script didn't handle. The archive's own top-level directory name and docker-compose.yml reflect whichever layout produced it (_asterisk_resolve_layout's two known layouts). The previous restore extracted straight into $PARENT_DIR, which recreates whatever name is baked into the archive -- restoring a droplet archive onto a fresh non-droplet install would land the data at a *second*, wrongly-named directory (asterisk-digital-ocean) alongside the freshly-installed one it was meant to replace, with docker-compose.yml still naming the old project/container(s). Every service that resolves Asterisk's layout by directory/container name (security-dashboard.sh, pstn-trunk.sh, CrowdSec's Asterisk acquisition, Caddy) would get confused by having two candidate layouts on disk, one of them stale and half-wired. Fix: extract into a scratch staging directory first. If the archived docker-compose.yml's container_name differs from this run's own $CONTAINER (baked in at generation time, so always correct for whichever layout THIS box's install actually uses), rewrite the project name, container name, and coturn container name in place (coturn's is always "$CONTAINER-coturn" on both known layouts, so no lookup table needed) before the data ever lands at $HERE -- never lets the archive's own naming leak through. A same-layout restore (most common case, or two droplet boxes, or two plain boxes) detects no mismatch and skips the rewrite entirely, unchanged from before. Verified against the real generated script (extracted from the heredoc, not reimplemented): a droplet-flavored archive restored onto a fresh plain-layout box lands at the correct single directory with no stray second directory, and docker-compose.yml's name/container_name/ coturn container_name all correctly rewritten to the plain layout (confirmed by diffing the actual restored file, not just checking for absence of errors); a same-layout restore (droplet archive onto a droplet box) confirmed to skip the rewrite entirely; the pre-existing external-IP patch (previous commit) still fires correctly stacked on top of the layout fix; and the extraction-failure rollback path (a corrupt/unreadable archive) still restores the pre-restore install untouched, verified via a marker file surviving the rollback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
6dc4f0d6c6 |
asterisk-standalone-backup.sh: fix external IP on cross-box restore
User confirmed via a live pjsip.conf on their actual DO box: both external_media_address and external_signaling_address are literal IPs, written by easy-asterisk at first container start (not something this repo's install script controls directly). A straight restore of a backup archive from a DIFFERENT box onto the new IONOS box would leave the OLD DigitalOcean IP baked into pjsip.conf — dialplan and PJSIP device credentials would come back fine, but RTP media (and likely SIP signaling/registration) would stay broken, silently, since nothing in the restore path previously touched these values. Fix: after extracting the archive, `restore` reads the archive's own external_signaling_address as "old IP", detects this host's actual current public IP (same DO-metadata -> ifconfig.me -> hostname -I fallback chain services/asterisk.sh's own install already uses), and if they differ, rewrites every occurrence across config/ and .env (fixed-string match, not a regex, so the IP's dots can't be misinterpreted). Deliberately does NOT touch spool/, logs/, or lib/ — those hold voicemail messages and call recordings, and a blind text substitution across binary audio would corrupt it. A restore onto the same host (e.g. rolling back a bad config change, no IP change) leaves every file untouched — the check only fires on an actual mismatch. Verified against the real generated script (extracted verbatim from the heredoc, not a reimplementation) with a full mock backup/restore cycle: built a fixture archive with pjsip.conf's three transport blocks (udp/tcp/tls) and .env's TURN_SERVER all hardcoded to a fake "old box" IP, plus a fake binary voicemail file; restored it onto a mocked "new box" with a different detected IP via a stubbed curl. Confirmed every occurrence in both pjsip.conf and .env was correctly rewritten to the new IP, and confirmed via byte-for-byte comparison that the binary voicemail file was completely untouched. Separately verified the same-IP case (mocked curl returning the archive's own IP) makes no changes at all, matching a same-host config rollback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
60238814f4 |
dr_bringup: support restoring from the offsite mirror on a brand-new box
This is what's actually needed for a DigitalOcean -> IONOS Asterisk migration (item #1 of the user's original 4-item list): move the whole stack to different hardware entirely, not restore onto the box the offsite mirror already targets. dr_bringup_kopia.sh only ever scanned DEST_NAMES for restorable snapshots — those are always local filesystem Kopia repos (services/backup.sh creates them with `repository create filesystem --path=...`), meaning they only exist on whichever box originally ran the backup. On a genuinely new box, every one of them fails to connect and there's nothing left to restore from — the script's own header comment only covered the case where "the spare box IS the box the primary's mirror targets" (i.e. already holds a copy of the repo data), not a fresh, unrelated box. Fix: also try REMOTE_TYPE/REMOTE_ARGS (the offsite Backblaze/S3 mirror, if configured) as a same-shaped destination named "offsite", reusing DEST_default_PASSWORD since sync-to always mirrors that exact same encrypted repo. Connects once into a fresh local config file scoped to this DR run, then folds into the existing per-destination scan/restore loop unchanged — "offsite" just becomes another entry in _DEST_ARR. Documented the actual migration workflow in the header comment, including the BACKUP_CONF override so copying the old box's backup.conf over doesn't clobber the new box's own freshly-configured one. Verified against the real script (not a reimplementation) with mocked kopia/docker binaries and a crafted backup.conf, covering: local dest unreachable + offsite connects successfully (the actual migration shape) with correct service/path discovery; a real (non---list) restore run confirmed it selects the latest of multiple snapshots by startTime and issues the correct `kopia restore <snapshot> <path>` call; offsite connect failing (bad REMOTE_ARGS) warns and degrades to "no restorable sources found" instead of crashing; REMOTE_TYPE=none skips the new code path entirely with no behavior change (regression check against the pre-existing local-only case). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
4717c3a080 |
Add garage-webui: browse Garage's buckets/objects like Backblaze's console
User's actual question: Backblaze B2's web console lets them browse a bucket as folders/files; Garage has no equivalent by default, so after switching an additional backup mirror from Backblaze-only to also target local Garage, they had no way to visually confirm data landed there the way they could on Backblaze. "S3 storage is opaque, you can't browse it" was true of Garage's *own* CLI, but wrong as a blanket statement — Backblaze's browsability comes from a client (its web console) layered on top of the same kind of object storage, and Garage has an actively-maintained equivalent (khairul169/garage-webui, 1.1k stars, "integrated objects/bucket browser") that gives the same experience against Garage's S3 API. services/garage-webui.sh (new): standard service-template Docker service. Requires an existing services/garage.sh install (checks for $DOCKER_DIR/garage/.env, errors with instructions if missing — this is a browser for an existing instance, not a replacement). Reaches Garage over host.docker.internal (both containers' ports are already published to the host — simpler and more robust than trying to join garage's own Compose-project-scoped default network by name). Has its own login (AUTH_USER_PASS, bcrypt via a throwaway `docker run --rm httpd:alpine htpasswd` — same $ -> $$ escaping services/wg-easy.sh already uses for its own bcrypt PASSWORD_HASH, verified here against a real docker compose config run: unescaped, Compose tries to interpolate $2y$05... as variable references and silently corrupts the value with a "not set" warning; escaped, it passes through intact with no warning), so it doesn't need Authelia gating by default. Prerequisite fix in services/garage.sh: its admin API (bucket/key management, object listing — the thing garage-webui talks to) has been running with zero authentication since this service was first built, because admin_token was never set in garage.toml. Nothing in this repo called that API before now, so it went unnoticed; adding a real consumer is what surfaced it. Fixed: generate admin_token (openssl rand -base64 32) alongside the existing rpc_secret, persist GARAGE_ADMIN_TOKEN/GARAGE_ADMIN_PORT to .env for garage-webui to read locally (never sent over SSH, unlike the S3 credentials backup.sh reads remotely). Update mode backfills admin_token into an existing garage.toml (+ restarts just the garage container to apply it) for anyone who installed before this change, same backfill-not-break approach as the GARAGE_S3_API_PORT fix from the previous commit. Verified: bash -n on both files; docker compose config against real Docker Compose for both the primary garage.toml/.env generation (with the new admin_token/GARAGE_ADMIN_PORT fields) and the new garage-webui docker-compose.yml; the bcrypt-escaping behavior specifically (proved via a minimal repro that unescaped $ corrupts the value with a warning, escaped does not); the admin_token/ GARAGE_ADMIN_PORT Update-mode backfill logic against old- and new-style .env/garage.toml fixtures, including idempotency (running it twice adds nothing a second time); and the credential-parsing regexes in garage-webui.sh against both a complete .env fixture and an old one missing the new fields (confirms the "run garage's Update first" error path actually triggers rather than proceeding with blanks). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
558fc3e75b |
Fix layout-apply crash on Garage reinstall, backfill missing .env fields
Live failure: `sudo ./setup.sh garage` → Full reinstall on a box that
already had a working Garage install crashed with:
Error: ApplyClusterLayout returned InternalError (500): Internal
error: Invalid new layout version
Root cause: a "Full reinstall" deliberately never wipes ./data or
./meta (that's real backup-mirror data — Kopia's sync-to s3 target —
and losing it silently on reinstall would be far worse than the
alternative), but the cluster-init step unconditionally re-ran `garage
layout assign` + `layout apply --version 1` every time it was reached.
Garage requires each apply to be exactly previous_version + 1; a node
that already has a committed layout (from the earlier install, still
sitting in the preserved ./meta) rejects a second "1". Fix: check
`garage status` for "NO ROLE ASSIGNED" first and only run the
assign/apply once, matching what the surrounding comment already
claimed happened ("Only ever run once") but the code didn't enforce.
Second, related issue this would have hit immediately after: the same
reused-./meta state almost always means an existing bucket + key from
the earlier install are still sitting in Garage's storage. The fresh
flow was about to silently create a brand-new bucket/key and overwrite
.env to point at those instead — orphaning any real data already in
the old bucket (nothing left on disk pointing at it, even though it's
still physically stored). Now: when the layout is already applied,
list existing buckets and require an explicit y/n (default n) before
creating new ones, with recovery instructions for reconnecting to an
existing bucket by hand instead.
Third, the actual reason a full reinstall was reached at all: Update
mode never backfills .env fields added to this script after someone's
initial install (GARAGE_S3_API_PORT, needed by services/backup.sh to
read an instance remotely) since Update deliberately never touches
.env otherwise — the only other path was the now-unsafe fresh
reinstall. Update now backfills just that missing key by reading the
real port back out of the already-written docker-compose.yml, so a
future .env schema addition doesn't force this tradeoff again.
Verified with standalone harnesses (not the live install, mocked
`garage status`/bucket-list output and .env/docker-compose.yml
fixtures): all four layout-state branches (fresh node, existing
buckets + decline, existing buckets + confirm, existing role but no
buckets), and both backfill cases (missing key added, existing key
left alone). Caught and fixed a real bug in the first draft of the
port-extraction regex during this testing — grep -oE '^[0-9]+' never
matched because the captured group still had its surrounding quotes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
f61a2717bf |
Offer retry/fall-back-to-SFTP/skip when Garage isn't found for a mirror
Previously, if the additional-mirror S3/Garage check couldn't find
~/docker/garage/.env on the remote box, it just warned and silently
dropped the mirror — forcing a full re-run (and re-entering every
already-answered prompt: destinations, passwords, schedule, B2,
DR-spare, etc.) once Garage was actually installed.
Wrap the S3/SFTP branch in a loop so the "Garage isn't installed yet"
case now offers a real 3-way choice:
1) install Garage in another session, then retry the same .env check
without leaving this script
2) fall back to SFTP for this one mirror, reusing the already-resolved
destination host/port/user/mirror-name with no re-prompting
3) skip just this mirror (default — safe for UNATTENDED, which
resolves to this automatically since prompt_text returns its
default without blocking)
Everything else install_backup() has already collected lives outside
this loop, so none of it is at risk regardless of which of the three
exits it via.
Verified against a standalone harness reproducing the state machine
with a mocked ssh (empty .env vs. populated .env after a simulated
install) and prompt_text, covering all three interactive choices, the
blank/Enter default, and UNATTENDED mode (confirms the blocking
"press Enter to retry" read is unreachable there since prompt_text
resolves choice 1's prompt to default "3" without waiting on stdin).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
74cab14f86 |
Let the additional-mirror setup read Garage credentials over SSH
Extends the "ADDITIONAL MIRROR" section (previously SFTP-only) with a
type choice: SFTP, or S3 against a Garage instance already running on
that box. For the S3 path, this script never asks the operator to retype
a bucket name or key — it SSHes to the destination, reads
~/docker/garage/.env directly (the real, currently-configured values,
generated once by services/garage.sh and never touched again on its own
Update runs), and uses those for the dry-run verification and the
persisted mirror args. If Garage isn't installed there yet, it says so
plainly with the exact install command instead of failing cryptically or
silently skipping.
Also removed the last hardcoded suggestions from services/garage.sh
itself ("kopia-backup" / "kopia" as fixed prompt defaults) — replaced
with a freshly-generated suggestion each run (timestamp-suffixed), so
nothing about the bucket/key name is a fixed string baked into this repo
at any point in the chain; it's always the operator's actual choice, read
back live wherever it's needed.
Verified end-to-end against a mocked ssh (returning realistic
~/docker/garage/.env content) covering both outcomes: Garage installed
with a real bucket/key correctly parsed, dry-run run, and persisted; and
Garage missing, correctly warning with the install command and leaving
backup.conf untouched either way.
|
||
|
|
01ca0a7b09 |
Fix leading-whitespace bug in garage key create output parsing
Garage's real CLI output pads labels with extra spaces for column
alignment ("Key ID: GKxxxx"), not a single space like the
mocked test used ("Key ID: GKxxxx") — the fixed ": " field separator left
that padding stuck to the parsed value, so .env ended up with access
key/secret strings carrying leading whitespace inside the quotes.
Confirmed live by the user right after install. This would have broken S3
auth outright once actually used, since access keys have to match exactly.
Switched to ':[[:space:]]+' as a regex field separator, which consumes
however many spaces are actually there instead of assuming exactly one.
Verified against both the single-space and padded/aligned formats — both
now produce the identical clean value with no leading whitespace.
|
||
|
|
129a97f34b |
Add Garage — MinIO CE's replacement — as a self-hosted S3 object store
MinIO's open-source community edition is dead: console GUI stripped May 2025, Docker images stopped publishing October 2025, repo formally archived April 2026, with MinIO redirecting everyone to their paid AIStor product. Verified this directly before building anything, since recommending a since-abandoned image would have been worse than the SFTP problem this was meant to solve. Garage (Deuxfleurs) is the actively-maintained small-scale self-hosted replacement — single Rust binary, purpose-built for exactly this "one lightweight node" use case (as opposed to SeaweedFS, which targets large object counts / large-scale deployments, more machinery than a single backup-mirror target needs). services/garage.sh follows this repo's standard service template: port scanning for the S3 API/RPC/admin ports, an RPC secret generated once and never touched again on Update, and a one-time cluster init sequence (layout assign/apply, bucket create, key create, bucket allow) gated on whether .env already has a saved access key — Update reruns skip all of it and just refresh the image. Primary intended use: a local S3-compatible target for services/backup.sh's additional-mirror Kopia sync, so a local mirror can reuse the exact same sync-to s3 code path already proven reliable for the Backblaze B2 mirror, instead of Kopia's separate, less-exercised SFTP backend that's been the source of today's connection troubleshooting. Verified end-to-end against a mocked environment (fake docker exec returning realistic `garage status`/`garage key create` output) — caught and fixed a real off-by-one in the status-output parsing this way (grabbed the column-header row's literal "ID" instead of the actual node ID; output has a title line, then a header line, then the data row). Also validated the generated docker-compose.yml with real `docker compose config` in both the no-network and network-created cases. |
||
|
|
b25724a61a |
Suggest a /kopia-data subdir of the DR-spare path for the SFTP mirror
The additional-mirror "Remote path for the repo" prompt always suggested a generic ~/backups/kopia-mirror default, unrelated to wherever the operator already pointed the DR-spare sync. Requested directly: default to that same location instead, in its own /kopia-data subdirectory so Kopia's repository files don't end up visually mixed in with the two plain config files (backup.conf, README.md) the DR-spare sync writes straight into DR_SYNC_PATH itself. Falls back to the original generic default when DR_SYNC_PATH isn't set (no DR-spare configured yet). Verified the path computation handles a DR_SYNC_PATH with or without a trailing slash correctly (no double slash), and the unset case still falls back as before. |
||
|
|
1bb7498567 |
Pass the resolved SSH port to Kopia's SFTP mirror, not just user/host
The additional-mirror setup already resolves user/hostname through ssh -G so a ~/.ssh/config alias works, but never extracted port — Kopia's sftp storage backend doesn't read ~/.ssh/config at all and defaults to 22 regardless of what the alias actually configures. Confirmed live: this produced "server unexpectedly closed connection: unexpected EOF" on the dry-run verification — Kopia connecting to the right host on the wrong port, not a credentials or host-key issue, which is exactly why plain `ssh main` kept working the entire time this was being debugged (it reads the alias's Port line correctly). Now parses `port` out of the same ssh -G output, defaults to 22 if absent (matching ssh's own default), and passes --port= through to both the dry-run check and the persisted EXTRA_MIRROR_ARGS string — the latter matters as much as the former, since that's what every actual scheduled sync reuses afterward, not just the one-time verification. Verified the parsing against three cases: a custom-port alias, a default-port alias, and an unresolvable alias — all three resolve to the correct port with no manual intervention needed. |
||
|
|
2b9ba85875 |
Merge remote-tracking branch 'origin/main' into claude/ionos-script-integration-x32ofw
# Conflicts: # README.md # lib/common.sh |
||
|
|
29db2da5fc |
Add fix_pikapods_dump.py to the repo; cross-reference it from the migration script
extras/fix_pikapods_dump.py patches two confirmed Adminer PostgreSQL-export bugs that otherwise make a PikaPods Mattermost migration fail outright: unquoted enum-label DEFAULT values (Postgres reads the bare label as a column reference and rejects the CREATE TABLE) and boolean columns serialized as bare 0/1 instead of true/false (Postgres doesn't implicitly cast integers to boolean). Boolean columns are discovered by actually parsing each CREATE TABLE in the dump rather than working from a hand-curated list — Postgres only reports the first bad column per failed row, so a list built from error output alone would likely be incomplete. Already verified earlier this session against a real local Postgres 16 instance; reviewed now for anything needing redaction before committing — it's a generic text-processing tool with no hostnames, credentials, file paths, or personal data in it, so nothing needed changing. Cross-referenced from the generated migrate-from-pikapods.sh's header comment (services/mattermost.sh) so anyone hitting a CREATE TYPE/CREATE TABLE or boolean-column import error is pointed at the fix instead of having to rediscover it. |
||
|
|
a7d3dc5b5d |
Fix imported file ownership in the generated PikaPods migration script
The generated migrate-from-pikapods.sh (services/mattermost.sh's existing "Migrating from an existing Mattermost instance?" prompt on fresh installs) already correctly parameterizes PROJECT_DIR/MM_CONTAINER/DB_CONTAINER per instance — no bug there. What it missed: after rsync/cp-ing files in from the export, it never touched ownership, so the imported ./data landed owned by whoever ran the script instead of the fixed UID 2000 mattermost/mattermost-team-edition runs as. Every file write then failed with permission denied — confirmed live as the actual cause of a client-side "stream closed" error on image/file uploads after a real migration. Adds chown -R 2000:2000 ./data right after the copy step, and a root check up front since chowning to an arbitrary UID needs it (docker/psql access already implied running as root in practice, just never enforced explicitly). Usage lines updated to say `sudo` to match. Verified by reconstructing the exact generated script from the real source heredocs (head + variable substitution + body, the same three pieces the actual cat/cat>> sequence produces) and syntax-checking the result — root check and chown both land in the right place, and PROJECT_DIR/MM_CONTAINER/DB_CONTAINER still resolve correctly per instance. |
||
|
|
767479d113 |
Self-heal Mattermost data/logs/config/plugins ownership on every start
Root-caused a live "stream closed" image-upload failure to data/20260814/.../mkdir: permission denied — the volumes weren't owned by the fixed UID 2000 mattermost/mattermost-team-edition runs as, most likely left that way by the PikaPods data import. The install script already chown -R 2000:2000's these on every run (fresh or update), so re-running the installer would have fixed it — but that still means remembering to re-run it every time ownership drifts for any reason, including causes this repo doesn't control (a future migration, a manual restore, anything that copies files in as a different UID). Added a small mattermost-fix-perms init container (busybox, chown, exit) that the mattermost service now depends on via condition: service_completed_successfully. Runs on every `docker compose up` — including a plain host reboot, since restart: unless-stopped brings the stack back on its own — so this self-heals permanently instead of needing a human to notice and fix it by hand again. Verified the generated compose file (with representative variable values) against real `docker compose config`: valid YAML, and the dependency graph correctly shows mattermost waiting on both db (service_healthy) and mattermost-fix-perms (service_completed_successfully). |
||
|
|
10b985dc95 |
Add standalone Pi-hole service; move wg-easy's default port off Netbird's
Two independent, requested changes: - services/pihole.sh: new standalone service, Pi-hole v6 (the image moved entirely to a TOML-based /etc/pihole config — the old WEBPASSWORD env var and separate /etc/dnsmasq.d volume are both gone; uses FTLCONF_webserver_api_password and FTLCONF_dns_listeningMode=ALL instead). Deliberately not wired into wg-easy or any other VPN — a device has to be pointed at it manually (per-device or via router DHCP). DNS itself (53/tcp+udp) is never scanned/moved since shifting it off the standard port would defeat the point; a port_in_use check warns instead of blocking, since the common case (systemd-resolved on 127.0.0.53 only) doesn't actually collide with Pi-hole binding the host's real interfaces. Web admin UI is Caddy-fronted like everything else in this repo. Added to the README services table. - services/wg-easy.sh: default VPN/web ports moved from 51820/51821 to 51830/51831. Netbird's own WireGuard listener also defaults to exactly 51820 — installing both on one box means wg-easy's existing scan-and-move logic would silently shift its port every time, which is harder to predict/document than just not starting on the collision in the first place. The scan itself is unchanged and still moves both ports further if even the new default is taken. Tested pihole.sh's full standalone install flow (no-Caddy and Caddy-present-locally cases) against a mocked environment, validating both generated docker-compose.yml files with `docker compose config`, and confirmed the reinstall-mode gate correctly no-ops on a second run in unattended mode. |
||
|
|
0f9ce0a82b |
Add a flag to permanently retire shared coturn without auto-reinstalling it
ensure_coturn_user() auto-installs shared coturn (services/coturn.sh) any time $DOCKER_DIR/coturn doesn't exist — correct behavior for "first service that needs TURN", wrong behavior for "an operator deliberately decided every consumer should run its own dedicated coturn instead and removed the shared one on purpose". The function had no way to tell those two states apart, so deleting ~/docker/coturn didn't actually retire it — the next service to call this function (a Mattermost reinstall, a fresh Asterisk install) would silently bring it right back. touch ~/docker/.coturn-retired now short-circuits the function straight to the existing "no TURN available, caller degrades gracefully" return path, before it ever looks at install_coturn. Every existing consumer (Asterisk, Mattermost) already handles that path correctly today — it's exactly what happens if shared coturn simply fails to install — so this needed no changes on the consumer side, only closing the gap in the shared function. Verified with a mock: with the flag present, install_coturn is never invoked and the function returns empty COTURN_HOST/rc=1 as expected. |
||
|
|
d2568ecb09 |
Read back existing destinations, ntfy, schedule, and B2 fields on rerun
Requested after a rerun silently reset DR_SYNC_PATH (fixed separately) — auditing the rest of install_backup() turned up the same class of bug in several other places, one of them worse than the one that prompted this: - Default destination repo path defaulted to $ACTUAL_HOME/backups/... even when the real configured repo was somewhere else entirely (this user's actual path is /root/backups/kopia-backup) — accepting the shown default on a rerun would have pointed the installer at the wrong location. - Extra (non-"default") destinations weren't preserved AT ALL on a rerun — skipping "Add more destinations?" silently dropped every extra destination, and anything mapped to it, from the rewritten backup.conf. - The per-service destination-assignment prompt always showed "[default]" regardless of the service's actual existing mapping. - ntfy URL/token always started blank, silently disabling notifications on any rerun where they weren't retyped. - The schedule prompt always defaulted to option 1 (daily 02:00) instead of reading back whatever OnCalendar was actually already running. - B2's four sub-fields (bucket/endpoint/key ID/secret) always started blank even when reconfiguring an already-working REMOTE_TYPE=s3 setup — a mispaste on any one of the four meant retyping all four blind, since there was nothing to fall back to per-field (the existing REMOTE_ARGS was already preserved as a whole on a blank/failed attempt, just not offered back as individual editable defaults). All six read the same way: pull the existing value from backup.conf (or, for the schedule, from the live systemd timer unit — schedule isn't stored in backup.conf) and use it as the prompt default, so accepting the default keeps what's already there instead of silently reverting it. Verified all six against a mock backup.conf + timer fixture with pre-existing values for every field this touches. Known remaining gap: KEEP_LATEST (retention count) still isn't read back — doing so correctly needs the repo already connected, which happens later in this same function's flow. Flagging rather than rushing a reorder here. |
||
|
|
b9152369ef |
Fix DR-spare path reset on reinstall and tilde-quoting in remote commands
Two stacked bugs, found together when re-running the backup installer to
add an SFTP mirror silently reverted a previously-set absolute
DR_SYNC_PATH back to the script's tilde-based default, which then failed
outright:
1. services/backup.sh never read DR_SYNC_HOST/DR_SYNC_PATH back from an
existing backup.conf before prompting (every other setting in this file
does — passwords, mirrors). Accepting the prompt defaults on a rerun
silently reset both to blank/"~/docker/backup" instead of keeping what
was already configured. Fixed by reading them back the same way
DEST_*_PASSWORD already does.
2. extras/backup_kopia.sh's DR-spare sync wraps the remote path in single
quotes for its `ssh host "mkdir -p '...'"` / `"chmod 600 '.../...'"`
commands. Single-quoting a leading ~ stops the remote shell from
expanding it at all, so it looked for a literal directory named "~"
instead of the home directory — breaking the script's own DEFAULT
DR_SYNC_PATH ("~/docker/backup") for anyone who actually used it.
rsync's own transfer step has separate, correct tilde handling, which is
why the sync itself "succeeded" while the follow-up chmod couldn't find
the file. Fixed with a small _dr_remote_quote() helper that keeps a
leading ~/ outside the quotes while still safely quoting the rest of
the path.
Verified the quoting fix by parsing the exact constructed command string
in bash directly — a plain '~/docker/backup' stays literal (the bug),
~/'docker/backup' correctly expands to $HOME/docker/backup (the fix).
|
||
|
|
9edd821349 |
Don't offer to generate a root SSH key when one already works for the host
_backup_ensure_root_ssh_key() only ever checked for /root/.ssh/id_ed25519 or id_rsa by exact filename. Root can already SSH to the DR-spare/mirror host just fine in practice (proven by this same script's own DR-spare sync succeeding), just via a key with some other name — so the function had no way to see that and always fell through to offering a copy-from-user-home or brand-new ssh-keygen, both unnecessary. Now takes the target host as an optional argument. When given, it tests root's SSH access to that host as-is first and resolves the actual key via `ssh -G <host>` (which expands ~/.ssh/config the same way the SFTP-dest resolution earlier in this file already does) before falling back to the copy/generate prompts. Both call sites (DR-spare, SFTP mirror) now pass their respective host. Verified against a mock ssh: an already-working non-default-named key gets detected and reused with no prompts, and the original copy/generate fallback still triggers correctly when SSH genuinely doesn't work yet. |
||
|
|
f6e5bb4ea3 |
Use rsync instead of scp for the DR-spare backup.conf/README sync
The freshly-added raw-error logging paid off immediately: the box's spare sync was failing every run with "scp: Connection closed" while plain ssh exec to the same host worked fine. That split (ssh exec OK, scp specifically rejected) matches modern OpenSSH's default scp-over-SFTP transfer hitting a restriction on the remote side that a plain exec or rsync's own protocol don't trigger. Swapped the scp step for rsync -a over the same ssh options, keeping the ssh mkdir -p before it (rsync doesn't create missing destination directories) and the ssh chmod after. Verified the exact command/quoting against mocked ssh/rsync binaries — array expansion and remote path handling both check out. |
||
|
|
63774e0500 |
Read $_ERR once per failure instead of twice, fixing lost raw-error text
Last night's fully-failed backup run (0/20, "repository not found" on every service) showed the real gap: categorize_error() clearly saw real content in $_ERR (it matched a specific pattern, not the generic fallback), but log_raw_error()'s separate re-read of the same file moments later came back empty on every single failure — so the raw-error logging added earlier this session produced nothing when it mattered most. Fixed by reading $_ERR into a variable exactly once per failure and passing that string to both categorize_error() and log_raw_error(), instead of two independent file reads. Verified against a mock harness reproducing the same call pattern (three simulated failures in a loop, single shared error file) — both the categorized reason and the raw stderr text now come through on every iteration. Doesn't explain why last night's repo access failed in the first place (disk and mount checks came back clean) — but the next time it happens, this will actually surface the real kopia error instead of losing it. |
||
|
|
7efd087993 |
Add standalone backup/restore script for Asterisk, independent of Kopia
Asterisk's whole state (dialplan, pjsip devices, voicemail, recordings, .env with its coturn credential, docker-compose.yml) already lives under one self-contained directory, so asterisk-standalone-backup.sh just tars it — with stop/restart safety around the tar since voicemail/spool write continuously, and a move-aside-then-extract restore that rolls back automatically if extraction fails. Written into the install directory at both fresh-install and update time via _asterisk_write_standalone_backup_script(). Output defaults to ~/asterisk-backups/, deliberately outside ~/docker/, so a Kopia backup of the box doesn't also back up a backup-of-itself. Meant for a quick pre-change snapshot or moving this PBX to a new host without standing up the full backup stack first. Documented in the generated README's new "Standalone backup/restore" section. Tested against a mocked EA_DIR (fake docker/docker compose, config/spool/voicemail files) confirming backup produces a correct tar and restore replaces content correctly with rollback on extraction failure. |
||
|
|
6339235781 |
Correct B2 application key guidance: use "All" bucket access, not one bucket
Confirmed live and cross-checked against a real, documented Kopia issue (kopia/kopia#5329): the walkthrough previously told the operator to scope the Application Key to just the bucket they created — the more security-conservative default, and correct for B2's own S3-compatible API in general. But Kopia specifically needs the listBuckets capability even though it only ever touches the one configured bucket, and B2's basic "Add a New Application Key" web form doesn't expose a way to grant listBuckets on a bucket-restricted key — only an account-wide ("All") key gets it through that form. Without it, the connection fails with B2's unhelpful "Cannot access bucket" error, which doesn't point at the actual missing capability at all. Updated the guidance to "All" with the reasoning inline, and a note that single-bucket scoping is still possible for anyone willing to create the key via B2's CLI/API directly (b2_create_key with an explicit capabilities list including listBuckets) rather than the basic web form this walkthrough is written for. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
9c5d8c32f7 |
Reuse the sudo user's SSH key for root; make B2 rejection unambiguous
Two separate fixes from a live report. 1. The DR-spare and SFTP-mirror sections both checked ONLY /root/.ssh for a key, missing the common case: the person running `sudo ./setup.sh backup` already has a key under their own home directory (used interactively, quite possibly already authorized on the target box), while root — who actually runs the scheduled systemd service — has none. Confirmed live: "the computer has the ssh key for the sudo user on the box" produced "No SSH key found for root" with no inline way to do anything about it beyond a pointer to go set one up elsewhere and re-run. Factored both call sites into one shared _backup_ensure_root_ssh_key() that checks root first, then offers to reuse the sudo user's existing keypair (copied into /root/.ssh with correct ownership/permissions, root:root 600) before falling back to generating a brand new one — reusing an existing key can work immediately if it's already authorized on the target, where a fresh key needs a new ssh-copy-id round-trip regardless. Verified all three branches (root already has a key, root has none but the user does and accepts reuse, neither exists and one gets generated) against a mocked filesystem. 2. The B2 dry-run failure message read like it could be about missing input even when every field was non-empty — confirmed there's no code path where non-blank-but-wrong values actually trigger the separate "Left blank" message (the two are on disjoint branches), but the dry-run failure text itself didn't rule that out or point at the actual likely cause. Now echoes back what was entered (bucket, endpoint, Key ID — never the secret) so it's easy to eyeball against B2's own confirmation screen, states plainly that this is a rejection of non-blank input, and names the most likely cause directly: pairing the Key ID from one Application Key with the Secret from a different one, which is easy to do after creating more than one while troubleshooting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
fd0c3f7050 |
Log the raw error text, not just categorize_error()'s bucket label
Confirmed live: a real failure ("WARNING: spare sync failed — error —
see system logs on ubuntu") didn't match any of categorize_error()'s
known patterns, fell into its generic catch-all bucket, and the actual
stderr text that would have explained it was sitting in a mktemp'd file
this script deletes on exit (trap ... EXIT) — so there was nothing in
"system logs" to actually go check. The categorization was silently
discarding the one piece of information that would have diagnosed the
problem.
Added log_raw_error(), called right after every categorize_error() site
(5 of them: two snapshot-failure paths, the primary REMOTE_TYPE mirror,
the new EXTRA_MIRROR_NAMES loop, and the DR-spare sync) — logs the raw
stderr text (truncated to 500 chars) into the same log stream as
everything else, so it survives past the run that produced it instead
of being deleted with the temp file. categorize_error()'s short bucket
label is untouched and still used for FAILED_SVCS/notification text,
which should stay concise — this adds the detail alongside it, not
instead of it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
350708ae10 |
Resolve ~/.ssh/config aliases before handing a host to Kopia's sync-to sftp
Confirmed: Kopia's sync-to sftp has its own SFTP client and doesn't read ~/.ssh/config the way the system ssh/scp binaries do — so an alias set up via wg-easy's sync-ssh-aliases.sh (or any ~/.ssh/config Host entry) worked fine for the DR-spare connectivity check (which shells out to real ssh) but silently failed for this mirror: a plain @-split on an alias like "main" (no @ present) produced --host=main, a name that only resolves inside ~/.ssh/config, not real DNS. The dry-run check correctly rejected it and the mirror was never saved — no error surfaced beyond that, so it looked like nothing happened. Now resolves the destination through `ssh -G` before building the Kopia flags — the same mechanism ssh itself uses to expand config aliases — and falls back to the previous plain @-split only if that comes back empty. Verified against three cases: a bare alias (resolves via a mock ~/.ssh/config Host block), an explicit user@ip (passes through unchanged), and an unrecognized name (falls back to a sane literal hostname rather than erroring). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
b99ca32438 |
Support multiple simultaneous offsite mirrors, not just one
Requested: mirror to Backblaze B2 AND directly to the IONOS spare box over Tailscale, at the same time, not one or the other. REMOTE_TYPE/ REMOTE_ARGS was hardcoded to a single mirror target — extending it to a list would have meant redesigning the one thing that already works and was already verified against real B2 credentials, so this adds a separate, additive mechanism instead: EXTRA_MIRROR_NAMES, a space- separated list, with per-entry MIRROR_<name>_TYPE/_ARGS (same argument shape as REMOTE_ARGS). An existing B2-only backup.conf keeps working completely unchanged if this new section is skipped. install_backup() gets a new "ADDITIONAL MIRROR" prompt after the existing B2 section: offers a direct SFTP mirror (Kopia's sync-to sftp, not the deprecated b2 provider — same reasoning as the S3/B2 choice already made), defaults the destination to whatever was typed at the DR-spare prompt above (same box, same purpose, no reason to ask twice), checks passwordless SSH and an SSH key exist first, then verifies with a --dry-run against the just-created 'default' repo before saving it — same "don't save something broken" discipline as the B2 flow. Verified against a mock backup.conf that install-side writes and worker-side reads agree on the exact format, and that reusing an existing mirror name reconfigures it instead of duplicating it in the name list. One correction while researching sync-to sftp's flags: unlike plain ssh, Kopia doesn't shell out to the system SSH client, so it needs an explicit --keyfile and --known-hosts path rather than picking up whatever `ssh` already trusts automatically — checked Kopia's own docs for the exact flags before writing this, same as the earlier S3 case. extras/backup_kopia.sh's worker loops through EXTRA_MIRROR_NAMES after the existing REMOTE_TYPE mirror step, running sync-to for each destination against each additional mirror and folding failures into the same FAILED_SVCS/notification reporting the primary mirror already uses. Verified end-to-end against a mock backup.conf and a stubbed kp_for: both the B2 and the new SFTP mirror get called in sequence with the correct arguments. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
f36cb2c394 |
Show character counts and which field was blank in B2 setup prompts
Requested after a live failure: pasting into the hidden Application Key field silently captured nothing (terminal/SSH-client dependent), and the only symptom was a generic "one or more fields left blank" warning after all four prompts had already gone by — no way to tell which field, or even that the paste itself was the problem rather than something else. Each of the four fields now echoes its character count right after entry (never the value for the hidden Application Key field, just its length), so a failed paste is visible immediately instead of discovered several prompts later. The blank-field warning now also names exactly which field(s) were empty instead of a generic message. Verified against the user's actual reported case: bucket/endpoint/key-ID entered normally, Application Key came back empty — reproduces as "(0 characters entered)" on that line and "Left blank: Application Key" in the warning, both confirmed against a second case where all four fields are present and it passes through cleanly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
ba9c31aeb1 |
Detect Netbird/Tailscale too before offering wg-easy at the DR-spare prompt
Requested: don't push the operator toward installing wg-easy if they already have a different mesh VPN (Netbird or Tailscale) running — detect any of the three first, and only offer a choice when none are present. Detection checks wg-easy's own directory (this repo's install marker), then falls back to checking whether the netbird/tailscale binaries exist AND their systemd services are actually active — not just installed, since an installed-but-never-connected client isn't a usable path to the spare box either. wg-easy takes priority if somehow more than one is present, since it's this repo's own chain-installable option. When none are detected, offers a numbered choice: wg-easy (chain-installs via the existing declare -F guard), Netbird, or Tailscale (both via their official curl-pipe-sh installers — verified the current URLs against each vendor's own docs rather than guessing, since a wrong URL here would be a bad thing to ship). Both third-party options still need a manual follow-up step this script can't complete unattended (Netbird needs a setup key from the operator's account, Tailscale needs an interactive auth link) — the success message says so rather than implying the install alone finishes the job. Verified the detection branching against all the cases that matter: nothing present, only wg-easy's directory, only Netbird active, only Tailscale active, and multiple present at once (wg-easy correctly wins). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
8dd66945dd |
Offer to actually set up VPN + SSH keys at the DR-spare prompt
Requested improvement: the disaster-recovery spare prompt in backup.sh already ran a live connectivity check and, on failure, printed manual instructions (set up wg-easy separately if the spare isn't reachable, run ssh-keygen/ssh-copy-id yourself) — but never offered to do any of it right there, even though every piece is safe to automate inline. Now, when the passwordless SSH check fails: - If wg-easy isn't installed yet, offers to chain-install it (guarded with declare -F install_wg-easy, same pattern asterisk.sh already uses for security-dashboard/pstn-trunk) — covers the common case where the spare is a home box with no port-forward and no path there at all yet, not just a missing key. - If root has no SSH key, offers to generate one (ssh-keygen -t ed25519). - Offers to run ssh-copy-id against the spare interactively right there — it prompts for the spare's login password itself, so this script never touches or sees that password, just invokes the real command inline instead of telling the operator to go run it themselves after. - Re-runs the connectivity check after ssh-copy-id succeeds, so the install flow reports the actual current state instead of the pre-fix failure message. Verified the has-a-key detection (the part most likely to have a subtle &&/|| precedence bug) against all four cases — no key, only id_ed25519, only id_rsa, both — behaves correctly in each. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
90b0508e19 |
Reuse an existing destination's repository password on re-run
Confirmed live: backup.sh has no update/fresh distinction and re-runs every prompt on every invocation, including the repository password prompt — which always minted a fresh (typed or auto-generated) password regardless of whether a repo already existed at that destination's path. Re-running the installer (to add a destination, configure the new B2 offsite mirror, or just by habit) then fails to connect to the real, already-populated repo with "invalid repository password", because the repo's actual password is permanently whatever was set the first time and nothing read that back. Each destination's password is now read back from the existing backup.conf (if that destination name was already configured there) before falling through to prompt/auto-generate — same pattern already applied to REMOTE_TYPE/REMOTE_ARGS, EMBEDDED_COTURN_SLOT, and everywhere else in this session that re-running a script with no update/fresh gate turned out to silently regenerate something it shouldn't have. Verified against a mock backup.conf: an existing destination's password is reused verbatim, and a genuinely new destination name still falls through to fresh generation correctly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
779a49b02a |
Merge main into ionos-script-integration-x32ofw, resolve coturn conflicts
Both branches independently solved the same Mattermost/Asterisk coturn relay-port collision problem. Kept main's find_free_coturn_range-based approach (documented in CLAUDE.md as the canonical pattern, and shared across every coturn-owning service) over this branch's earlier Mattermost-only port-slot scheme, and cleaned up the now-unused _MM_COTURN_PORT/_MM_COTURN_MIN/_MM_COTURN_MAX/EMBEDDED_COTURN_SLOT references that had auto-merged without conflict markers. |
||
|
|
c48ed039e0 |
Guide + automate Backblaze B2 offsite mirror setup in backup.sh
Answers a direct ask: offsite mirroring existed only as a REMOTE_TYPE/ REMOTE_ARGS placeholder in backup.conf with a comment pointing at `kopia repository sync-to --help` — no interactive setup at all, B2 or otherwise. Checked before building anything: Kopia's dedicated `sync-to b2` provider is marked [DEPRECATED] on kopia.io's own command reference. B2 also offers an S3-compatible endpoint (s3.<region>.backblazeb2.com, same application key works as the access/secret key pair), and Kopia's `sync-to s3` provider isn't deprecated — so this targets that path instead of building on a command on its way out. What's now automated vs. guided, deliberately split: - Bucket creation and the application key are walked through as console steps, not automated. Object Lock specifically is a one-time, bucket-creation-only decision with a real tradeoff (undeletable-by- design vs. genuinely can't delete early) that shouldn't be silently flipped either way by a script on someone's behalf. - Once the operator has a bucket + endpoint + scoped application key (B2 requires a key scoped to one bucket, not the account master key — noted in the walkthrough), this becomes mechanical: run a `sync-to s3 --dry-run` against the just-created 'default' repo to verify the credentials actually work, and only then write REMOTE_TYPE=s3 / REMOTE_ARGS into backup.conf. A bad bucket name or key leaves REMOTE_TYPE at "none" with a clear error instead of saving a broken config that fails silently at 2am. - Encryption isn't a separate step — Kopia already encrypts client-side with the repository password set earlier in this same flow; called that out explicitly since it was asked about as if it needed its own setup step. Also fixed a regression the new prompt would otherwise have caused: backup.sh has no update/fresh distinction and re-asks everything on every run, so an already-configured offsite mirror is now read back from the existing backup.conf and preserved by default — answering "no" on a re-run no longer silently resets REMOTE_TYPE to "none". Verified the control flow (not just bash -n) against a mock kopia binary and stubbed prompts: good credentials wire up REMOTE_TYPE/ REMOTE_ARGS correctly, a rejected credential leaves REMOTE_TYPE at "none" rather than saving something broken, an existing configured value survives a "no" answer on re-run, and blank fields skip cleanly without attempting a dry-run at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
93d5459a67 |
Make the backup restore-test schedule configurable, add a run-now option
Answers a direct ask: the automated restore-verify test (extras/test_backup_kopia.sh — verifies the latest snapshot, restores it over a moved-aside copy, compares, rolls back, reports PASS/FAIL, sends an ntfy notification) was already fully non-interactive and already wired to a systemd timer/cron fallback by install_backup() — it just had no schedule choice at all, hardcoded to weekly (Saturday 03:00). Every service in this test stops briefly while its data gets moved aside and restored back, same interruption profile as the main backup job — so the schedule is a real tradeoff (more frequent verification vs. more frequent blips), not a free "always pick the most frequent" choice. Gave it the same Weekly/Monthly/Custom shape the main backup schedule prompt above it already offers, instead of a single hardcoded option. Also added an explicit "run the first test now?" prompt right after scheduling it — otherwise choosing Monthly means waiting up to a month before finding out whether the test even works, rather than getting that initial confirmation immediately and then settling into the chosen cadence. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
624ae3d2f3 |
Stop Mattermost's WEB_PORT/CALLS_UDP_PORT rescanning on every update
Found while adding a live scanner to the coturn-slot code: WEB_PORT and CALLS_UDP_PORT were scanned unconditionally, before the reinstall-mode prompt even ran and before anything stopped the currently-running container. On an "Update" run that meant find_free_port would see this instance's OWN already-published port as occupied and silently shift it to the next free one — every plain update could have moved the service's port out from under already-configured Caddy routes, bookmarks, and the Calls plugin's client config, without the operator asking for that. services/asterisk.sh already gets this right for WEB_ADMIN_PORT: update reads the existing port back from .env (no rescan), fresh scans from the plain default only after stopping the old container. Brought Mattermost in line with the same shape — the port resolution moved from before the reinstall-mode block to after it, so MODE is known and, for a fresh install/"Full reinstall", the old containers are already stopped by the time it scans. WEB_PORT/CALLS_UDP_PORT are now also written to .env directly (they weren't before), with a fallback to parse them from the existing MM_SERVICESETTINGS_LISTENADDRESS / docker-compose.yml port mapping for installs made before this change — so an update on an already-running instance doesn't regress just because its .env predates the new variables. Verified the explicit-var, fallback-parse, and priority-order (explicit wins over fallback) cases against a mock before shipping. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
cce8147059 |
Live-verify a newly assigned Mattermost coturn slot isn't already bound
Requested check: the slot-allocation scheme added in the previous commit only checked against OTHER mattermost*/.env files on the box, not against what's actually listening. A slot whose numbers happen to be free by that bookkeeping could still be squatted by something this script doesn't track (a manually-run process, an unrelated service) — this box already learned that lesson once, from Asterisk and Mattermost's embedded coturn ranges overlapping without either side knowing. Only a NEWLY assigned slot gets the live check — an already-cached slot (read back from this instance's own .env) is trusted as-is, since a live conflict on an already-configured, already-running instance's own port is a real problem to report, not something to silently route around by moving that instance's TURN port out from under it. Can't scan the full 200-port relay range port-by-port (large ranges use the offset scheme instead of scanning per CLAUDE.md's port-collision section) — checks the control port plus both relay-range boundaries as the practical middle ground. Verified against a mock: a candidate slot whose control port is already bound gets skipped in favor of the next free one. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
4a09d3d1de |
Give each Mattermost instance's embedded coturn its own port slot
Follow-up to the Asterisk/Mattermost relay-range overlap fix: that fix only handled the two-service collision, and left a documented gap for what happens when a second (or third...) Mattermost instance also falls back to embedded coturn — they'd have collided with each other on the same fixed 3479/49253-49452 numbers, same bug, different pair. find_free_port-style scanning doesn't work for the relay range itself — it's a scan for a single free port, not a free contiguous 200-port block — so this follows the same fixed-offset-per-instance approach CLAUDE.md documents for traccar.sh's large port range instead. Each instance gets an integer slot (control port = 3479 + slot, relay range = 49253 + slot*200 through +199) computed once as the smallest slot number not already claimed by another mattermost*/.env on the box, then cached in that instance's own .env as EMBEDDED_COTURN_SLOT so it reads back the same value on every later update or full reinstall instead of potentially landing on a different slot (which would silently move an already-configured instance's TURN port out from under it — the same "never touch what's already the box's answer" rule everything else in update mode already follows). Verified the allocation logic against a mock: first instance gets slot 0, a second gets slot 1 without stepping on the first, both instances keep their own slot across a simulated re-run, and a third new instance correctly lands on the next free slot (2) rather than reusing either. Threaded the computed port/range through every place that used to hardcode 3479/49253/49452: the coturn compose block, the UFW rule (now also labeled with the instance suffix, matching this file's other UFW comments), and the Calls-plugin TURN config text in the generated README/System-Console instructions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
158df0d545 |
Fix embedded-coturn relay-port overlap between Asterisk and Mattermost
Confirmed live on a box that retired the shared coturn service in favor of each service running its own dedicated/embedded coturn permanently: Mattermost's embedded-coturn fallback used relay range 49153-49352, which overlaps Asterisk's embedded coturn range (49152-49252) by ~100 UDP ports. Both run network_mode: host, so with shared coturn out of the picture this is the exact same collision CLAUDE.md documents as the original, already-fixed-once bug that the shared coturn service was built to solve in the first place — reintroduced here because Mattermost's embedded-coturn fallback path apparently never got checked against Asterisk's numbers when it was written. Moved Mattermost's embedded relay range to 49253-49452 (same 200-port width, now contiguous with and non-overlapping Asterisk's 49152-49252). Updated the docker-compose command flags, the matching UFW rule, and added a comment explaining the offset so it doesn't drift back into collision — and noting the known residual gap this doesn't cover: two Mattermost instances *both* falling back to embedded coturn at once would still collide with each other on these same fixed numbers. Not fixed here since it requires more than one Mattermost instance to be running without shared coturn at the same time, which isn't this box's situation; flagged in-code for whoever hits it. Also made asterisk.sh's generated README port table stop unconditionally claiming a TURN relay range it isn't actually publishing when the shared coturn service (not this install's own container) is fronting TURN instead — it now branches on USE_EMBEDDED_COTURN, which the function already receives as a parameter. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
54f99d8403 |
security-dashboard: drop the port from Sipnetic QR's TURN server field
Sipnetic's own documented "st" field format is explicit that the value is a hostname/IP without a port -- its worked example (turn:user:pass@host) has no port anywhere, even in the URI-with-credentials form. Appending :3478 as this repo was doing gets silently truncated by the app: confirmed live, the FQDN came through on scan but the port after it did not. coturn's listening port in this repo is always the STUN/TURN-conventional 3478 anyway, which is what a portless address implies, so stripping it before building the st field costs nothing and matches the actual spec. |
||
|
|
c2e02f78ae |
asterisk vendor: always append TURN port even when TURN_SERVER is host-only
entrypoint.sh's turn_server fallback only appended :TURN_PORT when TURN_SERVER was completely unset — a TURN_SERVER carried over from an older install (or set to a bare host by hand) passed straight through with no port, so the Sipnetic QR export's "st=turn:user:pass@host" field ended up missing the port entirely. Append it whenever the configured value has no colon at all, not just when it's empty. |
||
|
|
ca239a3886 |
Retire the shared coturn service — every WebRTC/SIP service now runs its own
Shared coturn (services/coturn.sh, ensure_coturn_user in lib/common.sh) is no longer an installable or usable option anywhere in this repo. It's moved to attic/coturn.sh (with tools/coturn-test-check.sh alongside it), which is outside setup.sh's services/*.sh glob, so it never registers, never appears in the menu, and `sudo ./setup.sh coturn` now fails with "unknown service". Asterisk and Mattermost each already had an opt-out to run their own dedicated coturn instead of the shared one; that opt-out is now the only behavior — the shared-coturn preference, the opt-out prompt, and every ensure_coturn_user() call site are gone. find_free_coturn_range() (lib/common.sh) is what makes unconditional dedicated coturn safe: it scans every coturn-owning service's own .env on the box for already-claimed relay ranges and picks one that can't collide, so Asterisk + any number of Mattermost instances can each run their own coturn on one box without the relay-port collisions this repo's coturn history warns about. Existing installs still pointed at a shared coturn container are left running as-is on `update` (no silent migration attempt against a service that no longer exists to heal against) — a full/fresh reinstall is the migration path, which generates a new dedicated coturn with fresh credentials and says so. Also updates CLAUDE.md's coturn guidance for future service authors, attic/README.md with the retirement rationale, and stale services/coturn.sh path references in services/asterisk.sh, tools/pstn-test-check.sh, README.md, and docs/vps-sizing-recommendations.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Crt4ymNEHEbWqscB1qvZgC |
||
|
|
d835e0d734 |
mattermost, asterisk: dynamic, collision-safe dedicated coturn ranges
Mattermost's own embedded-coturn fallback hardcoded the same relay range (49153-49352) for every instance, with no per-instance offset -- running two Mattermost instances without the shared coturn service (or one alongside Asterisk's own dedicated coturn, now possible via the prior commit) would silently reproduce the exact pre-merge collision bug this repo's coturn history warns about, just among Mattermost instances instead of Asterisk/Mattermost. Adds find_free_coturn_range() (lib/common.sh, standalone-mode-stubbed in both services matching the existing port_in_use/find_free_port convention): a coturn relay range can't be collision-checked with live socket scanning the way a single port can -- coturn only opens ports inside its configured range on demand, so an idle range looks the same as an unclaimed one to ss/netstat. The only reliable check is reading what every other coturn-owning service's .env on the box actually claims (COTURN_MAX_PORT for the shared instance, TURN_MAX_PORT for each dedicated one) and picking a range starting safely past the highest. Also adds Mattermost's own opt-out prompt for the shared coturn preference, matching the one just added to Asterisk (fresh-install-only, never re-asked on update, same as every other coturn-shape decision in that file). An update now explicitly preserves its existing dedicated range from .env rather than silently recomputing a new one. Verified end-to-end: a shared instance + Asterisk's dedicated coturn + two independent Mattermost instances, each discovering and avoiding every range already claimed by the others, land on entirely non-overlapping port blocks. |
||
|
|
85fc172377 |
asterisk: avoid relay-port collision between dedicated and shared coturn
Answers "can Asterisk run its own coturn while Mattermost keeps using the shared one" -- yes, but not safely until now: Asterisk's dedicated coturn hardcoded relay ports 49152-49252, entirely inside the shared instance's own default range (49152-49452). Running both on the same box (now possible via the previous commit's opt-out prompt) recreated the exact pre-merge collision this repo's coturn history warns about. When a shared instance is present, read its actual configured COTURN_MAX_PORT from ~/docker/coturn/.env and pick a dedicated range starting safely past it, so the two can never overlap regardless of what the shared instance was configured with. No shared instance on the box means no collision risk, so the historical 49152-49252 default is left untouched in that case. Threaded the computed range through every place it was previously hardcoded: the coturn container's own --min-port/--max-port, the UFW rule, the DigitalOcean Cloud Firewall rule list, the non-DO firewall reminder, and the generated README's port table. Verified the range-shift arithmetic directly: a shared instance configured up to 49452 shifts the dedicated range to 49502-49602 (clear); no shared instance leaves it at the original default. |
||
|
|
3abc46d8a9 |
asterisk: add opt-out for shared coturn on fresh reinstall
ensure_coturn_user always preferred the shared coturn service with no override once it was reachable -- there was no way to deliberately run Asterisk's own dedicated coturn again short of stopping the shared service outright (which would also break every other consumer, e.g. Mattermost Calls). Useful for reproducing an older install's exact shape when troubleshooting anything that might be specific to the coturn-sharing path. Only offered when a shared instance actually exists, and only reachable via an explicit fresh reinstall, matching this repo's existing rule that coturn shape never changes silently on an update. |
||
|
|
65b7964411 |
asterisk: mount Caddy's cert store regardless of coturn mode
The read-only bind mount that lets Asterisk's entrypoint auto-sync a real Let's Encrypt cert from Caddy (instead of falling back to self-signed) was gated on USE_EMBEDDED_COTURN == true. That condition conflated two unrelated things: Asterisk's own SIP transport-tls cert (what this mount is actually for) and coturn's separate TURNS capability (which the shared coturn service genuinely doesn't support, but is irrelevant here). Confirmed live: on a shared-coturn install with a real Caddy-issued cert already sitting on disk for DOMAIN_NAME, Asterisk kept generating a self-signed cert on every restart anyway, because /caddy-data was never mounted into the container -- sync_caddy_cert() had no cert store to find. Most SIP/TLS clients refuse a self-signed cert outright with no clear error, which was the actual cause of a "port's open, cert domain matches, registration still silently fails" case where every other layer (firewall, coturn reachability, DNS, cert CN/SAN) had already checked out clean. |
||
|
|
b76230b341 |
asterisk: regenerate self-signed TLS cert when DOMAIN_NAME changes
Vendor's entrypoint.sh only regenerates the self-signed cert if the file is missing or lacks a SAN extension -- it never checks whether the SAN actually matches the currently configured DOMAIN_NAME. Since /etc/asterisk/certs is a bind-mounted host directory, neither an update nor a full reinstall ever wipes it, so a domain entered once (even a placeholder, or one later changed) sticks in the cert indefinitely. Confirmed live: a box kept presenting a cert for a stale, originally- entered domain long after DOMAIN_NAME had changed and a full reinstall had run in between. Most SIP/TLS clients refuse a mismatched cert outright with no clear error, which was the actual cause of a "port's open but registration still fails" case -- firewall, coturn, and DNS had all already checked out clean. Patches the vendored entrypoint.sh (same guarded-sed pattern as the existing logger.conf patch) to also regenerate when the existing cert's SAN doesn't include the current DOMAIN_NAME. Verified against a scratch copy: missing cert regenerates, a cert already matching the domain is left alone, a mismatched domain now correctly regenerates and then stabilizes. |
||
|
|
893ab19759 |
asterisk: warn non-DO public installs about provider-side firewalls
FQDN mode was already available outside DigitalOcean detection (the home/LAN path's "Networking mode" menu offers it), but only DO installs got any reminder about a network-edge firewall sitting in front of the box -- non-DO public VPS installs got no equivalent, and UFW being wide open gives no signal that a separate provider-managed firewall exists at all. Confirmed live on an IONOS VPS: UFW allowed every SIP/TURN/RTP port, Asterisk's own PJSIP logger showed zero incoming packets, and nothing in the installer's own output pointed at the cause -- IONOS's own network firewall (Cloud Panel -> Networking -> Firewall Policies) only allowed 22/80/443/8443/8447 and silently dropped the rest before it ever reached the box. Adds _asterisk_remind_non_do_firewall(), fired whenever a fresh install sets a public FQDN without being in DO/droplet mode: same port list as what UFW just opened, plus a pointer at the IONOS console location as a concrete example other providers can generalize from. |
||
|
|
2b926ebf0b |
security-dashboard: widen main container for the Extensions table
The previous "box too narrow" fix targeted the QR popup, but the actual complaint (confirmed by screenshot) was the Extensions table itself -- ten columns (Ext/Name/Mobile/Status/Transport/PSTN/Whitelist/Messaging/ Voicemail/actions) forced .table-wrap's horizontal scrollbar even on a normal desktop viewport because main was capped at 1180px. Bumped to 1600px; verified via headless render at 1280-1920px that the table no longer overflows. |
||
|
|
5d0b6355b1 |
security-dashboard: put TURN creds in the Sipnetic QR, widen the popup
- ea_device_sipnetic_string() now sets Sipnetic's documented st= field to an explicit turn:user:pass@host:port URI built from the same TURN_SERVER/TURN_USERNAME/TURN_PASSWORD Asterisk itself reads from its .env (the shared VPS coturn on a droplet, or whichever coturn Asterisk is actually configured against). Previously the QR carried no TURN info at all, silently falling back to Sipnetic's own default STUN server instead -- registration/media then depends on whatever got typed in by hand instead of what Asterisk is actually using. - Popup widened (192px content -> 320px card) and the QR rendered at 3x its displayed resolution (physical size unchanged): the longer TURN-inclusive account string needs a denser code, and verified via a headless render + OpenCV/pyzbar decode that the extra module density needs the resolution bump to stay reliably scannable. - Restored (and expanded) the plain-text-credentials warning that was dropped when the box became a modal, now covering TURN creds too. |
||
|
|
779afcba62 |
security-dashboard: add white quiet zone around Sipnetic QR code
The QR popup's code was unreadable by real scanners: qrcodejs draws modules edge-to-edge with no margin of its own, so the code sat directly against the modal's dark background with no quiet zone. Verified with a headless render + pyzbar/OpenCV decode that the raw generated image had the code running to its edge and failed OpenCV's detector outright, while wrapping it in a 20px white padded frame (still ~2in overall) fixed it. |
||
|
|
f0d34a0028 |
security-dashboard: show extension Sipnetic QR code in a popup modal
Converts the existing inline QR toggle on the Extensions tab's detail panel into a small (2in square) modal popup with an X close button, click-outside, and Escape-to-close, instead of an expanding inline box. |
||
|
|
aa65b5ef5b |
Make "Full reinstall" a real teardown for asterisk, mattermost, and coturn
Extends the security-dashboard prototype to the shared-coturn trio, since these three are exactly the case that pattern was built for — a fresh reinstall of any of them today just overwrote files in place without stopping old containers first, and coturn's own fresh path never made an informed choice about the consumer credentials/database it happens to leave alone (safe today, but by omission rather than design). - asterisk.sh / mattermost.sh: "Full reinstall" now stops the existing containers (`docker compose down`) before falling through to the normal install flow, and asks a single explicit question — delete stored data (PBX config/spool/voicemail for Asterisk; Postgres db/uploads/config/ plugins for Mattermost) — defaulting to preserve. Their shared-coturn TURN credential is deliberately left alone either way (reused from cache via ensure_coturn_user(), same as update) — it's not this service's own data, and coturn already handles that continuity. Mattermost's existing "_db_has_data" check already reads the filesystem to decide whether to reuse or regenerate DB_PASS, so the wipe/preserve choice composes with that for free — no separate flag needed. Asterisk's warns to re-run pstn-trunk afterward if data is wiped, since that's what actually goes stale (its dialplan patch), not the fabricated "AMI secret" framing an earlier draft of this warning used before I checked the actual code. - coturn.sh: "Full reinstall" now lists which consumers are currently registered (from users/*.env) and asks explicitly whether to also wipe TURN credentials and the user database, instead of silently preserving them as an unexamined side effect of never deleting the directory. Defaults to preserve. If the operator does choose to wipe, the running container is restarted afterward — it holds the old, now-deleted turndb file open, so new turnadmin writes to the fresh file would otherwise go unseen until a restart anyway. Every affected consumer already self-heals a missing credential on its own next Update run via ensure_coturn_user()'s existing cache-miss path — no changes needed there, just confirmed it covers this case. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
7a1f09a0f7 |
Rename reinstall-mode prompt; make security-dashboard's "Full reinstall" a real teardown
Two-part change discussed and scoped in this session before touching
anything:
1. Rename "Reinstall in place" (r) -> "Update" (u) and "Full install" (f)
-> "Full reinstall" everywhere the prompt appears: lib/common.sh's
shared prompt_reinstall_mode(), plus the three services that carry
their own duplicated standalone-stub copy of it for standalone
execution (asterisk.sh, coturn.sh, wordpress.sh — per this repo's
documented standalone-bootstrap pattern). Internal state values
(update/fresh/cancel) are unchanged, so no other service's case
statement needed touching. docs/anveo-direct-setup-guide.md's `r`
reference updated to `u` to match. attic/asterisk-digital-ocean.sh
deliberately left alone — this repo's own policy is to not backport
fixes into attic/.
2. security-dashboard.sh's "Full reinstall" now does a real teardown
before reinstalling — stops and removes the systemd unit, sudoers
grant, Caddy site block, and secdash system user, then proceeds
through the normal fresh-install flow — instead of just overwriting
files in place while leaving the old service running underneath.
Prototype for a pattern discussed for other services later: split the
destructive question out explicitly ("also delete
dashboard-admins.conf — per-admin extension scoping?", default n) so
full reinstall doesn't silently discard state a plain "start over"
request wouldn't expect to lose. Verified the backup/restore mechanics
(mktemp, copy out before teardown, copy back after) against a mock
under `set -u` for both the preserve and wipe paths before shipping.
Update mode was already the strongest existing example of surfacing
newer optional prompts (its "Reconfigure Caddy protection?" /
"Reconfigure per-admin scoping?" sub-prompts already cover every setting
fresh-install offers) — no changes needed there for this service.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
d55a7cc81a |
Reach the coturn self-heal check from asterisk.sh's update path too
Direct follow-up to the previous commit's ensure_coturn_user() fix: that
fix is useless for Asterisk specifically unless something actually calls
ensure_coturn_user("asterisk") again, and the update ("Reinstall in
place") branch returns 0 well before the fresh-install path's call to it
— only "Full install" reached it, which re-prompts everything (droplet
detection, domain, etc.) just to fix a credential re-registration.
Added the same call to the update path, gated on NOT having an embedded
coturn (checked via the existing _HAD_EMBEDDED_COTURN detection) — calling
it unconditionally would silently chain-install the shared coturn service
for a box deliberately running Asterisk's own dedicated coturn, exactly
the kind of silent update-time migration CLAUDE.md's coturn guidance
warns against. .env stays untouched either way (self-heal re-registers
with the same cached password, never generates a new one).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
f94f43786f |
Self-heal orphaned coturn credentials in ensure_coturn_user()
Answers a direct question from this session: no, reinstalling asterisk/mattermost did NOT fix a coturn user missing from the live database, because ensure_coturn_user() only ever calls turnadmin -a in the else branch — reached only when the cache file (users/<consumer>.env) is MISSING. A stale-but-present cache file (exactly what a coturn container/volume recreation without preserving ./db leaves behind, per this session's real diagnosis) looked identical to a healthy one and was trusted blindly, so every consumer's installer kept silently reusing credentials that no longer existed in coturn's database. Now checks the cached username against coturn's actual live user list on every call, and re-registers it with the same cached password if it's missing — the same self-heal pattern this repo already applies elsewhere (Beszel's compose patch, Vaultwarden's SMTP half-state, FMD's chown). Re-uses the turnadmin -l log-noise filter from tools/coturn-test-check.sh (a real "user[realm]" line never contains a space; at least one coturn build writes its own startup log lines to stdout, not stderr, so a bare 2>/dev/null doesn't catch them). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
cd1d0e3b40 |
Filter turnadmin -l log noise before parsing usernames
Live run surfaced it: this coturn build writes its own startup log lines
("INFO SQLite connection was closed.", "INFO log file opened: ...") to
turnadmin -l's STDOUT, not stderr — 2>/dev/null never caught them, so
they got parsed as if they were usernames, producing nonsensical
"Database has user '2026-...INFO SQLite connection was closed.'" warnings
on a real run. A genuine "user[realm]" line never contains a space; every
log line does, so filtering on that is a simple, build-independent fix.
Also diagnosed the actual underlying failure this surfaced: coturn's live
user database was genuinely empty (both 'asterisk' and 'mattermost' had
cached credential files but neither was registered in the DB) — exactly
the container/volume-recreated-without-db drift this script's consumer
cross-check exists to catch, confirmed against a real run.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
8ee018ec10 |
Distinguish timeout from real TURN test failure, bump timeout to 20s
Latest live run showed the test getting killed by its own `timeout 10` before turnutils_uclient printed any result — just two startup INFO lines, no error. That's the coturn/coturn Docker image's turnutils_uclient (apparently a newer build with structured "LEVEL component: message" logging, different from the older packaged version available for local testing) taking longer than 10s to complete, not a real failure. Bumped both scripts' timeout to 20s, and now check for timeout(1)'s own exit code (124) separately from a real reported error — reported as WARN with a suggested manual command to re-run with more time and see the full result, instead of lumping "still running" in with "actually failed." Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
1624b76a82 |
Switch TURN test from -e <peer> to -y — real verification this time
Not another guess: installed coturn locally (apt-get install coturn) and
ran the actual server + turnutils_uclient against it to verify this
before shipping, since the last two rounds shipped based on reading the
usage text alone and both turned out incomplete.
-e 127.0.0.1 satisfies turnutils_uclient's "-e or -y required" check, but
then fails allocation with "channel bind: error 403 (Forbidden IP)" —
services/coturn.sh never sets --allow-loopback-peers, so loopback as a
peer address is correctly rejected by a real coturn instance, and the
previous fix's own comment about "loopback is always reachable" missed
that reachable and permitted aren't the same thing.
-y ("client-to-client") sidesteps this: it negotiates both ends of a real
relay through the server itself, needs no separate peer address, and
works fine over loopback. Verified directly against a real local
instance: exits 0 with real packet-loss/RTT stats on valid credentials,
and correctly fails ("Cannot complete Allocation", exit 255) on a wrong
password — so it's still a meaningful pass/fail, not just "didn't crash."
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
9f4f779220 |
Fix "Either -e peer_address or -y must be specified" in TURN allocation test
Another real failure from a live run: turnutils_uclient refuses to run at all without either -e <peer> or -y — a bare auth-only invocation isn't enough for it to actually attempt anything. Add -e 127.0.0.1 to both tools/pstn-test-check.sh's and tools/coturn-test-check.sh's invocations; loopback is always reachable since the test already runs via `docker exec` inside the coturn container itself, and it lets the test actually prove data relays through the allocation, not just that auth succeeded. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
2ac00c946b |
Merge remote-tracking branch 'origin/main' into codex/fix-drum-rhythm-game-audio-issues
# Conflicts: # services/drum-rhythm-game.sh |
||
|
|
6b199f1b22 |
Merge remote-tracking branch 'origin/main' into codex/fix-drum-rhythm-game-audio-issues
# Conflicts: # services/drum-rhythm-game.sh |
||
|
|
5dbfcbd120 |
Fix false TURN allocation failure, add attention recap, one-at-a-time reprint
Real bug caught from a live run: the coturn allocation test passed -t -T
(TCP/TLS) to turnutils_uclient, but services/coturn.sh always starts
coturn with --no-tls --no-dtls — requesting an encrypted/TCP transport
against a server that never offered one fails the allocation outright
("Cannot complete Allocation"), misreporting a config problem that didn't
exist. Dropped both flags in both tools/pstn-test-check.sh and
tools/coturn-test-check.sh so the test matches what the server actually
supports (plain UDP).
Also, from user feedback on the same run:
- warn()/fail() now collect their messages into arrays; the Summary
section prints a "Needs attention" recap of every FAIL/WARN together
at the end, instead of leaving the user to scroll back through a long
run to find what needs fixing.
- The softphone-setup block now offers to reprint itself one extension
at a time (paced with a keypress between each) after the main run, so
a long device list isn't lost in the scrollback either. Factored the
per-extension print into print_ext_info() so the full run and this
reprint can't drift apart. Guarded with `[ -t 0 ]` so it's skipped
automatically when the script isn't run interactively.
Verified via a fuller mock harness (fake docker/curl/systemctl/getent,
non-TTY stdin) that: the corrected turnutils_uclient invocation reports
success, the recap correctly lists FAIL before WARN, and the interactive
reprint prompt is skipped without hanging when stdin isn't a terminal.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
bb7aece023 |
Test Asterisk's own coturn config and print softphone setup info
Two follow-ups on the PSTN health check:
- New "coturn (TURN relay for Asterisk)" section reads Asterisk's own
TURN_* values from its .env (not re-derived) and runs a live TURN
allocation against whichever coturn Asterisk is actually configured to
use — the shared instance, or its own embedded per-Asterisk coturn if
that's what this box has (detected via the same "grep -q '^ coturn:'
docker-compose.yml" check CLAUDE.md's migration guidance describes).
Proves what Asterisk itself would use at call time, complementing
tools/coturn-test-check.sh's broader multi-consumer check.
- New "Softphone setup" section parses pjsip.conf directly and prints
per-extension SIP server/username/password/port/transport, plus TURN
credentials for any extension with ice_support=yes — the same values
Sipnetic's "Add Account" screen needs, computed here so a client isn't
installed just to read them out of the Security Dashboard.
Also fixed a bug caught while building a mock test harness to verify both
additions: the extension-registration parser grabbed state via a fixed
field position ($3), silently truncating multi-word states like "Not in
use" down to "Not". Replaced with a regex that captures everything
between the extension and the trailing "N of inf" — verified against both
single- and multi-word states.
And a real syntax bug caught by bash -n before this ever shipped: an
apostrophe inside a ${VAR:-default} expansion ("this box's IP") opens an
unterminated single-quote context even inside double quotes — reworded
to avoid the apostrophe entirely rather than fight bash's parser.
Full mock run (fake docker/curl/systemctl/getent, real pjsip.conf/.env
fixtures matching the actual generated format) confirmed both new
sections and the registration fix all produce correct output end to end.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
5c18cfa364 |
Add SMS test reminder, "which box handled it" note, and a coturn health check
Three follow-ups from live testing on this session's actual VPS: - tools/pstn-test-check.sh's SMS section printed the Forward-to-URL value to configure but never said what to do next — add the "text this DID, then watch journalctl -u sms-inbound -f" step right after it. - docs/pstn-sms-test-checklist.md: the "which box actually handled this" question has a simple answer (a DID's inbound routing targets exactly one IP:port, so there's no ambiguity to resolve, only a portal setting to confirm) — written up so it doesn't need re-deriving. Also fixed the --list example to cd into the repo first; ./setup.sh is a relative path and silently fails with "command not found" from any other directory, confirmed live in this session. - New tools/coturn-test-check.sh: health-checks the shared coturn instance (services/coturn.sh) and every consumer registered against it (Asterisk, any number of Mattermost instances) — container/identity, each cached consumer credential cross-checked against coturn's own live user database (catches the container/volume-recreated-without-db drift case), UFW rules for both the TURN port and the relay range, a capacity explanation reasoned from the actual port-range math instead of a guess, and a real TURN allocation test per consumer via turnutils_uclient — the only way to prove credentials + port range + firewall all actually work together, not just that each looks right in isolation. Deliberately does not attempt a concurrent load test, since that would consume real relay ports other services may be actively using. Verified the turnadmin -l output parsing, UFW rule matching, and the empty-array-under-set–u loop pattern against mock data before shipping. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
efb86fd92d |
Add provider-portal checklist with real values to the PSTN health check
Server-side config was fully verifiable already; what wasn't is the provider-account side (Anveo's authorized-IP list, DID routing, SMS forward-URL) since that lives entirely outside this box. Rather than leave "go check the portal" as a vague pointer, compute and print the exact values each portal field needs to match: this box's public IP, the trunk DID (from .pstn-trunk.env), and the SMS forward URL read straight from /opt/sms-inbound/settings.env (SMS_FORWARD_URL) instead of making the user reconstruct or hunt for a value the installer already generated and stored. Anveo-specific field-by-field checklist when PROVIDER_NAME matches; generic fallback otherwise. Verified the .pstn-trunk.env / settings.env sourcing against mock files matching the real generated format, including the literal $[from]$-style Anveo placeholders in SMS_FORWARD_URL, which must survive `source` under `set -u` without triggering bash's legacy $[...] arithmetic expansion — same guard pattern services/pstn-trunk.sh's own update path already uses. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
52fd75f679 |
Add automated PSTN/trunk/SMS health-check script
docs/pstn-sms-test-checklist.md's manual steps (registration, trunk
reachability, dialplan contexts, kill-switch state, usage-alert timer
health, recent call/message activity) are all things a script can check
directly instead of re-typed by hand each time — and re-typing them is
exactly what led to the container-name mistake in the prior commit.
tools/pstn-test-check.sh auto-detects the container/directory the same
way the checklist doc now does, runs every automatable check, and prints
PASS/WARN/FAIL per item plus a summary. What it can't cover — actually
placing a call or sending a text — still needs the checklist doc.
Caught during testing against real command output pasted in this
session: the endpoint-parsing loop matched pjsip's own column-header
line ("<Endpoint/CID...> <State...>") as if it were a real endpoint row,
producing a bogus result. Fixed by skipping any row whose parsed
extension starts with "<".
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
e03907882c |
Auto-detect the Asterisk container/dir in the PSTN test checklist
$CONTAINER/$EA_DIR only lived in the shell session where they were typed by hand — a new terminal or enough time between test steps left them empty, and an empty $CONTAINER silently collapsed "docker exec -it $CONTAINER asterisk -rx ..." into "docker exec -it asterisk -rx ...", failing with "No such container: asterisk" instead of an obviously-unset-variable error. Confirmed live. Replaced the manual pick with a docker ps auto-detect so a stale/forgotten variable can't silently break every command in the checklist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
d9f6190da1 |
Add PSTN/DID/SMS end-to-end test checklist
docs/pstn-calling-voipms-plan.md (design log) and docs/anveo-direct-setup-guide.md (account + droplet setup) already cover getting a trunk/DID/SMS working from scratch, but neither is a quick top-to-bottom checklist for verifying an already-installed setup still works — registration, trunk reachability, tiers, outbound/inbound calls (shared DID and personal DID), the spend-cap kill-switch, international calling, internal SIP messaging, and SMS inbound, in order, with what to check when each step fails. Pulls known gotchas (Commit Changes required after dashboard tier edits, mobile vs geographic DIDs for verification codes, the SIP-based SMS path Anveo doesn't actually offer) from the existing docs so they're not missed mid-test. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
8c52376495 |
Remove stale Caddy block before rewriting it on a fresh security-dashboard reinstall
The fresh-install path called _secdash_configure_caddy directly with no prior removal, unlike the update/reconfigure path which already calls _secdash_remove_caddy_block first. Re-running a "Full install" over an existing dashboard on the same domain therefore appended a second site block instead of replacing the first — and since Caddy serves whichever block comes first in the file, the old one (old Authelia address, old Basic Auth settings) kept winning even after answering the prompts with new values. Confirmed live: reconfiguring a dashboard from a local to a remote Authelia address left the old forward_auth target still in effect until the stale block was deleted by hand. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
142897f84c |
Self-heal Beszel agent compose files on update
install_beszel() and install_beszel-agent()'s "update" branches only did a pull+restart, never touching docker-compose.yml — so an already-installed box would never pick up the systemd/dbus/sensor mounts or apparmor:unconfined fixes without a manual edit or a disruptive fresh reinstall. Add _beszel_patch_agent_compose(), called from both update branches, that idempotently patches an existing docker-compose.yml with whichever of the two fixes it's still missing. Anchors on `network_mode: host` and the docker.sock mount line, both unique to the beszel-agent service and present in either compose shape (combined hub+agent or agent-only), so one function covers both install paths. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
e299b37b0c |
Add apparmor:unconfined to Beszel Docker agents - dbus mount alone isn't enough
The systemd/dbus mounts added last commit aren't sufficient by themselves on an AppArmor-enabled host (Ubuntu/Debian by default): the dbus "Hello" handshake fails with "An AppArmor policy prevents this sender from sending this message to this recipient", since the container has no AppArmor label the host's dbus-daemon profile recognizes. Only visible at LOG_LEVEL=debug - silent otherwise, which is why the mounts alone looked like they should have worked but didn't. Confirmed live against a real box hitting exactly this error. security_opt: apparmor:unconfined is Beszel's own documented fix (beszel.dev/guide/systemd#apparmor-error) for this exact error string. Added to both Docker-based agent compose generators (install_beszel's combined hub+agent, and install_beszel-agent's remote-only variant). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
933b5b00f4 |
Mount systemd/dbus/sensors into Docker-based Beszel agents for Services/Temp
The hub's "Services" column is systemd unit monitoring (CPU/memory per unit), and "Temp" is hardware sensor readings — neither is Docker container stats, which is what the existing docker.sock mount actually provides. A container is isolated from the host's systemd/dbus and most of /sys by default, so a Docker-deployed agent silently showed both columns empty, with nothing anywhere pointing at why. Confirmed live: a natively-installed agent (no Docker, a plain systemd service) gets both for free just by running as a normal host process, which is what surfaced the gap — a Docker-deployed agent sitting right next to it on another box showed nothing in either column. Added read-only mounts for /var/run/systemd/private, dbus's system_bus_socket, and /sys/class/hwmon + /sys/class/thermal to both Docker-based agent compose generators (install_beszel's combined hub+agent, and install_beszel-agent's remote-only variant). All four are best-effort: if a path doesn't exist on a given host, Docker mounts an empty directory rather than failing the container, so the worst case on an unusual host is an empty column, not a regression or a crash risk. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
88aac103c9 |
Fix FMD crash-loop: bind-mounted db dir needs UID 1000, not $ACTUAL_USER
fmd-server's image runs as a fixed, non-configurable UID:GID 1000:1000 baked into its own Dockerfile (useradd --uid 1000 fmd-server) - nothing like PUID/PGID to override it. The install script chowned the bind-mounted ./data dir to $ACTUAL_USER instead, which only happens to work when that user's host UID is coincidentally 1000. Confirmed live: the container crash-loops forever on "permission denied" creating its sqlite db otherwise - same root-cause shape as the Mattermost UID/GID bug fixed earlier this session, different fixed UID. Fixed at both points a container start can happen: the fresh-install path (chown -R 1000:1000 "$FMD_DIR/data" right after the existing $ACTUAL_USER chown, ordered after it since that one is recursive over the whole directory and would otherwise overwrite this) and the update path (previously unguarded - re-asserted before every docker compose up so a box already stuck in this state self-heals on next update instead of staying broken forever, same self-heal precedent as the Vaultwarden SMTP fix earlier this session). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
2a2aba0c9d |
Add Authelia OIDC provider + register-a-client flow for ActualBudget/Vaultwarden/other apps
Authelia's forward_auth (what this repo already sets up) gates a whole site behind a login page before the request reaches it. This is the opposite direction: an app with its own "Enable OpenID"/SSO setting delegating ITS login to Authelia, via Authelia's separate OIDC PROVIDER feature, which this repo had no support for at all. _authelia_ensure_oidc_provider() enables it once, idempotently: generates an HMAC secret (injected via a _FILE env var, same convention as the existing jwt/session/storage secrets) and an RSA signing keypair, then writes identity_providers.oidc into configuration.yml. The RSA private key has to be inlined as PEM directly in that file — Authelia's jwks schema has no file-path or env-var option for it — so configuration.yml gets chmod 600 once OIDC is enabled, unlike before when it held no raw secrets. _authelia_add_oidc_client() registers an app: presets for ActualBudget (/openid/callback) and Vaultwarden (/identity/connect/oidc-signin, and confirmed its SSO support is now native/upstream, not fork-only) fill in the redirect URI automatically; "Other/custom" covers anything else. Each app gets its own Client ID and a random secret (shown once, only the pbkdf2 hash is stored), and the output tells the operator exactly what to paste back into that app's own OpenID dialog or .env — including Vaultwarden's exact SSO_* env vars, not just generic OIDC endpoint URLs. Wired into the existing "Authelia already exists" menu as a new option, alongside "add another protected domain" and "reconfigure from scratch". Exact CLI output formats, default filenames, and YAML schema were verified against Authelia's own CLI source/docs (crypto rand's "Random Value: " label, crypto hash generate pbkdf2's "Random Password:"/"Digest:" labels, crypto pair rsa generate's private.pem/public.pem defaults) rather than guessed, since a wrong assumption here means a cryptic startup failure or broken secret extraction. The YAML manipulation (client-list insertion, domain extraction from session.cookies) was tested end-to-end against the real mikefarah/yq binary against a realistic mock config, which caught a real bug (extracting the wrong awk field for the domain, "domain:" instead of the actual value) before it shipped. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
59f82c57a6 |
Self-heal a half-set SMTP_HOST/SMTP_FROM in Vaultwarden's .env
Vaultwarden crash-loops outright if exactly one of SMTP_HOST/SMTP_FROM is
set ("Both SMTP_HOST and SMTP_FROM need to be set for email support
without USE_SENDMAIL"). The fresh-install prompt flow already avoids ever
writing that half-state, but "update" mode deliberately never touches
.env (same rule as everywhere else in this repo), so a box whose .env was
written before that prompt-side fix existed - or hand-edited since - stays
stuck crash-looping on every future update too, since nothing ever
re-checked it. Confirmed live on a real box.
New _vaultwarden_fix_smtp_halfstate() detects the half-set state and
blanks the whole SMTP block (matching what the fresh-install prompt does
when SMTP is skipped) rather than leaving it broken. Called right before
every docker compose up this file does - the update path (previously
unguarded) and the fresh-install start prompt (defense in depth, since
that path is already safe by construction) - so it self-heals regardless
of how a box got into this state.
Audited every other services/*.sh for the same half-set-required-pair
pattern (SMTP, MAIL_*, SMTP_HOST-style naming) - Vaultwarden is the only
one that actually writes paired config where a partial state crashes the
container. Authelia's SMTP is mandatory-with-defaults (a different,
non-crashing risk); Mattermost/frigate-notify only mention SMTP in
generated docs, never in config they write.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
d1d234b4d2 |
Fix Gatus false-positive red on Authelia-protected and stale-synced sites
The auto-sync condition "[STATUS] < 400" reads as red for any site behind Authelia's forward_auth: Gatus's probe is never logged in, so it correctly gets a 401 back every time — the site is completely healthy, Authelia is just doing its job, but that 401 fails the condition. Confirmed live: every site the user actually logs into showed permanently red. That single condition also had the opposite bug in reserve: on a genuine outage (connection refused, DNS failure, TLS failure), Gatus reports [STATUS] as 0, and 0 < 400 is true — a fully unreachable site would have silently read as "up". Fixed to two conditions together: "[CONNECTED] == true" (catches the actual outage case) and "[STATUS] < 500" (accepts any real response, including 401/403/redirects from an auth gate, only failing on Caddy's own 502/503/504 when the backend itself is unreachable). Also changed the sync loop to refresh conditions on already-synced endpoints, not just add-missing-ones — the old add-if-missing-only logic meant this fix would only apply to newly discovered domains, leaving every already-synced site (which is most of them, on a live box) stuck on the broken condition forever until removed and re-added by hand. Now every sync run (every 15 minutes via the existing systemd timer, or the one that happens immediately on a Gatus reinstall) self-heals all of them. Verified end-to-end against the real mikefarah/yq binary: an existing caddy-sync entry gets its conditions rewritten in place, an unrelated manually-added endpoint is left untouched, and a newly-discovered domain gets the corrected conditions from the start. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
459bde0f38 |
Make tab completion + backup pruning setup unconditional in setup.sh
Both were only ever wired up from inside install_base(), so a box that went straight to a direct single-service install (sudo ./setup.sh beszel-agent, or any other service) without first explicitly running `sudo ./setup.sh base` never got either — the direct-install branch exits before the guided flow's own `run_service base` call is ever reached. Confirmed live: tab completion doesn't work on a fresh box that installed beszel-agent first. Moved the call site to setup.sh itself, right after the --list/--status early exits (which stay read-only and don't require root) and before every other branch (configure, --remove, direct install, guided flow) — all of which are downstream of that point regardless of which one actually runs. Both helpers are idempotent and already no-prompt by design, so calling them unconditionally on every invocation is safe; skipped under --dry-run (with an equivalent [DRY-RUN] message) so a preview run doesn't write real files. install_base()'s own calls to both are now fully redundant (base.sh has no standalone-bootstrap block, so install_base() is only ever reached downstream of setup.sh's new call site) and removed, along with the two DRY-RUN preview lines that described them there. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
f7cc0e5bc4 |
Add beszel-agent: agent-only Beszel install for remote/homelab boxes
For monitoring a box that isn't the VPS (e.g. a homelab machine): only the agent needs to run there, and it connects OUTBOUND to the hub over HTTPS using the same key + universal token flow the hub-side installer already uses — no VPN, no router port-forwarding, and no FQDN needed on that box, since nothing on it ever needs to be reached FROM the hub. New register_service beszel-agent in services/beszel.sh (a second registration in the same file, precedented by base.sh's base+glow) reuses _beszel_configure_agent's paste/parse UX for the key/token instead of duplicating it — that function's signature changed from a bare hub port to a full login-URL string so both the local-hub path and this new agent-only path can share it. Run on the remote box: sudo ./setup.sh beszel-agent Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
c719197b20 |
Admin-scoping setup: show live extension list, auto-include owned DIDs
Two refinements to the per-admin extension scoping added last commit: - The setup prompt now shows the dashboard's current extensions (pulled from its own running /api/pstn-permissions, reusing list_extensions()'s already-correct pjsip.conf parsing instead of a second implementation in bash) before asking for each admin's list, with a real example built from actual extension numbers instead of a generic placeholder. Shown fresh for every admin added, one at a time. - An admin scoped to an extension now automatically sees that extension's directly-assigned personal DID's call/text history too, not just its internal activity — parse_pstn_calls()/parse_texts() key inbound rows by the DID that was dialed, not the owning extension, so without this a scoped admin would see their own extension's outbound calls but not inbound calls to their own number. New _dids_for_extensions()/ _admin_scope_for_calls() resolve this per-request from pstn-personal-dids.conf's direct (non-ring-group) owner field. Voicemail scoping is unaffected — a mailbox is always keyed by extension number regardless of which DID rang it. Also removed a dead DASHBOARD_ADMINS_HEADER Python constant left over from before the file-writing responsibility settled on the bash side only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
d124902b8d |
Add fail-closed per-admin extension scoping for Calls & Texts and Voicemail
Two admins sharing one dashboard can now each be scoped to their own extensions on the Calls & Texts and Voicemail tabs, while Security Log, CrowdSec, and Extensions stay fully visible to both — Authelia already provides real per-person identity here (Remote-User, forwarded by Caddy's existing forward_auth/import authelia wiring), this just teaches app.py to finally read it for these two tabs instead of ignoring it. New dashboard-admins.conf ([username] -> extensions=), configured via CLI prompts in security-dashboard.sh (offered at install and on reconfigure), read-only from app.py's side — no write access needed since the file is root-managed. allowed_extensions_for_user() is fail-closed by design: an empty/missing file means unrestricted (today's default, unchanged), but the moment one admin is configured, every other identity — an unlisted admin, a typo, or no Authelia identity at all — sees nothing on those two tabs until added. /voicemail/audio checks the same scope directly (not just the list route) so a guessed or copied URL can't bypass the filter. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
8950810cea |
Add voicemail: dialplan/mailboxes in Asterisk, Extensions toggle + Voicemail tab with click-to-play in the dashboard
Asterisk side (services/asterisk.sh): a [voicemail-access] context reachable from every extension (*97 checks your own mailbox, *98<ext> drops a message into another mailbox directly), gated live via AST_CONFIG() on a new "voicemail" flag in pstn-permissions.conf. voicemail.conf gets a skeleton [general]+[default] at install/update, then stays dashboard-owned from there — mailbox lines are never regenerated wholesale by asterisk.sh once the file exists, matching every other install-time-vs-dashboard-owned file split in this repo (.env, firewall rules, etc). Vendor files (entrypoint.sh, easy-asterisk.sh) get patched the same way messaging-dialplan.conf already does, including the live-extensions.conf patch for boxes with existing devices. While tracing the right #include anchor for this, found and then reverted a theoretical "fix" to messaging's own #include position: pstn-trunk.sh's own comment (live-confirmed 2026-07-24) directly contradicts the textbook Asterisk #include semantics I'd assumed, so the safer move was keeping messaging's anchor exactly as already verified working and using the same position for voicemail's own #include. Dashboard side (services/security-dashboard.sh): write_voicemail()/ _apply_voicemail_flag() toggle the flag and a PIN (generated once, kept across future toggles), regenerate_voicemail_conf() keeps voicemail.conf's [default] section in sync, and a module reload takes effect without a full Asterisk restart. Extensions tab gets a Voicemail column next to Messaging, showing the PIN once generated. New Voicemail tab lists every mailbox's messages (parsed from Asterisk's own msgNNNN.txt sidecars) with an inline <audio> player per row — /voicemail/audio validates ext/msg against strict regexes plus a resolved-path containment check before ever opening a file. Dashboard gets read-only ACL + systemd ReadOnlyPaths access to the voicemail spool dir, and a new sudoers-scoped module-reload command. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
3584ad6499 |
Make backup pruning fully automatic, no prompt
The safety net (only ever touches disposable *.backup.* files, never the newest one for any given file) makes this low-stakes enough to just set up unprompted, the same way tab completion already is — matches the user's own read on it. Still fully idempotent (skipped if the timer already exists), so a rerun doesn't re-ask or redo anything. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
c50704e1b3 |
Add automatic tab-completion setup and old config-backup pruning
Two things surfaced from actual use this session: 1. Tab completion (tools/setup-completion.bash, added earlier) required manually editing ~/.bashrc — easy to skip or get wrong (confirmed live: the source line never actually landed the first time). base now wires it in automatically (idempotent, checked by grep first), matching how it already touches ~/.bashrc for SSH Host aliases. 2. No pruning existed anywhere for the *.backup.<timestamp> files ~60 different services create before overwriting a live config (Caddyfile, /etc/fstab, etc) — every one of them backs up, none clean up, so they accumulate forever on a box reconfigured regularly. tools/prune-old-backups.sh prunes by file mtime (not by parsing the timestamp out of the filename — robust to the %Y%m%d-%H%M%S vs %Y%m%d_%H%M%S inconsistency across services), always keeping the single newest backup per distinct file regardless of age. Verified both the normal case (mixed old/new, prunes only the old ones) and the edge case (every backup for a file is old, keeps the newest one anyway) against real fixtures. base offers it as a daily systemd timer (prompted, since it deletes files — unlike the tab-completion wiring, which doesn't). Also added logrotate for Caddy's own access logs (/var/log/caddy/*.log), which had no rotation at all and grow unbounded on an active box. Uses copytruncate specifically: the log directory is bind-mounted into the running Caddy container and read live by CrowdSec, so truncating in place avoids either of them needing to notice or react to a rotation happening. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
cd003dbaf3 |
Add Gatus auto-sync from Caddyfile, promote ensure_yq to lib/common.sh
Adds one Gatus endpoint per Caddy site block automatically, tagged group: caddy-sync — the sync only ever adds/removes entries in that exact group, so anything added by hand (the default external checks, a custom endpoint) is never touched regardless of what the Caddyfile looks like. Offered at install time (syncs once immediately) and, if systemd is available, scheduled via a timer every 15 minutes so a site added or removed later gets picked up without re-running the installer — matches the "schedule that checks the Caddyfile" shape asked for. Domain extraction tracks actual brace depth (reusing the same approach as remove_service's Caddy block removal) rather than a naive line-by-line scan, so it correctly skips the global options block and parenthesized snippet definitions like (authelia) without needing to special-case them by name. Verified end-to-end against a real Caddyfile/config.yaml fixture with the actual mikefarah/yq binary: initial sync adds the right entries and leaves the default "external" group alone, a second run with no Caddyfile changes is a true no-op (0 added, 0 removed), and changing the Caddyfile (removing one site, adding another) correctly adds the new endpoint and removes only the stale one. Also fixes a real gap surfaced while building this: ensure_yq (used by both gatus.sh now and onlyoffice.sh already) checked `command -v yq` alone, which a box can satisfy with a completely different, incompatible yq — confirmed live in this environment, Debian/Ubuntu's own `yq` apt package is kislyuk/yq (a Python jq-wrapper) which silently errors on mikefarah/yq's `e '.path' file` syntax every caller here depends on. Now checks the version string actually identifies as mikefarah's before trusting it, installing to /usr/local/bin (which precedes /usr/bin on Ubuntu's default PATH) if not. Promoted ensure_yq itself from onlyoffice.sh (its only previous user) to lib/common.sh now that gatus.sh needs the same thing, so both share one implementation. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
7c85f326c8 |
Accept Beszel's own "copy for docker compose" snippet directly
Confirmed live: Beszel's Settings -> Tokens & Fingerprints page surfaces
a "copy for docker compose" shortcut as the prominent way to grab the
key/token — not a bare string — so the previous two-prompt flow (paste
plain key, paste plain token) didn't match what people actually have
in their clipboard. The user pasted that YAML snippet into .env by
hand afterward, using the container's raw KEY/TOKEN names and YAML
`NAME: 'value'` syntax instead of what the compose file's own
${AGENT_KEY:-}/${AGENT_TOKEN:-} substitution actually reads — the
agent then failed with "no key provided" since AGENT_KEY was never
actually set.
_beszel_extract_field pulls KEY/TOKEN out of whatever shape the paste
arrives in — YAML mapping (`KEY: 'value'`), compose list style
(`- KEY=value`), or plain `KEY=value` — regardless of quoting. The
agent prompt now accepts a multi-line paste (the whole snippet, or
just the two lines) instead of asking for two separately pre-extracted
values; if no labeled KEY/TOKEN line is found at all, it falls back to
treating the paste as a bare key and asks for the token separately, so
a Beszel version that really does just show plain strings still works.
Verified all three paths directly: the exact mixed KEY:/TOKEN= paste
the user had, the skip path (blank first line leaves .env untouched),
and the bare-value fallback with no labels at all.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
e10e1e90b4 |
Fix invalid docker-compose.yml when Beszel is installed with local Caddy
Confirmed live: "networks.beszel-agent additional properties ... not allowed" — the root-level networks: block (_CADDY_NET_SECTION) was placed between the two services instead of after both. Since it sits at 0 indentation, YAML parsed the following beszel-agent: line as a continuation of the networks: mapping instead of a new services: entry, so the whole beszel-agent service definition got swallowed as if it were a (invalid) child of networks.caddy_net. gatus.sh's identical _CADDY_NET_BLOCK/_CADDY_NET_SECTION pattern never hit this because it only ever has one service, so the same placement is always the last content in the file there. Moved _CADDY_NET_SECTION (the root-level networks: definition) to after both services; _CADDY_NET_BLOCK (the per-service "join caddy_net" snippet) stays right after the hub's own volumes, where it correctly nests under the beszel: service only. Verified by regenerating the compose file with local Caddy present and parsing it with PyYAML: services.beszel and services.beszel-agent are now proper siblings, beszel-agent keeps its image/environment/volumes keys, beszel's own networks: is scoped to just that service, and the root networks: definition is separate and correctly placed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
fc8f57ab51 |
Add bash tab-completion for setup.sh
./setup.sh mat<TAB> now completes to ./setup.sh mattermost, same for flags. Service names are read fresh from services/*.sh on every completion — never a hardcoded list, which would go stale the moment a new service file gets added (matches this repo's own "adding a service = adding one file, nothing generated" rule from CLAUDE.md). Verified live: sourced the script and confirmed completions for "mat" and "--li", and specifically confirmed "bes" resolves to "beszel" — the service added earlier this same session — with zero changes needed to the completion script itself, proving the list is genuinely dynamic rather than something that looked right once and then rotted. Self-locating via its own BASH_SOURCE path rather than a hardcoded install directory, so it keeps working regardless of where the repo is cloned. Works through a leading `sudo` via bash-completion's standard sudo pass-through (enabled by default on Ubuntu). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
7938399e99 |
Add Beszel for lightweight server + Docker monitoring
Answers "what's the best way to see CPU/RAM/disk usage on this box"
(IONOS's own dashboard doesn't expose it) and "does Gatus cover this" —
it doesn't, Gatus is a black-box HTTP check (is the site responding
from the outside), Beszel is white-box host/process monitoring (is the
box under memory/disk pressure, is a container actually running vs.
crash-looping). Complements Gatus rather than replacing it.
Mirrors the hub+agent same-system layout from beszel's own
supplemental/docker/same-system/docker-compose.yml (fetched from the
actual upstream repo, not reconstructed from memory) — hub is the web
dashboard, agent reads /var/run/docker.sock (read-only) to report every
currently-running container automatically, no per-service config
needed as containers get added or removed.
Genuinely a two-phase install: the hub's SSH keypair and universal
token only exist after logging into its web UI once, so this starts
the hub, walks through where to find both values, and finishes wiring
the agent once provided — skipping is fine, a rerun in "update" mode
detects the agent was never connected and offers to finish it.
Verified the generated docker-compose.yml/.env by running the actual
file-writing code path with docker/configure_caddy_for_service mocked
out — confirmed TOKEN/KEY are correctly left as literal
${AGENT_TOKEN:-}/${AGENT_KEY:-} for Docker Compose's own substitution
at "up" time, not prematurely expanded by the heredoc itself.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
eeaa2e6c64 |
Accept plain "remove"/"uninstall" as aliases for --remove
./setup.sh filebrowser remove (no dashes) fell through to the normal install dispatch instead of removing anything, since only the --remove flag form was recognized. Accept the bare words too — order-independent either way (./setup.sh remove filebrowser works the same). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
efcad48755 |
Add a generic service removal command: ./setup.sh <name> --remove
No removal path existed anywhere in this repo — manually removing a
service meant hand-editing docker-compose.yml, the Caddyfile, and UFW
rules yourself, or just leaving orphaned config behind.
remove_service (lib/common.sh) handles the common case: stop/remove
the service's containers (with an explicit y/n on whether to also wipe
data volumes, default no), find and remove its Caddy site block if one
exists, remove any UFW rule tagged with its name, and optionally
delete its ~/docker/<name> directory (default no — keep data as a
safety net unless explicitly confirmed).
The Caddy site block removal (_remove_caddy_site_block) tracks actual
brace depth rather than scanning to the next blank line or EOF — the
same class of bug this repo already hit once with a naive Samba
config edit. Verified against a multi-block test Caddyfile with nested
log{}/header{} blocks: removes exactly the targeted block, leaves
every other block (including ones with their own nested braces)
byte-for-byte intact, and is a safe no-op when nothing matches.
Also fixes the UFW rule-number extraction: ufw status numbered pads
single-digit rule numbers with a leading space ("[ 3]" vs "[10]") to
align columns, which the regex didn't account for — every single-digit
rule would have silently never matched and never gotten deleted.
Wired into setup.sh as a new --remove flag, resolving SERVICE_ALIAS
and validating the name the same way run_service already does.
Scoped to the common case (a Docker service at $DOCKER_DIR/<name> with
a standard configure_caddy_for_service site block); a hand-built Caddy
block or non-standard layout may need manual cleanup for the parts
this can't find. Non-Docker services (base, ssh-key-import, etc.)
report cleanly that they're not handled rather than erroring
confusingly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
ad011b2db8 |
Add check_container_health helper, wired into mattermost.sh as reference
A service's "Started" message after docker compose up -d doesn't mean the app is actually working — it can still crash-loop (bad DB password, missing required env var, etc.) with no visible sign until someone separately runs docker ps -a much later, exactly what happened repeatedly this session (mattermost, koha-db, homebox, vaultwarden, filebrowser all showed a clean "Started" message while crash-looping). check_container_health (lib/common.sh) waits briefly, checks the container's actual status and restart count via docker inspect, and prints recent logs automatically if it's not running or has already restarted — instead of a misleading one-line success message. Wired into mattermost.sh's own start step as the reference implementation, guarded by declare -F so standalone runs (no lib/common.sh sourced) degrade gracefully. Not retrofitted across every other service in one pass — this establishes the shared helper so other services can adopt it incrementally. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
74628f3736 |
Fix vaultwarden SMTP false-positive and mattermost DB password mismatch
vaultwarden: SMTP_PORT defaulted to "587" and SMTP_SECURITY was a
hardcoded "starttls" literal in the .env template, written
unconditionally regardless of whether SMTP_HOST was ever provided.
Confirmed live: skipping SMTP entirely (blank SMTP_HOST) still wrote
real values for those two, and Vaultwarden reads that as "some SMTP
config is present," refusing to start ("Both SMTP_HOST and SMTP_FROM
need to be set") even with host/from genuinely blank. Both now stay
empty unless SMTP_HOST is actually set.
mattermost: DB_PASS/MM_SECRET were only reused from the existing .env
when MODE=update — a "fresh" reinstall always generated a new
POSTGRES_PASSWORD. Confirmed live: choosing fresh after removing only
the mattermost app container (not the whole directory) regenerates the
password in .env while db/'s existing Postgres data still enforces the
OLD one from its first init (the entrypoint skips re-init on existing
data), causing "password authentication failed for user mattermost" on
every start. Whether db/ already has real data is what actually
determines whether the old password is still live, not which reinstall
mode was chosen — reuse the existing secrets whenever db/ is non-empty,
regardless of MODE.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
c177312947 |
Fix three crash-looping services: koha-db, homebox, vaultwarden
koha-db: compose used MYSQL_ROOT_PASSWORD/MYSQL_DATABASE/MYSQL_USER/
MYSQL_PASSWORD, but this mariadb:11 image version's entrypoint doesn't
recognize MYSQL_ROOT_PASSWORD as any of its accepted root-password
options at all. Confirmed live: "Database is uninitialized and password
option is not specified" on every start, even though DB_ROOT_PASS was
correctly generated and present in .env the whole time. Switched all
four to their MARIADB_* equivalents.
homebox: a newer homebox release requires HBOX_AUTH_API_KEY_PEPPER (at
least 32 bytes) or the container panics on startup — this installer
never set it. Generate one with generate_password 48 and wire it
through .env + the compose environment block.
vaultwarden: the SMTP setup prompts let you enter a host but leave
"SMTP from address" blank (no default), writing a half-configured state
Vaultwarden refuses to start with ("Both SMTP_HOST and SMTP_FROM need
to be set"). Validate after prompting — if SMTP_HOST is set but
SMTP_FROM came back empty, disable SMTP entirely instead of writing a
config known to crash the container.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|
||
|
|
77440c8f75 |
Live-scan Traccar's device-protocol range for collisions, not just Asterisk's ports
The 5000-5150 range was only ever checked against Asterisk's hardcoded fixed ports (5038/5060/5061) for a first instance — no live scan of the rest of the range, because the directory-count-based offset mechanism only triggers for an explicit additional instance. Confirmed live: this range sat unclaimed at the OS level while this Traccar instance's container had never actually started, so an unrelated service's own find_free_port scan found port 5007 genuinely free (nothing was listening there yet) and took it — invisible to any check until Traccar itself tried to bind its declared range for the first time, failing with "port is already allocated". Add a live scan across the whole intended range (skipping Asterisk's expected carve-outs at the base 5000-5150 range) and shift by 1000, same step the multi-instance path already uses, until genuinely clear. Also fixed the compose-block and README generation, which keyed off INSTANCE_SUFFIX being empty to decide whether Asterisk's exclusions were needed — now keyed off whether the range is still the unshifted default (PROTO_MIN -eq 5000), since a first instance can now end up shifted too. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
8bdf4c0a07 |
Add explicit UFW rules for Caddy's 80/443 — Docker was silently bypassing UFW
services/caddy.sh never called ufw allow for any of its published ports. Confirmed live: ufw status showed no rule for 80 or 443 on a box with UFW active (default deny incoming), yet HTTPS sites were reachable fine — Docker manipulates iptables directly for published container ports (the ports: mapping in Caddy's own compose file), which bypasses UFW's filtering entirely regardless of what ufw status reports. This wasn't an actual exposure gap — 80/443 are supposed to be open to everyone, that's the point of a reverse proxy — but it means ufw status was actively misrepresenting this box's real firewall state on its two most externally-facing ports, which is exactly the kind of thing that looks like a problem (and did, when investigating an unrelated Let's-Encrypt failure) even though nothing was actually unprotected. Add explicit ufw allow rules for 80/tcp, 443/tcp, and 443/udp (HTTP/3) so ufw status reflects reality, matching every other service in this repo managing its own firewall rules instead of relying on undocumented Docker/iptables interaction. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn |
||
|
|
060f288107 |
Fix authelia.sh skipping the real Caddy snippet due to a commented example
grep -q "(authelia)" matches caddy.sh's starter Caddyfile's own commented-
out example block ("# (authelia) {", included as documentation), so
authelia.sh believed the real snippet already existed and never wrote
it. Any later service adding `import authelia` to its own site block
then references a snippet that only exists as a comment.
Confirmed live: this takes Caddy down completely, not just the
Authelia-protected site — "Error: adapting config using caddyfile:
File to import not found: authelia" is a load-time failure, so Caddy
restart-loops and every site it fronts goes with it.
Anchor the check to an actual uncommented snippet definition
(^\(authelia\)\s*\{) instead of a bare substring match.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
|