New installs now name the container "asterisk", matching every other
service's container_name == service name convention, instead of
reusing the vendored easy-asterisk CLI tool's own name (which stays
/usr/local/bin/easy-asterisk inside the container, unrelated and
unchanged).
An existing "easy-asterisk" install is never silently renamed: every
place that resolves the container name (_asterisk_resolve_layout in
asterisk.sh, plus the duplicated copies in security-dashboard.sh,
sms-inbound.sh, pstn-trunk.sh, and tools/pstn-test-check.sh's docker ps
detection) now reads it from the box's own docker-compose.yml instead
of assuming it, falling back to "asterisk" only when there's no
existing install to read. Migrating a live box to the new name is a
one-time manual action (edit docker-compose.yml's container_name for
Asterisk and its coturn sidecar, docker compose down + up -d); every
sibling service then picks it up automatically on its next run.
The DigitalOcean-droplet layout (asterisk-digital-ocean directory,
easy-asterisk-do container) is untouched by this - that naming stays
exactly as documented for pre-merge droplet installs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
keep_alive_interval is a type=global pjsip.conf option, not a
type=transport option -- it never existed on [transport-tls] on any
Asterisk version. The IONOS TLS-keepalive mitigation was inserting it
there, which made sorcery reject the whole transport-tls object
("Could not find option suitable for category 'transport-tls' named
'keep_alive_interval'"), silently killing TLS SIP entirely instead of
just adding a keepalive.
Both _asterisk_patch_keepalive_vendor_files (deployed vendor copies)
and _asterisk_ensure_live_keepalive (live pjsip.conf) now target
[global]/type=global, and both self-heal a box that already picked up
the bad placement by removing it from [transport-tls] first.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
The stack health check's Security Dashboard/sms-inbound/ntfy checks only
had visibility into this box's own local Caddyfile — but all three can
legitimately be fronted by a Caddy (and Authelia) on a completely
different box instead, the same remote-Caddy pattern sms-inbound.sh and
ntfy.sh's own installers already support via CADDY_MODE/CADDY_REMOTE_HOST.
A site explicitly configured for remote Caddy was getting a false "Caddy
has no site block for it" for each of them, with a fix offer that would
have been actively wrong: adding a redundant local Caddy block for
something deliberately fronted elsewhere.
Now resolves the same site-wide CADDY_MODE the affected services'
installers themselves use before treating "not found locally" as a real
issue — only counts it, and only offers a fix, when the site is actually
in local Caddy mode. Remote (or no-Caddy) mode gets a plain informational
line instead: not wrong, just not something this box can verify.
Verified: default/local mode still flags a genuinely unwired dashboard as
an issue with a fix prompt; CADDY_MODE=remote (even with a local Caddy
directory also present) correctly downgrades the same finding to
informational with no prompt and no issue counted.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
"update" mode's whole promise is leaving already-configured things
alone — but that assumption silently breaks when something was
configured but never fully wired up, and update never re-asks the
questions that would reveal it. This session hit three separate
instances of exactly that on one droplet revert: a domain set with no
Caddy block, a Caddy block with no synced TLS cert (transport-tls fails
to bind — "Unable to retrieve PJSIP transport 'transport-tls'", breaking
every call), and a baked-in external IP left over from before the box
moved. sms-inbound and (potentially) Security Dashboard/ntfy can have
the identical "domain set, nothing serving it" gap with no way to
discover it either, since their own update modes don't re-ask.
_asterisk_run_stack_health_check(), called every "update", replaces the
narrower Caddy-only check added last time:
- Compares pjsip.conf's baked external_signaling_address against this
box's actual current public IP; offers to rewrite it and restart.
- Checks Asterisk's own DOMAIN_NAME has both a Caddy site block and a
matching TLS cert in the container; offers to fix each independently.
- Checks Security Dashboard / sms-inbound / ntfy (whichever are
installed) for a matching Caddy site block, via a new lib/common.sh
helper (caddy_domain_for_upstream) that finds the block without
needing to already know the domain — none of these three services
persist it anywhere. Points at that service's own "Full reinstall"
(the only mode that re-asks) since fixing their config isn't this
file's to script.
The cert-sync fix needed a non-interactive hook into the vendored
easy-asterisk CLI, which only exposed it as an interactive menu item
(Server Settings -> Force re-sync Caddy certs). Added a --sync-caddy-cert
flag via _asterisk_patch_cert_sync_cli(), patching the deployed vendor
copy the same way _asterisk_patch_voicemail_vendor_files and friends
already do — never vendor/ in git.
Every check runs unconditionally (never opt-in, so a gap is never missed
by nobody thinking to ask); every fix is individually opt-in and named
as a real config change, unlike the rest of "update"'s no-side-effects
default.
Also factored the DO-metadata/ifconfig.me/hostname-I public-IP detection
chain (previously duplicated 3 times) into _asterisk_current_public_ip().
Verified: caddy_domain_for_upstream against a multi-block Caddyfile
(distinguishes same-prefix upstreams correctly); the full health check
against fake docker/curl across every combination (all wired, IP
mismatch declined/accepted, cert mismatch declined/accepted, dashboard
unwired, sms-inbound wired vs. placeholder-domain, multi-instance ntfy
with one wired and one not); and _asterisk_patch_cert_sync_cli's
idempotency + resulting syntax against a real copy of the vendor script.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
"update" mode never re-asks the domain/networking/Caddy questions, on the
assumption there's already Caddy/Authelia config in place to leave alone.
That assumption breaks for an install where a domain was set at some
point (DOMAIN_NAME in .env) but Caddy never actually got a site block for
it — declined at install time, DNS wasn't ready yet, or Caddy was
reinstalled/reset separately since. Previously the only way back was a
full reinstall, which re-generates a dedicated coturn container with new
TURN credentials (every already-configured phone needs its QR re-scanned)
— a lot of blast radius just to add one missing Caddy block, and enough
that reaching for it risks the extensions/voicemail data a "fresh"
reinstall can also wipe if the wrong prompt is answered.
"update" mode now detects this specific gap (domain set, no matching
Caddyfile block) and offers to run _asterisk_configure_caddy_public()
right there — the same function "fresh" installs use, but it only ever
touches the Caddyfile and .env's WEB_ADMIN_AUTH_DISABLED line, never
coturn/extensions/anything else "update" already promises not to touch.
Verified in isolation: offers and calls the fix when the domain is set
with no matching Caddyfile block, stays silent when a block already
exists.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Confirmed live: neither this repo nor the vendored easy-asterisk script
ever sets a PJSIP endpoint's `mailboxes=` field. add_device()'s own
device_config template never writes it, and write_voicemail() only ever
touched voicemail.conf — so recording a voicemail worked fine
(voicemail.conf + the dialplan's VoiceMail() call), but no phone ever
actually subscribed to be told about it, regardless of whether the
voicemail flag was on. Matches the exact symptom of "voicemail records
fine, but no notice comes up on the phone."
Add _ea_set_endpoint_mailboxes(), called from write_voicemail(): adds/
updates mailboxes=<ext>@default in that extension's PJSIP endpoint stanza
when voicemail is enabled, removes it when disabled, and reloads
res_pjsip so it takes effect immediately. Bounded to just the
type=endpoint stanza (pjsip.conf reuses the same [ext] bracket name for
type=endpoint/type=auth/type=aor) the same way lib/common.sh's
_remove_caddy_site_block is bounded for Caddy blocks — verified against a
two-device pjsip.conf that editing one extension's mailboxes= never
touches its own auth/aor stanzas or another extension's stanzas, that a
repeat enable doesn't duplicate the line, and that disabling removes it
cleanly.
Existing extensions with voicemail already enabled won't get this
retroactively — the Extensions tab's voicemail toggle has to actually run
again (off then back on) to apply it, since this only fires on the
enabled/disabled transition itself.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
When the "Public domain for the webhook" prompt was left blank (DNS not
ready yet, or just missed), the installer built the Forward-to-URL as
literal https://<your-domain>/sms/... and persisted that placeholder to
settings.env as if it were real. "Update" mode never re-prompts for the
domain (by design — it's meant to leave already-configured settings
alone), so every later re-run silently re-served the same unusable
placeholder, with nothing indicating anything was wrong. A DID provider
(Anveo) correctly rejects it — it isn't a resolvable hostname.
- Only build FORWARD_URL when a real domain was entered; leave it empty
otherwise instead of substituting the placeholder.
- Fresh-install summary and README now say plainly that setup isn't
complete and how to finish it, instead of printing an empty/bogus URL.
- Update-mode now detects a missing/placeholder domain and tells you to
re-run with "Full reinstall" to be asked again, instead of reporting
success with a broken URL.
Verified with a direct test of _sms_write_readme() and the FORWARD_URL
construction for both the blank- and real-domain cases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Follow-up to Frigate's Authelia integration: both of these can also skip
their own login entirely once Authelia is doing the gating, each with a
different trust model appropriate to what the app actually supports.
- gitea: new _gitea_offer_reverse_proxy_auth(), a second Authelia
integration alongside the existing OIDC "Sign in with Authelia" button.
Enables Gitea's own ENABLE_REVERSE_PROXY_AUTHENTICATION so it auto-logs
in from a trusted Remote-User header — no click, no separate Gitea
session to expire on its own. Trust is IP-range based
(REVERSE_PROXY_TRUSTED_PROXIES), computed from caddy_net's real subnet
the same way ufw_allow_from_caddy_net does; refuses to enable the
feature at all if that can't be determined rather than fall back to a
permissive default — Gitea's own Docker image has shipped an unscoped
default before (GHSA-f75j-4cw6-rmx4, any IP could impersonate any user).
Rewires Gitea onto caddy_net and re-points Caddy at gitea:3000, since it
previously only reached Caddy via its published host port. Gitea's own
login stays available as a fallback, so unlike Frigate there's no
"native login off with nothing gating it" state to guard against.
- uptimekuma: sets DISABLE_AUTH=true only once Caddy's "import authelia"
gate is confirmed in front of it. Uptime Kuma already joined caddy_net
unconditionally, so this only needed the env var plus moving the
Authelia-gated Caddy call earlier (before docker-compose.yml is
written); the existing unconditional call at the end now only runs as a
fallback when the Authelia path wasn't used or wasn't completed. Kuma's
DISABLE_AUTH has no IP-scoping or secret check left once set — the
strictest of the three to get the ordering right on, since a mistake
here means wide open, not just spoofable.
Verified with a local test harness (fake Authelia/Caddy/docker-network
state): both the happy path and the "Caddy declined" safety fallback
produce the expected docker-compose.yml/.env/Caddyfile output for each
service, and Gitea's subnet-detection refusal + idempotent-rerun guard
were exercised directly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Frigate has its own built-in login separate from Authelia's session, so
just adding `import authelia` in front of it (the pattern used for
no-built-in-auth services) would leave two independent logins stacked,
defeating the point of Authelia's "remember me" on mobile. Frigate has a
`proxy` auth mode built for exactly this — trust Remote-User/Remote-Groups
from an upstream forward_auth proxy and disable its own login entirely.
- Extend configure_caddy_for_service() with an optional 5th arg for
sub-directives inside the reverse_proxy block itself (header_up), needed
to pin an X-Proxy-Secret header so Frigate's proxy-auth trust can't be
spoofed by a request reaching its published port directly, bypassing
Caddy/Authelia. Backward compatible — every other caller is unaffected.
- services/frigate.sh: prompt to protect with Authelia when installed;
wires import authelia + the X-Proxy-Secret header_up into Caddy, and
only writes config.yml's auth.enabled: False + proxy block once Caddy
actually confirms it's fronting the domain (never disables the native
login with nothing else gating access). Reuses the secret across
reinstalls instead of rotating it. Calls _authelia_scope_access() so
access can be restricted to specific users instead of every Authelia
account. Fixed a latent bug in the standalone-mode Caddy stub where the
auth block was placed after reverse_proxy instead of before it (dead
code — the same "Authelia never prompts" bug class CLAUDE.md documents
for the real helper).
- CLAUDE.md: document the new configure_caddy_for_service parameter and
Frigate's hybrid built-in-auth/forward_auth pattern.
Verified end-to-end against a local test harness (fake Authelia/Caddy
dirs): config.yml, .env, and the generated Caddyfile block all agree on
the shared secret and header names, auth is skipped cleanly when Caddy
isn't configured, and the secret is reused on a second run.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
User feedback: wanted chromaDB/rag-server/mcp-server included in the
optional-services picker added last commit, not just Gitea/Portainer/
Kiwix/InvokeAI/ComfyUI/Aider. Confirmed against the compose file's
depends_on chain before adding: open-webui only depends_on ollama (its
OLLAMA_BASE_URL connection works standalone), so none of these three are
actually required for regular chat — only Open WebUI's separate RAG tab
(routed through rag-server) and MCP tool-calling need them. Mealie's own
Ollama usage never touches this stack at all.
Bundled chromadb+rag-server+mcp-server as one option (7), not three
separate numbers — mcp-server depends_on rag-server depends_on chromadb,
so stopping only one of the three would leave the others running against
a dead dependency instead of a clean stop. Also added a cascade for the
existing kiwix option: mcp-server depends_on kiwix too (not just
rag-server), so stopping kiwix without also stopping mcp-server has the
same problem — now handled automatically with a dedup pass in case both
the kiwix cascade and option 7 add mcp-server to the stop list.
Verified all four cases in isolation: kiwix-only correctly cascades to
mcp-server, option 7 alone stops the right three, choosing both dedupes
to one clean list, and unrelated choices (gitea/portainer) are unaffected.
User feedback: local-ai-setup.sh always brings up the entire stack
unconditionally (Gitea, Portainer, Kiwix, InvokeAI, ComfyUI, Aider,
alongside the core Ollama/Open WebUI/ChromaDB/RAG/MCP) with no way to opt
out — e.g. Gitea when you already run git elsewhere, or Portainer when you
manage Docker some other way.
Didn't touch local-ai-setup.sh's own compose generation for this (it's
vendored upstream code, and other services reference these by container
name/network in ways that would need individual auditing to make safely
conditional). Instead, added a post-install picker in the wrapper: after
the full stack starts, offer to `docker compose stop` whichever of the six
non-core services aren't wanted. Images are already pulled either way, so
anything stopped comes back later with a plain `docker compose up -d
<name>` — no reinstall needed.
Verified the choice-parsing loop in isolation: "1 4 9 3" correctly warns
on the invalid "9" and resolves to gitea/invokeai/kiwix.
Confirmed live: a bare hostname in DOMAIN= (typing "vault.example.com"
instead of "https://vault.example.com" at the install prompt — easy to do
despite the example text showing the scheme) crash-loops the container
with no clear startup error, and re-running the installer doesn't fix an
already-written .env since "update" mode deliberately never touches it.
Two changes, mirroring how the existing SMTP half-state bug is already
handled in this file:
- Normalize VW_DOMAIN at prompt time — missing scheme gets https://
prefixed automatically instead of writing it verbatim.
- New _vaultwarden_fix_domain_scheme() self-heal, called at the same two
sites as _vaultwarden_fix_smtp_halfstate() (the "update" path and the
fresh-install "start now" path), so a box that already has a scheme-less
DOMAIN self-heals on its next start instead of staying stuck.
Verified the self-heal function in isolation: vault.mydomain.com ->
https://vault.mydomain.com.
local-ai-setup.sh runs as whoever invoked this wrapper — root, since
setup.sh itself runs under sudo — so every file it generates
(docker-compose.yml, .env, requirements.txt, server.py, mcp_server.py,
pull-models.sh, start/stop/status.sh) came out root-owned. Nothing handed
that back to ACTUAL_USER unconditionally: the only existing
ensure_docker_dir_ownership call was inside the cloud-provider wiring
block, so it silently never ran at all for anyone who skipped cloud
providers.
Confirmed live: this repo's own "Skipped. Run later: cd $AS_DIR && bash
local-ai-setup.sh" message tells the user to re-run it directly later as
themselves (no sudo) — which then fails with "Permission denied" on any
file root created during the original sudo run, e.g. requirements.txt.
Same root cause class as a stray root-owned .git/FETCH_HEAD blocking a
plain `git pull` — a root-run leaving files a later unprivileged run can't
touch.
Fix: call ensure_docker_dir_ownership "$AS_DIR" unconditionally right
after the installer-run block, not only on the cloud-provider path.
Confirmed live: `docker compose pull` failed with "yaml: line 44, column
29: mapping values are not allowed in this context" during the "Starting
Stack" phase of local-ai-setup.sh. Root cause: two healthcheck blocks
(ollama, chromadb) crammed interval/timeout/retries onto one
semicolon-separated line —
interval: 30s; timeout: 10s; retries: 5
— which isn't valid YAML; a scalar value can't contain a second `key:`
token like that unless quoted. Split each into three separate properly
indented keys, matching how every other multi-key block in this same file
is written.
Verified by generating the actual docker-compose.yml via the real heredoc
(same one docker-stack.md's variables would produce) and parsing the
result with PyYAML — line 44 is exactly the fixed `interval: 30s` line,
and the full file now parses as valid YAML.
Pre-existing bug in the vendored source, unrelated to this session's
earlier ai-stack.sh/local-ai-setup.sh changes (those only touched the
pull-models.sh heredoc and the cloud-provider prompt, both well before
this point in the install) — first surfaced now because this is the first
run in this session to actually reach the "Starting Stack" step rather
than stopping earlier.
Blank already meant skip, but user feedback wanted a keystroke that says
so explicitly rather than just leaving the input empty. Added "0) Skip —
stay fully local" to the menu, updated the prompt to mention it, and
handled "0" as a silent no-op in the choice loop (previously it would
have fallen through to the "Ignoring unknown choice" warning).
The instruction was only in explanatory text a few lines above the actual
prompt (prompt_text "Cloud providers to add []:") — easy to miss once
that's scrolled past, especially since the bracketed default shows empty
but doesn't say what empty means. User feedback: the screen itself should
say it, not just text above it. Now reads "Cloud providers to add (blank =
skip, stay fully local):".
None of local-ai-setup.sh's tier-selected models (CHAT_MODEL/CODE_MODEL/
EMBED_MODEL) can read an image — there was no way to get vision support out
of this stack at all before now. Added a numbered pick-list to the
generated pull-models.sh, right after the existing DeepSeek-R1 optional
pull, matching that same read -rp pattern:
1) moondream ~1.7 GB by Moondream AI — tiny, built for
CPU-only or weak/old-GPU hardware
2) llava:7b ~4.7 GB general-purpose vision
3) qwen2.5vl:7b ~6 GB stronger accuracy, more RAM/VRAM
4) llama3.2-vision:11b ~7.9 GB heaviest of the four
moondream is the recommended default — sized for exactly the "6 vCPU, 8GB
RAM, no GPU" case this was asked for, unlike the other three which assume
real GPU/RAM headroom.
Verified by actually running the heredoc that generates pull-models.sh
(with EMBED_MODEL/CHAT_MODEL/CODE_MODEL stood in) and syntax-checking the
resulting output script, not just the source — the outer heredoc is
unquoted so $-escaping mistakes wouldn't show up as a bash -n failure on
local-ai-setup.sh itself, only on what it generates.
services/ai-stack.md gets a matching "Vision models" section (sizes, the
manual pull command, and how to point an app's OPENAI_MODEL at one).
laptop_full_setup.sh's separate, non-interactive pull-models.sh generator
is untouched — it's not invoked anywhere in this repo's own install flow
(only local-ai-setup.sh is, from install_ai-stack()), so it's out of
scope here.
Confirmed live: install_mealie() pre-computes BASE_URL as
recipes<suffix>.$SITE_DOMAIN before ever asking about Caddy, then
configure_caddy_for_service() separately prompts for a domain — which the
user can freely override (e.g. typing mealie.mydomain.com instead of
accepting the recipes.mydomain.com default). Nothing fed that choice back
into BASE_URL, so it stayed stale. Since BASE_URL is exactly what
_mealie_offer_authelia_oidc() registers as the OIDC redirect URI, this
produced Authelia's "redirect_uri does not match any of the OAuth 2.0
Client's pre-registered redirect_uris" — Caddy and DNS were both correctly
pointed at the new domain, but the client Authelia had on file still said
the old one.
Added CADDY_SERVICE_DOMAIN as a new configure_caddy_for_service() out-param
(lib/common.sh) — the same out-param convention as the existing
CADDY_SERVICE_CONFIGURED/CADDY_SERVICE_MODE, set right after the domain
prompt is accepted. install_mealie() now reconciles BASE_URL against it
immediately after the Caddy call, before the Authelia OIDC step reads
BASE_URL back out of .env. ActualBudget's equivalent OIDC offer asks for
its own domain fresh each time rather than reading a pre-computed BASE_URL,
so it isn't affected by this class of bug and needs no equivalent fix.
Verified the reconciliation logic in isolation against a synthetic .env.
Root cause of the recurring Mealie OIDC "unexpected character '/' in
variable name" failure, confirmed against the user's actual
configuration.yml byte content: this repo's own scripts write
authelia_url/domain unquoted, but YAML makes quoting optional, and a
hand-edited config can add single or double quotes around the value
(here: authelia_url: 'https://authelia.example.com.'). awk's
`print $2`/`print $3` is a naive whitespace-split token grab that doesn't
know about YAML quoting, so it captured the value WITH the literal quote
characters attached. The generated discovery URL then came out
`'https://authelia.example.com.'/.well-known/openid-configuration` —
Docker Compose's env parser closed the quoted value at that embedded
closing quote and choked on the trailing text as an invalid new token.
The earlier \r-stripping commit was a real but different fix (a
CRLF-tainted line fails to match these anchored awk patterns at all) —
it didn't cause and couldn't have fixed this. Both guards are needed and
now both apply, in both _authelia_provision_oidc_client() (domain and
portal URL) and the same latent bug in _authelia_add_oidc_client()'s
domain parse.
Verified end-to-end: reconstructed the user's exact reported byte
content (od -c dump) in a synthetic configuration.yml, ran the actual
_authelia_ensure_oidc_provider/_authelia_provision_oidc_client/
_mealie_offer_authelia_oidc functions against it (docker calls stubbed),
and confirmed the generated .env line is now a single clean line with no
embedded quotes or split.
The previous commit added OIDC_AUTHELIA_PORTAL_URL parsing but only
tr -d '\r'-sanitized the awk output, not the input. That's insufficient
for a CRLF-tainted file (confirmed live: a configuration.yml line
hand-edited by something that saves Windows line endings) — every
line-anchored awk pattern here fails to match at all against a line like
" cookies:\r", since $ anchors end-of-string and the \r is still part of
it, not just leaves a stray \r in the captured value. Symptom was Mealie's
generated OIDC_CONFIGURATION_URL line getting split mid-string, which
Docker Compose's env parser (bare \r treated as a line break too) reported
as "unexpected character '/' in variable name".
Fixed by piping the file through tr -d '\r' before awk sees it, for both
the domain and portal-URL parses. Verified against a synthetic CRLF config
that reproduces the exact failure — both fields now parse clean.
_authelia_provision_oidc_client() gained a new out-param,
OIDC_AUTHELIA_PORTAL_URL, read back from the instance's own
configuration.yml (session.cookies[].authelia_url) — the actual source of
truth for where the portal lives — instead of every caller separately
assuming "https://auth.$domain".
install_authelia() and add_authelia_domain() both still default new
instances to "auth." as before (unchanged), but that's just a default, not
a guarantee: it's plain text in configuration.yml and gets hand-edited on
some boxes (e.g. a dedicated instance renamed to "authelia." to avoid
colliding with another instance's "auth." on a different machine). Mealie,
ActualBudget, and Gitea's native-OIDC wiring all independently hardcoded
"auth." when building their discovery URL, so a renamed portal silently
produced a discovery URL pointing at a host that doesn't serve Authelia —
surfacing as an opaque 500 during the OIDC token exchange with no useful
client-side error.
Verified the new awk parse against both a default ("auth.") and a renamed
("authelia.") cookies block before trusting it.
edit_authelia_user() previously only ever let you select one user, act on
them, and then returned all the way out of install_authelia() (which calls
it with an immediate `return 0`) — deleting several users meant re-running
`sudo ./setup.sh authelia` and re-navigating to option 3 from scratch for
every single one.
Restructured: the per-user action menu (edit/reset-password/2FA/admin/
service-access/delete) is now _authelia_manage_one_user(), and
edit_authelia_user() drives it in a loop — numbered multi-select up front
("2 4" deletes/edits both), then "Manage more users?" to go again with a
freshly re-read user list instead of exiting. Guards against acting on a
user who was already deleted earlier in the same batch.
Verified against a synthetic users.yml: selecting two users by number and
deleting both in one pass removes exactly those two, leaves the others
untouched.
tr -cs 'a-z0-9_-' '-' only allowed lowercase letters, so any uppercase
leading character (e.g. "Bob") got converted to a dash and then stripped
by the paired leading-dash sed, silently truncating the username. Widened
to a-zA-Z0-9_- in both add_authelia_user() and _authelia_scope_access().
Also extends the "Manage an existing user" menu (still numbered-selection
throughout) with:
- option 6: toggle a user's membership in any existing "<service>-only"
scoped-access group, picked by number, via two new helpers
(_authelia_list_scoped_groups, and re-resolving the user's line range
before each toggle since a prior toggle in the same pass shifts it)
- option 7: delete a user outright (_authelia_delete_user_block), with confirmation
Verified against a synthetic users.yml (add/remove toggling across
multiple groups, block deletion, uppercase-username round-trip) before
touching the live file.
Confirmed live: re-running ActualBudget's Update path produced zero
output for the Authelia SSO step — no prompt, no message, straight back
to the shell. Root cause: the idempotency guards in
_actualbudget_offer_authelia_oidc / _mealie_offer_authelia_oidc /
_gitea_offer_actions_runner were plain `grep -q ... && return 0` — silent
by construction. Indistinguishable from the step not running at all,
which is exactly what it looked like.
ActualBudget and Mealie's OIDC offers now explain what they found and
ask whether to reconfigure (registers a fresh Authelia client + secret,
clearing the old env vars first) instead of silently bailing. Gitea's
Actions-runner offer explains what it found and how to check its status
(reconfiguring that one means editing a docker-compose service block,
not just a couple of env vars, so it just informs rather than offering
to redo it).
Also: _authelia_scope_access now shows existing Authelia usernames as a
numbered list before asking who should have access — picking by number
works alongside typing new names directly (mix freely, e.g. "1 3
newperson"), rather than requiring exact usernames typed from memory
with no reference and no protection against a typo silently creating a
duplicate account.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
_authelia_scope_access asked for usernames to grant access to a service
without ever showing who already exists — confirmed live, the prompt
just showed a blank "Usernames:" line with nothing to reference. A typo
against an existing name doesn't fail or warn, it silently creates a new,
separate account instead of matching the intended one.
Now lists existing Authelia users (reusing _authelia_list_usernames,
already used elsewhere in this file) right before the prompt, and warns
about the typo/duplicate-account risk explicitly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Fixes two things found while answering a question about staying logged
in across every Authelia-protected service:
1. CLAUDE.md's own "stay logged in" instructions referenced
remember_me_duration — renamed to remember_me in Authelia 4.38, this
repo pins 4.39.20. Authelia doesn't error on an unknown key, it just
silently ignores it, so following that guidance as written would have
done nothing. install_authelia() itself already uses the correct
`remember_me` key at install time (default 7d) and was never affected
— only the hand-edit instructions in the docs were stale.
2. There was no way to change it afterward without hand-editing the file,
contrary to this repo's own "no manual config editing" direction.
Added _authelia_set_remember_me() (new menu option 7): prompts for a
new duration (12h/7d/1M/1y/-1 to disable), writes it, restarts.
Tested the sed replacement against a synthetic session block before
trusting it on real config. Also documented clearly (both in the
function's own prompt and in CLAUDE.md) that this only controls
Authelia's own session — a native-OIDC app's own session/token expires
on its own separate schedule, which this setting doesn't touch.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Confirmed live: ActualBudget's new automated "Sign in with Authelia" offer
hit a client_id ("actualbudget") already registered from an earlier use of
the interactive "Register an app" menu — that older flow only registers
the client in Authelia and prints instructions to paste the secret into
the app's own settings manually; if that paste step never happened,
ActualBudget's .env never got the OIDC vars, but Authelia still considered
the client_id taken. _authelia_provision_oidc_client's duplicate check
just warned and returned 1, permanently blocking the automated offer with
no path forward — the stale registration's secret was shown once and
already gone, so there was nothing to recover, only reasons to replace it.
Added _authelia_remove_oidc_client() (tested against a synthetic
multi-client config, both mid-list and last-in-list removal) and changed
the duplicate-client_id check to remove-and-replace instead of failing.
Every automated caller (gitea/mealie/actualbudget's SSO offers) uses a
fixed, service-specific client_id, so a collision there means "this same
service was already registered," not a different app's ID being
clobbered. The interactive menu's own earlier duplicate check (a distinct
code path, one step before this one) is untouched — it still warns and
stops before prompting further, since a human-typed ID colliding with an
unrelated app is a different, more ambiguous situation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Researched which of the "has built-in auth" services actually support
native OIDC before wiring anything in (checked live docs, not assumed) —
two services turned out to contradict general assumption: Portainer's
OAuth/OIDC is Business Edition only (this repo installs portainer-ce, which
doesn't have it), and ntfy has no auth-oauth2-* support at all despite it
seeming like the kind of thing a modern self-hosted tool would have added
by now. Full findings recorded in CLAUDE.md so this doesn't need
re-researching.
Two real, verified wins wired up, both entirely env-var driven — no
manual config file editing, matching this repo's "no manual wizard"
philosophy and reusing the exact _authelia_provision_oidc_client /
_authelia_scope_access machinery already built for Gitea:
- mealie: OIDC_AUTH_ENABLED/OIDC_CLIENT_ID/OIDC_CLIENT_SECRET/
OIDC_CONFIGURATION_URL appended to the existing .env (env_file: .env is
already how mealie.sh's compose reads it). Also adds a
--forwarded-allow-ips entrypoint override when Caddy-fronted — confirmed
against Mealie's own issue tracker that without it, the generated OIDC
redirect URI comes out http:// even when actually served over https://,
which providers reject as a scheme mismatch.
- actualbudget: ACTUAL_OPENID_DISCOVERY_URL/CLIENT_ID/CLIENT_SECRET/
SERVER_HOSTNAME, same pattern. Redirect path (/openid/callback) matches
the preset already used by authelia.sh's own "Register an app" menu for
this same service.
Both offered on fresh installs and Update reruns, default no, and both
call _authelia_scope_access() afterward so access can be restricted to
specific users instead of every Authelia user, same as Gitea.
Immich has real OIDC + a system-config API but needs one more
verification pass on the exact request payload before automating — not
guessing that part. Jellyfin and Home Assistant only have third-party
plugin/HACS-based OIDC, a bigger lift than an env-var toggle — noted but
not attempted this pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Every domain with an access_control rule was reachable by any Authelia
user by default (the existing catch-all *.${AUTHELIA_DOMAIN} rule) — no
way to restrict a specific service to a subset of users without hand-
editing configuration.yml and users.yml directly.
_authelia_scope_access(SERVICE_ID, DOMAIN) is a new generic, reusable
helper: call it after any service finishes being protected by Authelia
(forward_auth gate or native OIDC alike — it only cares about the domain).
Offers universal vs. specific-users access; if scoped, creates a
"<service_id>-only" group, adds every listed username to it (creating
accounts on the fly for names that don't exist yet, via the new
_authelia_create_user_noninteractive — a non-interactive sibling to
add_authelia_user, same extraction pattern already used for
_authelia_provision_oidc_client), and inserts an allow+deny rule pair
above the general catch-all. Idempotent on rerun.
_authelia_report_access_scope() (new menu option 6) is the read side —
lists who has universal vs. service-scoped access, and offers to promote
a scoped user back to universal by removing their "-only" group
membership.
services/gitea.sh's _gitea_offer_authelia_sso() is the reference
integration, calling _authelia_scope_access after successfully wiring up
Gitea's OIDC login. The other services with a plain "Protect X with
Authelia?" prompt (magicmirror, wolf-pair, js99er, drum-rhythm-game,
iopaint, paintplus, stirling-pdf, wolf) are natural follow-ups once this
is confirmed working live — each just needs one added call.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Sipnetic clients on the user's home WiFi/VLAN lose SIP/TLS registration
every ~5s when Asterisk runs on IONOS, but never on DigitalOcean or over
mobile data — CrowdSec, OPNsense firewall/IDS, TURN-for-registration, and
raw packet loss have all been ruled out via live testing. Leading theory
is an idle-connection timeout inside IONOS's network virtualization layer.
keep_alive_interval sends a periodic double-CRLF over the TLS transport to
keep it from going idle, the standard mitigation for this failure class.
Follows the existing dual-patch pattern (vendor-template copy + live file)
since transport objects aren't picked up by `pjsip reload` and need a
container restart to apply, same as the live_dangerously fix.
Confirmed live: a repo's local bare mirror clone got stuck on a stale
commit indefinitely even though every sync run reported [ok] — fetch
never failed, it just wasn't updating refs/heads/* the way this script
assumes. GitHub had the real current commit; the bare clone (and
therefore what got pushed to Gitea) stayed frozen on an old one.
`fetch --all --prune` trusts remote.origin.fetch as stored in the bare
repo's own git config rather than asserting what that mapping actually
is — if it drifted from the +refs/heads/*:refs/heads/* convention a
fresh `git clone --bare` sets up (for whatever reason — this specific
repo's local clone directory's history is unclear), fetch would
"successfully" land new commits somewhere this script never reads
(refs/remotes/origin/*) while refs/heads/* — the ref that actually gets
mirrored — never moves. Deleting and re-cloning the affected repo's
local directory fixed it immediately, consistent with a refspec-drift
theory, though the exact original cause wasn't pinned down further.
Now pins `+refs/heads/*:refs/heads/*` explicitly on every fetch instead
of relying on `--all` plus whatever's configured, in both sync
directions. Also stopped redirecting stderr to /dev/null on every
clone/fetch/push call — a real auth or network failure now shows up in
the log instead of a bare "Failed to X" with no reason, which is what
made this bug take three rounds of manual ls-remote/rev-parse forensics
across two machines to actually pin down.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Gitea Actions is Gitea's own CI, largely GitHub-Actions-workflow-compatible
(.gitea/workflows/*.yml). Off by default; this Gitea install is otherwise
just a passive GitHub pull mirror, so the main value here is resilience —
.gitea/workflows/*.yml can still run something like a GitHub Actions build
if GitHub itself is ever unreachable.
_gitea_offer_actions_runner(), offered on fresh installs and Update reruns
(idempotent — no-ops if already set up):
- Enables GITEA__actions__ENABLED / DEFAULT_ACTIONS_URL in the compose
file's environment, restarts to apply
- Generates a runner registration token via `gitea actions
generate-runner-token`
- Appends an act_runner service to the same docker-compose.yml, using
the host's Docker socket to launch a fresh container per job — the
same pattern this repo already uses for portainer/watchtower/
uptimekuma/beszel/traccar's autoheal
- Falls back to printing manual setup instructions if token generation
fails, rather than losing the attempt silently
_gitea_fix_ownership()'s data/-exclusion (added when we fixed the earlier
SQLite readonly-database bug) now also skips runner-data/, so a future
reinstall/update doesn't clobber the runner's own state the same way.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Gitea has its own built-in login, so it was never wired into the
forward_auth/Caddy pattern the rest of this repo uses to gate apps with
no auth of their own — that's still correct and unchanged. But Gitea
also supports adding an OAuth2/OpenID Connect authentication source
natively, and Authelia can act as an OIDC provider — a genuinely
different, additive integration: an extra "Sign in with Authelia" button
on Gitea's own login page, alongside local login, not a Caddy-level gate.
Refactored services/authelia.sh's _authelia_add_oidc_client() to split
out its non-interactive core as _authelia_provision_oidc_client() — same
behavior for the existing ActualBudget/Vaultwarden/Immich/custom-app menu
flow, but now callable directly by other services with explicit args
instead of walking a human through the menu, returning the plaintext
secret and Authelia's domain via out-params.
services/gitea.sh's new _gitea_offer_authelia_sso() uses that to fully
automate both sides when accepted: registers Gitea as an OIDC client in
Authelia, then runs `gitea admin auth add-oauth` itself to add Authelia
as an authentication source — no manual web-UI copy-paste on either side,
matching how this installer already avoids manual wizards for the admin
account/token. Falls back to printing the values for a manual add if the
Gitea-side CLI call fails. Offered on fresh installs and on Update
reruns (default no, so a plain Update stays silent), so it can be added
later without a full reinstall.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Root cause of "attempt to write a readonly database (1544)" on repo
creation: Gitea's container always runs internally as UID 1000
(USER_UID/USER_GID are fixed in docker-compose.yml, independent of
whoever's running this installer) — the image chowns /data to that UID
itself at startup. install_gitea()'s three ensure_docker_dir_ownership
calls recursively chown the *entire* service directory, data/ included,
to $ACTUAL_USER. On a box where the installer runs as root directly
(ACTUAL_USER=root), that resets a live data/ back to UID 0. If the
container doesn't happen to restart right after — confirmed live: Update
mode against an already-running container just no-ops instead of
restarting — nothing ever re-fixes it, and every subsequent write to
Gitea's own SQLite DB fails.
Added _gitea_fix_ownership(), which chowns everything in the service
directory except data/, and swapped it in at all three call sites. The
container continues to own data/'s permissions exclusively, as it always
has on first boot; this installer no longer fights it on every rerun.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
((pull_count++)) evaluates to the PRE-increment value — 0 on the very
first successful pull/push — and under this script's `set -euo pipefail`,
an arithmetic command evaluating to 0 counts as a failing command and
kills the script immediately. Confirmed live: a real, fully successful
GitHub -> Gitea pull (visible in sync.log as "PULL ... OK") still made
the whole run exit non-zero and get reported as "Sync run failed", purely
because it was the first repo to sync (0 -> 1). Any subsequent repo in
the same run would have been fine, but most real installs only have a
handful of repos, so this could look like sync is just broken.
Switched all three counters (pull_count, push_count, fail_count) to
assignment form (`count=$((count + 1))`), which always exits 0 regardless
of the resulting value. page++ elsewhere in the file starts at 1, not 0,
so it isn't affected by this and was left as-is.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
"Failed to reach GitHub API. Check GITHUB_TOKEN." (and the equivalent
Gitea message) pointed at the token every time, even when the real cause
was something else entirely — confirmed live twice in one debugging
session: once a GitHub-side 503 outage, once a Gitea account locked
behind a must-change-password 403. Both times the fix was to run the
same curl by hand to see the actual status/response.
Fold that same probe into the script itself: on failure, re-request with
-i and print the HTTP status and response body directly, so the failure
mode (bad token vs. remote outage vs. account lock vs. network/DNS) is
visible immediately instead of requiring a manual curl round-trip.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Root cause of the "Failed to reach Gitea API" / 403 errors on every retry:
`gitea admin user change-password` (used in the already-exists branch to
sync the account's password to what the user just entered) defaults to
setting must_change_password=true, unlike `user create` which was already
pinned to --must-change-password=false. Once set, Gitea rejects every API
call — including the sync script's own token-authenticated calls — with
403 "You must change your password", even though the token itself and
GITEA_URL were both completely correct. Confirmed live via a direct curl
against /api/v1/user.
Pin the same flag on change-password that create already used.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Live symptom: pasting the GitHub PAT over SSH showed the token text
landing on the terminal *after* "No GitHub token entered" had already
printed — the prompt's read() returned empty a beat before the paste
actually arrived (a paste/Enter race that isn't specific to this box,
just common over higher-latency SSH sessions). A single empty answer
was treated as "user has no token" and the install moved on silently.
Both token prompts (GitHub token, and the Gitea-token manual fallback)
now retry up to 3 times interactively before giving up, and strip
whitespace from what was captured in case the paste carried a stray
leading/trailing newline. Unattended installs still take one shot, same
as before, since nobody's there to retry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
The sync direction step only ever set up the timer (or printed manual
instructions) — there was no way to actually confirm tokens/config work
without waiting for the first scheduled run or invoking the script by
hand afterward. Add a post-configure prompt: dry-run preview (--list),
run for real right now, or skip. Defaults to dry-run interactively;
defaults to skip under UNATTENDED so a headless install with no GitHub
token configured doesn't spam preflight errors.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
generate-access-token used a fixed --token-name "sync", which Gitea
rejects on a second call for the same user (e.g. a retry against an
already-existing admin account, now common after the readiness-wait
fix). The failure was silent: it fell through to a manually-labeled
"Paste the Gitea token here" prompt appearing immediately before the
real "GitHub token:" prompt, so a pasted GitHub PAT could land on the
wrong prompt and leave GITHUB_TOKEN empty with no clear reason why.
Token name now includes a timestamp so it's always unique, and the
fallback prompt is relabeled to make clear it wants a Gitea token, not
the GitHub one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Previously the admin account was silently auto-generated (username =
ACTUAL_USER, random password) and only created after a fixed 60s
readiness probe — a slow first boot (SQLite init on a slower disk) timed
the whole install out with no account ever created, leaving the user to
create one by hand with a raw docker exec.
Now the install prompts for admin username/password up front, then folds
account creation into the same retry loop used to detect readiness (up to
2 minutes), so a slow-but-eventually-successful boot no longer dead-ends
the install. A retry against a partially-completed prior run (account
already exists) is treated as success and syncs the password instead of
failing outright.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
CONTAINER_NAME was correctly re-detected from the current layout
(ASTERISK_KIND) at the top of install_pstn-trunk(), then immediately
clobbered by sourcing the saved .pstn-trunk.env settings file — and
_pstn_apply_settings writes CONTAINER_NAME straight back into that
same file every run, so a stale value self-perpetuates forever once
it's wrong.
Confirmed live: a box migrated from a DigitalOcean droplet (where the
settings file correctly saved CONTAINER_NAME=easy-asterisk-do) to a
plain/home install kept that droplet-era value indefinitely, since
only "update" runs happened after the migration. Every update
reloaded/restarted a container that no longer existed instead of the
one actually running Asterisk, so dialplan changes (including
pstn-trunk-inbound-dialplan.conf) never reached the live process even
though the config files themselves were written correctly.
Re-asserts the freshly-detected value after the source instead of
trusting whatever was saved, so it can't drift from the box's actual
current layout and self-heals the persisted file on the next update.
Gitea previously only existed bundled inside the full ai-stack service
(Ollama/ComfyUI/InvokeAI/etc. all together) — no way to get just a git
server without the rest of that heavy stack. This adds it as its own
lightweight service, reusing the vendored gitea-github-sync.sh but not
any of ai-stack's other components.
- Auto-creates a Gitea admin account and API token via the container's
own CLI (no manual web setup wizard).
- Asks GitHub token, sync direction (GitHub->Gitea / Gitea->GitHub /
both), and whether to install a systemd timer for automatic sync —
prints manual commands instead if declined.
- Own systemd unit for the timer rather than the vendor script's
built-in --install-timer, since that always runs both directions
with no way to pin a single direction.
pstn_personal_ring dialed the owner and hung up regardless of
DIALSTATUS, so no-answer/busy calls to a personal DID just dropped
silently instead of offering voicemail. Now falls to
VoiceMail(<owner>@default,u) on anything but ANSWER, gated by the
owner's existing voicemail=yes/no flag in pstn-permissions.conf.
The shared ring-group and group-owned personal DID inbound paths have
the same gap but no single owning extension to pick a mailbox for —
left as-is pending a decision on what that should do.
Adds live_dangerously = yes to asterisk.conf automatically on install/
update, fixing silent PSTN call denial (AST_CONFIG() returning empty
with no error when this option is off).
AST_CONFIG() silently returns an empty string instead of erroring when
asterisk.conf's [options] section lacks live_dangerously = yes, so a
correct tier_out=full in pstn-permissions.conf still evaluates as no
permission — every outbound/ring-group call gets denied with nothing
in the logs pointing at the real cause. Easy Asterisk's vendor default
ships without this set. Now applied automatically on fresh install and
on every "update" rebuild, restarting only when the file actually
changes.
Immich has full native OIDC support (its own docs list Authelia as a
supported provider), but needed more than the single-redirect-URI
model _authelia_add_oidc_client() previously supported: it requires
three redirect_uris at once (web login, account-linking page, and the
mobile app's app.immich:///oauth-callback custom-scheme redirect).
Generalized redirect-URI handling from a scalar REDIRECT_PATH to two
arrays (domain-relative REDIRECT_PATHS, plus already-complete
EXTRA_REDIRECT_URIS for non-domain-based ones like the mobile scheme)
and build the YAML redirect_uris list from however many are present.
ActualBudget/Vaultwarden/Other still resolve to a single-entry array,
so their generated config is unchanged. Verified the multi-entry YAML
generation against a python yaml parser before wiring it in, and the
case-statement/array logic in isolation against the real file's code.
Confirmed live (not from docs): Authelia's in-portal Settings -> Change
Password also emails a one-time code to confirm, same as Forgot
Password — it is not a no-SMTP path as earlier text here assumed.
Reworded all three spots in authelia.sh that claimed otherwise to
point at the admin-side "Edit an existing user" -> "Reset password"
action instead, which never touches email.
Also corrected CLAUDE.md's "No built-in auth — should be protected"
list per an actual grep of services/*.sh: it was missing
drum-rhythm-game, iopaint, paintplus, stirling-pdf, wolf, and the
unconditionally-protected security-dashboard/asterisk, and wrongly
included sky-cam (a non-Docker batch script with no web UI or Caddy
integration at all, nothing for Authelia to protect).
New menu option lists existing users by number; picking one opens a
submenu to edit email/display name, force a password reset, reset a
2FA device (authelia storage user totp delete), toggle a one_factor
exemption for that user via a subject-scoped access_control rule, and
promote/demote admin group membership.
All the YAML-editing helpers (line-range lookup, scoped field/group
edits, and the access_control rule insertion/removal used by the 2FA
exemption toggle) are line-range-scoped to the target user only, and
were verified against single- and multi-domain/multi-user fixtures
before wiring them into the interactive flow — a bad edit to
access_control here would break every protected domain, not just one
user's account.
30 characters, guaranteed at least 5 uppercase, 5 digits, and 5 special
characters, shuffled. Scoped to add_authelia_user() only via a small
local generator — deliberately not routed through lib/common.sh's
shared generate_password, since that one is alphanumeric-only by
design (its paired validate_password rejects special characters) and
plenty of other services embed its output unescaped into .env/YAML/URLs.
ActualBudget requires inviting additional OpenID users from its own
"Server Online" screen before their login is accepted, separate from
Authelia authenticating them successfully. Companion doc gets appended
to the generated README automatically (write_readme convention).
Adding a user previously required hand-editing users.yml and generating
the argon2 hash manually. New menu option (2) on an existing Authelia
install prompts for username/email/display name/admin group, generates
the hash and temp password, inserts the users.yml block, and restarts
Authelia — mirroring the existing add_authelia_domain/OIDC-client flows.
Asterisk DO->IONOS migration (resolved), web-based extension messaging
(not started), Pi-hole (done), VPS-as-VPN-endpoint with encrypted DNS
(idea stage). Temporary -- delete once these are finished or turned into
real issues/PRs.
Prompts on fresh/new installs and writes MM_SERVICESETTINGS_COLLAPSEDTHREADS
into .env; update reruns read the existing value back instead of
re-prompting, since it may have been changed later via System Console.
Documents the setting and how to flip it later in the generated README.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GKXTDG1ivYASN9fvDZ9Rfv
restore replaces every file under the Asterisk directory with a fresh
extraction from the archive (rm -rf + mv in the new tree) -- every file is
a new inode, so any POSIX ACL grants the Security Dashboard holds on them
(_secdash_grant_asterisk_access's setfacl access to .env/config/logs/spool)
are gone, and were never on the new files to begin with. Confirmed live:
surfaces as "No permission to read .../.env" in the Extensions tab right
after a restore, on a box where the dashboard was already working fine.
The restore script can't fix this itself (runs standalone, no access to
services/*.sh's functions or the dashboard's service-user name), so it now
prints a reminder to re-run `sudo ./setup.sh security-dashboard` at the end
of a restore when the dashboard's systemd unit is present, plus a matching
note in the generated README.
Filters by IP/range, scenario, ASN, carrier name, or country in a single
free-text field -- asked for so a phone's current IP or its network/carrier
name can be searched directly instead of scanning the full ban list by eye.
DEVICE_MARKER_RE expected "; === Device: NAME [AA:marker] (category) ===",
but device_config's own template (further down this file) generates
"; === Device: NAME (category) [AA:marker] ===" -- category parens before
the AA tag, not after. The regex never matched a real device comment, so
list_extensions() silently returned [] for every device on every install,
and /api/pstn-permissions served {"extensions": []} regardless of what was
actually in pstn-permissions.conf. That's why the Extensions tab's
Messaging/Voicemail checkboxes always rendered unchecked after a save +
reload even though the file itself had messaging=yes/voicemail=yes written
correctly -- the JS falls back to an all-default row when the endpoint
returns nothing. ea_list_devices() and the rename-device code parse the
same comment via string-splitting/a differently-shaped regex and were
already correct; this was the one broken parser.
The prior fix (930233c) wrote noload lines for app_voicemail_imap.so/
app_voicemail_odbc.so into config/asterisk/modules.conf at install time, but
vendor's docker/entrypoint.sh regenerates /etc/asterisk/modules.conf
unconditionally on every container start (same bind-mounted file) and
clobbers it within seconds — confirmed live on a fresh install with the
prior fix in place. Move the noload patch into
_asterisk_refresh_vendor_files()'s sed pass over the vendor template
instead, alongside the existing logger.conf/cert-regen patches, so it
survives the container's own regeneration. Drop the now-dead host-side
_asterisk_write_modules_conf() and its call sites.
Live-discovered bug, present on every install using this script, not
specific to any one box or extension: the easy-asterisk image ships
app_voicemail.so, app_voicemail_imap.so, and app_voicemail_odbc.so all
autoloading by default -- three alternative storage backends for the
SAME application (VoiceMail, VoiceMailMain, VMAuthenticate,
VoiceMailPlayMsg, VMSayName, the VM_INFO function, several AMI
actions), which collide registering those names against each other on
every single Asterisk start. This box's own container log showed the
exact signature on every restart: "Already have an application
'VoiceMail'" (and every sibling) followed by "app_voicemail.c:15897
load_module: Failure registering applications, functions or tests" --
app_voicemail never actually finished loading. Confirmed against
Asterisk's own documentation this session rather than assumed: this
is a known multi-backend conflict ("administrators should enable only
one module at a time"), not something specific to this repo's config.
Fix: new _asterisk_write_modules_conf, called from both the fresh-
install and update paths (matching voicemail-dialplan.conf's own
call-site pattern) alongside the other config/asterisk files, all
sharing the already-bind-mounted ./config/asterisk:/etc/asterisk
volume -- no new mount needed. noloads the two backends this repo
never configures (no IMAP/ODBC settings are ever written anywhere in
this script), leaving only the plain file-based app_voicemail.so
(the one voicemail.conf's [default] mailboxes actually target) to
load cleanly. Regenerated on every install/update, unlike
voicemail.conf, since modules.conf carries no per-install state of
its own -- consistent with how messaging-dialplan.conf/voicemail-
dialplan.conf are already handled, and added to their same chmod 644
line.
Requires a container restart to take effect on an existing install
(re-run `sudo ./setup.sh asterisk` -> Update, which regenerates this
file, then restart the container once).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Live blocker: user's actual migration is DigitalOcean droplet
(asterisk-digital-ocean/, container easy-asterisk-do) -> IONOS (plain
asterisk/, container easy-asterisk) -- exactly the case this script
didn't handle. The archive's own top-level directory name and
docker-compose.yml reflect whichever layout produced it
(_asterisk_resolve_layout's two known layouts). The previous restore
extracted straight into $PARENT_DIR, which recreates whatever name is
baked into the archive -- restoring a droplet archive onto a fresh
non-droplet install would land the data at a *second*,
wrongly-named directory (asterisk-digital-ocean) alongside the
freshly-installed one it was meant to replace, with docker-compose.yml
still naming the old project/container(s). Every service that resolves
Asterisk's layout by directory/container name (security-dashboard.sh,
pstn-trunk.sh, CrowdSec's Asterisk acquisition, Caddy) would get
confused by having two candidate layouts on disk, one of them stale
and half-wired.
Fix: extract into a scratch staging directory first. If the archived
docker-compose.yml's container_name differs from this run's own
$CONTAINER (baked in at generation time, so always correct for
whichever layout THIS box's install actually uses), rewrite the
project name, container name, and coturn container name in place
(coturn's is always "$CONTAINER-coturn" on both known layouts, so no
lookup table needed) before the data ever lands at $HERE -- never
lets the archive's own naming leak through. A same-layout restore
(most common case, or two droplet boxes, or two plain boxes) detects
no mismatch and skips the rewrite entirely, unchanged from before.
Verified against the real generated script (extracted from the
heredoc, not reimplemented): a droplet-flavored archive restored onto
a fresh plain-layout box lands at the correct single directory with
no stray second directory, and docker-compose.yml's name/container_name/
coturn container_name all correctly rewritten to the plain layout
(confirmed by diffing the actual restored file, not just checking for
absence of errors); a same-layout restore (droplet archive onto a
droplet box) confirmed to skip the rewrite entirely; the pre-existing
external-IP patch (previous commit) still fires correctly stacked on
top of the layout fix; and the extraction-failure rollback path (a
corrupt/unreadable archive) still restores the pre-restore install
untouched, verified via a marker file surviving the rollback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
User confirmed via a live pjsip.conf on their actual DO box: both
external_media_address and external_signaling_address are literal
IPs, written by easy-asterisk at first container start (not something
this repo's install script controls directly). A straight restore of
a backup archive from a DIFFERENT box onto the new IONOS box would
leave the OLD DigitalOcean IP baked into pjsip.conf — dialplan and
PJSIP device credentials would come back fine, but RTP media (and
likely SIP signaling/registration) would stay broken, silently, since
nothing in the restore path previously touched these values.
Fix: after extracting the archive, `restore` reads the archive's own
external_signaling_address as "old IP", detects this host's actual
current public IP (same DO-metadata -> ifconfig.me -> hostname -I
fallback chain services/asterisk.sh's own install already uses), and
if they differ, rewrites every occurrence across config/ and .env
(fixed-string match, not a regex, so the IP's dots can't be
misinterpreted). Deliberately does NOT touch spool/, logs/, or lib/ —
those hold voicemail messages and call recordings, and a blind text
substitution across binary audio would corrupt it. A restore onto the
same host (e.g. rolling back a bad config change, no IP change)
leaves every file untouched — the check only fires on an actual
mismatch.
Verified against the real generated script (extracted verbatim from
the heredoc, not a reimplementation) with a full mock backup/restore
cycle: built a fixture archive with pjsip.conf's three transport
blocks (udp/tcp/tls) and .env's TURN_SERVER all hardcoded to a fake
"old box" IP, plus a fake binary voicemail file; restored it onto a
mocked "new box" with a different detected IP via a stubbed curl.
Confirmed every occurrence in both pjsip.conf and .env was correctly
rewritten to the new IP, and confirmed via byte-for-byte comparison
that the binary voicemail file was completely untouched. Separately
verified the same-IP case (mocked curl returning the archive's own
IP) makes no changes at all, matching a same-host config rollback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
This is what's actually needed for a DigitalOcean -> IONOS Asterisk
migration (item #1 of the user's original 4-item list): move the
whole stack to different hardware entirely, not restore onto the
box the offsite mirror already targets.
dr_bringup_kopia.sh only ever scanned DEST_NAMES for restorable
snapshots — those are always local filesystem Kopia repos
(services/backup.sh creates them with `repository create filesystem
--path=...`), meaning they only exist on whichever box originally ran
the backup. On a genuinely new box, every one of them fails to
connect and there's nothing left to restore from — the script's own
header comment only covered the case where "the spare box IS the box
the primary's mirror targets" (i.e. already holds a copy of the repo
data), not a fresh, unrelated box.
Fix: also try REMOTE_TYPE/REMOTE_ARGS (the offsite Backblaze/S3
mirror, if configured) as a same-shaped destination named "offsite",
reusing DEST_default_PASSWORD since sync-to always mirrors that exact
same encrypted repo. Connects once into a fresh local config file
scoped to this DR run, then folds into the existing per-destination
scan/restore loop unchanged — "offsite" just becomes another entry in
_DEST_ARR. Documented the actual migration workflow in the header
comment, including the BACKUP_CONF override so copying the old box's
backup.conf over doesn't clobber the new box's own freshly-configured
one.
Verified against the real script (not a reimplementation) with mocked
kopia/docker binaries and a crafted backup.conf, covering: local dest
unreachable + offsite connects successfully (the actual migration
shape) with correct service/path discovery; a real (non---list)
restore run confirmed it selects the latest of multiple snapshots by
startTime and issues the correct `kopia restore <snapshot> <path>`
call; offsite connect failing (bad REMOTE_ARGS) warns and degrades to
"no restorable sources found" instead of crashing; REMOTE_TYPE=none
skips the new code path entirely with no behavior change (regression
check against the pre-existing local-only case).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
User's actual question: Backblaze B2's web console lets them browse a
bucket as folders/files; Garage has no equivalent by default, so after
switching an additional backup mirror from Backblaze-only to also
target local Garage, they had no way to visually confirm data landed
there the way they could on Backblaze. "S3 storage is opaque, you
can't browse it" was true of Garage's *own* CLI, but wrong as a
blanket statement — Backblaze's browsability comes from a client (its
web console) layered on top of the same kind of object storage, and
Garage has an actively-maintained equivalent (khairul169/garage-webui,
1.1k stars, "integrated objects/bucket browser") that gives the same
experience against Garage's S3 API.
services/garage-webui.sh (new): standard service-template Docker
service. Requires an existing services/garage.sh install (checks for
$DOCKER_DIR/garage/.env, errors with instructions if missing — this
is a browser for an existing instance, not a replacement). Reaches
Garage over host.docker.internal (both containers' ports are already
published to the host — simpler and more robust than trying to join
garage's own Compose-project-scoped default network by name). Has its
own login (AUTH_USER_PASS, bcrypt via a throwaway `docker run --rm
httpd:alpine htpasswd` — same $ -> $$ escaping services/wg-easy.sh
already uses for its own bcrypt PASSWORD_HASH, verified here against a
real docker compose config run: unescaped, Compose tries to interpolate
$2y$05... as variable references and silently corrupts the value with
a "not set" warning; escaped, it passes through intact with no
warning), so it doesn't need Authelia gating by default.
Prerequisite fix in services/garage.sh: its admin API (bucket/key
management, object listing — the thing garage-webui talks to) has
been running with zero authentication since this service was first
built, because admin_token was never set in garage.toml. Nothing in
this repo called that API before now, so it went unnoticed; adding a
real consumer is what surfaced it. Fixed: generate admin_token
(openssl rand -base64 32) alongside the existing rpc_secret, persist
GARAGE_ADMIN_TOKEN/GARAGE_ADMIN_PORT to .env for garage-webui to read
locally (never sent over SSH, unlike the S3 credentials backup.sh
reads remotely). Update mode backfills admin_token into an existing
garage.toml (+ restarts just the garage container to apply it) for
anyone who installed before this change, same backfill-not-break
approach as the GARAGE_S3_API_PORT fix from the previous commit.
Verified: bash -n on both files; docker compose config against real
Docker Compose for both the primary garage.toml/.env generation (with
the new admin_token/GARAGE_ADMIN_PORT fields) and the new
garage-webui docker-compose.yml; the bcrypt-escaping behavior
specifically (proved via a minimal repro that unescaped $ corrupts
the value with a warning, escaped does not); the admin_token/
GARAGE_ADMIN_PORT Update-mode backfill logic against old- and
new-style .env/garage.toml fixtures, including idempotency (running
it twice adds nothing a second time); and the credential-parsing
regexes in garage-webui.sh against both a complete .env fixture and
an old one missing the new fields (confirms the "run garage's Update
first" error path actually triggers rather than proceeding with
blanks).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Live failure: `sudo ./setup.sh garage` → Full reinstall on a box that
already had a working Garage install crashed with:
Error: ApplyClusterLayout returned InternalError (500): Internal
error: Invalid new layout version
Root cause: a "Full reinstall" deliberately never wipes ./data or
./meta (that's real backup-mirror data — Kopia's sync-to s3 target —
and losing it silently on reinstall would be far worse than the
alternative), but the cluster-init step unconditionally re-ran `garage
layout assign` + `layout apply --version 1` every time it was reached.
Garage requires each apply to be exactly previous_version + 1; a node
that already has a committed layout (from the earlier install, still
sitting in the preserved ./meta) rejects a second "1". Fix: check
`garage status` for "NO ROLE ASSIGNED" first and only run the
assign/apply once, matching what the surrounding comment already
claimed happened ("Only ever run once") but the code didn't enforce.
Second, related issue this would have hit immediately after: the same
reused-./meta state almost always means an existing bucket + key from
the earlier install are still sitting in Garage's storage. The fresh
flow was about to silently create a brand-new bucket/key and overwrite
.env to point at those instead — orphaning any real data already in
the old bucket (nothing left on disk pointing at it, even though it's
still physically stored). Now: when the layout is already applied,
list existing buckets and require an explicit y/n (default n) before
creating new ones, with recovery instructions for reconnecting to an
existing bucket by hand instead.
Third, the actual reason a full reinstall was reached at all: Update
mode never backfills .env fields added to this script after someone's
initial install (GARAGE_S3_API_PORT, needed by services/backup.sh to
read an instance remotely) since Update deliberately never touches
.env otherwise — the only other path was the now-unsafe fresh
reinstall. Update now backfills just that missing key by reading the
real port back out of the already-written docker-compose.yml, so a
future .env schema addition doesn't force this tradeoff again.
Verified with standalone harnesses (not the live install, mocked
`garage status`/bucket-list output and .env/docker-compose.yml
fixtures): all four layout-state branches (fresh node, existing
buckets + decline, existing buckets + confirm, existing role but no
buckets), and both backfill cases (missing key added, existing key
left alone). Caught and fixed a real bug in the first draft of the
port-extraction regex during this testing — grep -oE '^[0-9]+' never
matched because the captured group still had its surrounding quotes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Previously, if the additional-mirror S3/Garage check couldn't find
~/docker/garage/.env on the remote box, it just warned and silently
dropped the mirror — forcing a full re-run (and re-entering every
already-answered prompt: destinations, passwords, schedule, B2,
DR-spare, etc.) once Garage was actually installed.
Wrap the S3/SFTP branch in a loop so the "Garage isn't installed yet"
case now offers a real 3-way choice:
1) install Garage in another session, then retry the same .env check
without leaving this script
2) fall back to SFTP for this one mirror, reusing the already-resolved
destination host/port/user/mirror-name with no re-prompting
3) skip just this mirror (default — safe for UNATTENDED, which
resolves to this automatically since prompt_text returns its
default without blocking)
Everything else install_backup() has already collected lives outside
this loop, so none of it is at risk regardless of which of the three
exits it via.
Verified against a standalone harness reproducing the state machine
with a mocked ssh (empty .env vs. populated .env after a simulated
install) and prompt_text, covering all three interactive choices, the
blank/Enter default, and UNATTENDED mode (confirms the blocking
"press Enter to retry" read is unreachable there since prompt_text
resolves choice 1's prompt to default "3" without waiting on stdin).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Extends the "ADDITIONAL MIRROR" section (previously SFTP-only) with a
type choice: SFTP, or S3 against a Garage instance already running on
that box. For the S3 path, this script never asks the operator to retype
a bucket name or key — it SSHes to the destination, reads
~/docker/garage/.env directly (the real, currently-configured values,
generated once by services/garage.sh and never touched again on its own
Update runs), and uses those for the dry-run verification and the
persisted mirror args. If Garage isn't installed there yet, it says so
plainly with the exact install command instead of failing cryptically or
silently skipping.
Also removed the last hardcoded suggestions from services/garage.sh
itself ("kopia-backup" / "kopia" as fixed prompt defaults) — replaced
with a freshly-generated suggestion each run (timestamp-suffixed), so
nothing about the bucket/key name is a fixed string baked into this repo
at any point in the chain; it's always the operator's actual choice, read
back live wherever it's needed.
Verified end-to-end against a mocked ssh (returning realistic
~/docker/garage/.env content) covering both outcomes: Garage installed
with a real bucket/key correctly parsed, dry-run run, and persisted; and
Garage missing, correctly warning with the install command and leaving
backup.conf untouched either way.
Garage's real CLI output pads labels with extra spaces for column
alignment ("Key ID: GKxxxx"), not a single space like the
mocked test used ("Key ID: GKxxxx") — the fixed ": " field separator left
that padding stuck to the parsed value, so .env ended up with access
key/secret strings carrying leading whitespace inside the quotes.
Confirmed live by the user right after install. This would have broken S3
auth outright once actually used, since access keys have to match exactly.
Switched to ':[[:space:]]+' as a regex field separator, which consumes
however many spaces are actually there instead of assuming exactly one.
Verified against both the single-space and padded/aligned formats — both
now produce the identical clean value with no leading whitespace.
MinIO's open-source community edition is dead: console GUI stripped May
2025, Docker images stopped publishing October 2025, repo formally
archived April 2026, with MinIO redirecting everyone to their paid AIStor
product. Verified this directly before building anything, since recommending
a since-abandoned image would have been worse than the SFTP problem this
was meant to solve.
Garage (Deuxfleurs) is the actively-maintained small-scale self-hosted
replacement — single Rust binary, purpose-built for exactly this "one
lightweight node" use case (as opposed to SeaweedFS, which targets large
object counts / large-scale deployments, more machinery than a single
backup-mirror target needs).
services/garage.sh follows this repo's standard service template: port
scanning for the S3 API/RPC/admin ports, an RPC secret generated once and
never touched again on Update, and a one-time cluster init sequence
(layout assign/apply, bucket create, key create, bucket allow) gated on
whether .env already has a saved access key — Update reruns skip all of it
and just refresh the image.
Primary intended use: a local S3-compatible target for services/backup.sh's
additional-mirror Kopia sync, so a local mirror can reuse the exact same
sync-to s3 code path already proven reliable for the Backblaze B2 mirror,
instead of Kopia's separate, less-exercised SFTP backend that's been the
source of today's connection troubleshooting.
Verified end-to-end against a mocked environment (fake docker exec
returning realistic `garage status`/`garage key create` output) — caught
and fixed a real off-by-one in the status-output parsing this way (grabbed
the column-header row's literal "ID" instead of the actual node ID; output
has a title line, then a header line, then the data row). Also validated
the generated docker-compose.yml with real `docker compose config` in both
the no-network and network-created cases.
The additional-mirror "Remote path for the repo" prompt always suggested
a generic ~/backups/kopia-mirror default, unrelated to wherever the
operator already pointed the DR-spare sync. Requested directly: default
to that same location instead, in its own /kopia-data subdirectory so
Kopia's repository files don't end up visually mixed in with the two
plain config files (backup.conf, README.md) the DR-spare sync writes
straight into DR_SYNC_PATH itself.
Falls back to the original generic default when DR_SYNC_PATH isn't set
(no DR-spare configured yet). Verified the path computation handles a
DR_SYNC_PATH with or without a trailing slash correctly (no double slash),
and the unset case still falls back as before.
The additional-mirror setup already resolves user/hostname through ssh -G
so a ~/.ssh/config alias works, but never extracted port — Kopia's sftp
storage backend doesn't read ~/.ssh/config at all and defaults to 22
regardless of what the alias actually configures. Confirmed live: this
produced "server unexpectedly closed connection: unexpected EOF" on the
dry-run verification — Kopia connecting to the right host on the wrong
port, not a credentials or host-key issue, which is exactly why plain
`ssh main` kept working the entire time this was being debugged (it reads
the alias's Port line correctly).
Now parses `port` out of the same ssh -G output, defaults to 22 if absent
(matching ssh's own default), and passes --port= through to both the
dry-run check and the persisted EXTRA_MIRROR_ARGS string — the latter
matters as much as the former, since that's what every actual scheduled
sync reuses afterward, not just the one-time verification.
Verified the parsing against three cases: a custom-port alias, a
default-port alias, and an unresolvable alias — all three resolve to the
correct port with no manual intervention needed.
extras/fix_pikapods_dump.py patches two confirmed Adminer PostgreSQL-export
bugs that otherwise make a PikaPods Mattermost migration fail outright:
unquoted enum-label DEFAULT values (Postgres reads the bare label as a
column reference and rejects the CREATE TABLE) and boolean columns
serialized as bare 0/1 instead of true/false (Postgres doesn't implicitly
cast integers to boolean). Boolean columns are discovered by actually
parsing each CREATE TABLE in the dump rather than working from a
hand-curated list — Postgres only reports the first bad column per failed
row, so a list built from error output alone would likely be incomplete.
Already verified earlier this session against a real local Postgres 16
instance; reviewed now for anything needing redaction before committing —
it's a generic text-processing tool with no hostnames, credentials, file
paths, or personal data in it, so nothing needed changing.
Cross-referenced from the generated migrate-from-pikapods.sh's header
comment (services/mattermost.sh) so anyone hitting a CREATE TYPE/CREATE
TABLE or boolean-column import error is pointed at the fix instead of
having to rediscover it.
The generated migrate-from-pikapods.sh (services/mattermost.sh's existing
"Migrating from an existing Mattermost instance?" prompt on fresh installs)
already correctly parameterizes PROJECT_DIR/MM_CONTAINER/DB_CONTAINER per
instance — no bug there. What it missed: after rsync/cp-ing files in from
the export, it never touched ownership, so the imported ./data landed
owned by whoever ran the script instead of the fixed UID 2000
mattermost/mattermost-team-edition runs as. Every file write then failed
with permission denied — confirmed live as the actual cause of a
client-side "stream closed" error on image/file uploads after a real
migration.
Adds chown -R 2000:2000 ./data right after the copy step, and a root
check up front since chowning to an arbitrary UID needs it (docker/psql
access already implied running as root in practice, just never enforced
explicitly). Usage lines updated to say `sudo` to match.
Verified by reconstructing the exact generated script from the real
source heredocs (head + variable substitution + body, the same three
pieces the actual cat/cat>> sequence produces) and syntax-checking the
result — root check and chown both land in the right place, and
PROJECT_DIR/MM_CONTAINER/DB_CONTAINER still resolve correctly per instance.
Root-caused a live "stream closed" image-upload failure to
data/20260814/.../mkdir: permission denied — the volumes weren't owned by
the fixed UID 2000 mattermost/mattermost-team-edition runs as, most likely
left that way by the PikaPods data import. The install script already
chown -R 2000:2000's these on every run (fresh or update), so re-running
the installer would have fixed it — but that still means remembering to
re-run it every time ownership drifts for any reason, including causes
this repo doesn't control (a future migration, a manual restore, anything
that copies files in as a different UID).
Added a small mattermost-fix-perms init container (busybox, chown, exit)
that the mattermost service now depends on via
condition: service_completed_successfully. Runs on every `docker compose
up` — including a plain host reboot, since restart: unless-stopped brings
the stack back on its own — so this self-heals permanently instead of
needing a human to notice and fix it by hand again.
Verified the generated compose file (with representative variable values)
against real `docker compose config`: valid YAML, and the dependency graph
correctly shows mattermost waiting on both db (service_healthy) and
mattermost-fix-perms (service_completed_successfully).
Two independent, requested changes:
- services/pihole.sh: new standalone service, Pi-hole v6 (the image moved
entirely to a TOML-based /etc/pihole config — the old WEBPASSWORD env var
and separate /etc/dnsmasq.d volume are both gone; uses
FTLCONF_webserver_api_password and FTLCONF_dns_listeningMode=ALL
instead). Deliberately not wired into wg-easy or any other VPN — a device
has to be pointed at it manually (per-device or via router DHCP). DNS
itself (53/tcp+udp) is never scanned/moved since shifting it off the
standard port would defeat the point; a port_in_use check warns instead
of blocking, since the common case (systemd-resolved on 127.0.0.53 only)
doesn't actually collide with Pi-hole binding the host's real interfaces.
Web admin UI is Caddy-fronted like everything else in this repo. Added to
the README services table.
- services/wg-easy.sh: default VPN/web ports moved from 51820/51821 to
51830/51831. Netbird's own WireGuard listener also defaults to exactly
51820 — installing both on one box means wg-easy's existing scan-and-move
logic would silently shift its port every time, which is harder to
predict/document than just not starting on the collision in the first
place. The scan itself is unchanged and still moves both ports further if
even the new default is taken.
Tested pihole.sh's full standalone install flow (no-Caddy and
Caddy-present-locally cases) against a mocked environment, validating both
generated docker-compose.yml files with `docker compose config`, and
confirmed the reinstall-mode gate correctly no-ops on a second run in
unattended mode.
ensure_coturn_user() auto-installs shared coturn (services/coturn.sh) any
time $DOCKER_DIR/coturn doesn't exist — correct behavior for "first
service that needs TURN", wrong behavior for "an operator deliberately
decided every consumer should run its own dedicated coturn instead and
removed the shared one on purpose". The function had no way to tell those
two states apart, so deleting ~/docker/coturn didn't actually retire it —
the next service to call this function (a Mattermost reinstall, a fresh
Asterisk install) would silently bring it right back.
touch ~/docker/.coturn-retired now short-circuits the function straight to
the existing "no TURN available, caller degrades gracefully" return path,
before it ever looks at install_coturn. Every existing consumer (Asterisk,
Mattermost) already handles that path correctly today — it's exactly what
happens if shared coturn simply fails to install — so this needed no
changes on the consumer side, only closing the gap in the shared function.
Verified with a mock: with the flag present, install_coturn is never
invoked and the function returns empty COTURN_HOST/rc=1 as expected.
Requested after a rerun silently reset DR_SYNC_PATH (fixed separately) —
auditing the rest of install_backup() turned up the same class of bug in
several other places, one of them worse than the one that prompted this:
- Default destination repo path defaulted to $ACTUAL_HOME/backups/... even
when the real configured repo was somewhere else entirely (this user's
actual path is /root/backups/kopia-backup) — accepting the shown default
on a rerun would have pointed the installer at the wrong location.
- Extra (non-"default") destinations weren't preserved AT ALL on a rerun —
skipping "Add more destinations?" silently dropped every extra
destination, and anything mapped to it, from the rewritten backup.conf.
- The per-service destination-assignment prompt always showed "[default]"
regardless of the service's actual existing mapping.
- ntfy URL/token always started blank, silently disabling notifications on
any rerun where they weren't retyped.
- The schedule prompt always defaulted to option 1 (daily 02:00) instead of
reading back whatever OnCalendar was actually already running.
- B2's four sub-fields (bucket/endpoint/key ID/secret) always started
blank even when reconfiguring an already-working REMOTE_TYPE=s3 setup —
a mispaste on any one of the four meant retyping all four blind, since
there was nothing to fall back to per-field (the existing REMOTE_ARGS was
already preserved as a whole on a blank/failed attempt, just not offered
back as individual editable defaults).
All six read the same way: pull the existing value from backup.conf (or,
for the schedule, from the live systemd timer unit — schedule isn't stored
in backup.conf) and use it as the prompt default, so accepting the default
keeps what's already there instead of silently reverting it. Verified all
six against a mock backup.conf + timer fixture with pre-existing values for
every field this touches.
Known remaining gap: KEEP_LATEST (retention count) still isn't read back —
doing so correctly needs the repo already connected, which happens later
in this same function's flow. Flagging rather than rushing a reorder here.
Two stacked bugs, found together when re-running the backup installer to
add an SFTP mirror silently reverted a previously-set absolute
DR_SYNC_PATH back to the script's tilde-based default, which then failed
outright:
1. services/backup.sh never read DR_SYNC_HOST/DR_SYNC_PATH back from an
existing backup.conf before prompting (every other setting in this file
does — passwords, mirrors). Accepting the prompt defaults on a rerun
silently reset both to blank/"~/docker/backup" instead of keeping what
was already configured. Fixed by reading them back the same way
DEST_*_PASSWORD already does.
2. extras/backup_kopia.sh's DR-spare sync wraps the remote path in single
quotes for its `ssh host "mkdir -p '...'"` / `"chmod 600 '.../...'"`
commands. Single-quoting a leading ~ stops the remote shell from
expanding it at all, so it looked for a literal directory named "~"
instead of the home directory — breaking the script's own DEFAULT
DR_SYNC_PATH ("~/docker/backup") for anyone who actually used it.
rsync's own transfer step has separate, correct tilde handling, which is
why the sync itself "succeeded" while the follow-up chmod couldn't find
the file. Fixed with a small _dr_remote_quote() helper that keeps a
leading ~/ outside the quotes while still safely quoting the rest of
the path.
Verified the quoting fix by parsing the exact constructed command string
in bash directly — a plain '~/docker/backup' stays literal (the bug),
~/'docker/backup' correctly expands to $HOME/docker/backup (the fix).
_backup_ensure_root_ssh_key() only ever checked for /root/.ssh/id_ed25519
or id_rsa by exact filename. Root can already SSH to the DR-spare/mirror
host just fine in practice (proven by this same script's own DR-spare sync
succeeding), just via a key with some other name — so the function had no
way to see that and always fell through to offering a copy-from-user-home
or brand-new ssh-keygen, both unnecessary.
Now takes the target host as an optional argument. When given, it tests
root's SSH access to that host as-is first and resolves the actual key via
`ssh -G <host>` (which expands ~/.ssh/config the same way the SFTP-dest
resolution earlier in this file already does) before falling back to the
copy/generate prompts. Both call sites (DR-spare, SFTP mirror) now pass
their respective host.
Verified against a mock ssh: an already-working non-default-named key gets
detected and reused with no prompts, and the original copy/generate
fallback still triggers correctly when SSH genuinely doesn't work yet.
The freshly-added raw-error logging paid off immediately: the box's spare
sync was failing every run with "scp: Connection closed" while plain ssh
exec to the same host worked fine. That split (ssh exec OK, scp specifically
rejected) matches modern OpenSSH's default scp-over-SFTP transfer hitting a
restriction on the remote side that a plain exec or rsync's own protocol
don't trigger.
Swapped the scp step for rsync -a over the same ssh options, keeping the
ssh mkdir -p before it (rsync doesn't create missing destination
directories) and the ssh chmod after. Verified the exact command/quoting
against mocked ssh/rsync binaries — array expansion and remote path
handling both check out.
Last night's fully-failed backup run (0/20, "repository not found" on
every service) showed the real gap: categorize_error() clearly saw real
content in $_ERR (it matched a specific pattern, not the generic
fallback), but log_raw_error()'s separate re-read of the same file moments
later came back empty on every single failure — so the raw-error logging
added earlier this session produced nothing when it mattered most.
Fixed by reading $_ERR into a variable exactly once per failure and
passing that string to both categorize_error() and log_raw_error(),
instead of two independent file reads. Verified against a mock harness
reproducing the same call pattern (three simulated failures in a loop,
single shared error file) — both the categorized reason and the raw
stderr text now come through on every iteration.
Doesn't explain why last night's repo access failed in the first place
(disk and mount checks came back clean) — but the next time it happens,
this will actually surface the real kopia error instead of losing it.
Asterisk's whole state (dialplan, pjsip devices, voicemail, recordings,
.env with its coturn credential, docker-compose.yml) already lives under
one self-contained directory, so asterisk-standalone-backup.sh just tars
it — with stop/restart safety around the tar since voicemail/spool write
continuously, and a move-aside-then-extract restore that rolls back
automatically if extraction fails. Written into the install directory at
both fresh-install and update time via _asterisk_write_standalone_backup_script().
Output defaults to ~/asterisk-backups/, deliberately outside ~/docker/, so
a Kopia backup of the box doesn't also back up a backup-of-itself. Meant
for a quick pre-change snapshot or moving this PBX to a new host without
standing up the full backup stack first.
Documented in the generated README's new "Standalone backup/restore"
section. Tested against a mocked EA_DIR (fake docker/docker compose,
config/spool/voicemail files) confirming backup produces a correct tar and
restore replaces content correctly with rollback on extraction failure.
Confirmed live and cross-checked against a real, documented Kopia issue
(kopia/kopia#5329): the walkthrough previously told the operator to scope
the Application Key to just the bucket they created — the more
security-conservative default, and correct for B2's own S3-compatible
API in general. But Kopia specifically needs the listBuckets capability
even though it only ever touches the one configured bucket, and B2's
basic "Add a New Application Key" web form doesn't expose a way to grant
listBuckets on a bucket-restricted key — only an account-wide ("All")
key gets it through that form. Without it, the connection fails with
B2's unhelpful "Cannot access bucket" error, which doesn't point at the
actual missing capability at all.
Updated the guidance to "All" with the reasoning inline, and a note that
single-bucket scoping is still possible for anyone willing to create the
key via B2's CLI/API directly (b2_create_key with an explicit
capabilities list including listBuckets) rather than the basic web form
this walkthrough is written for.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Two separate fixes from a live report.
1. The DR-spare and SFTP-mirror sections both checked ONLY /root/.ssh
for a key, missing the common case: the person running `sudo
./setup.sh backup` already has a key under their own home directory
(used interactively, quite possibly already authorized on the target
box), while root — who actually runs the scheduled systemd service —
has none. Confirmed live: "the computer has the ssh key for the sudo
user on the box" produced "No SSH key found for root" with no
inline way to do anything about it beyond a pointer to go set one up
elsewhere and re-run.
Factored both call sites into one shared _backup_ensure_root_ssh_key()
that checks root first, then offers to reuse the sudo user's existing
keypair (copied into /root/.ssh with correct ownership/permissions,
root:root 600) before falling back to generating a brand new one —
reusing an existing key can work immediately if it's already
authorized on the target, where a fresh key needs a new ssh-copy-id
round-trip regardless. Verified all three branches (root already has
a key, root has none but the user does and accepts reuse, neither
exists and one gets generated) against a mocked filesystem.
2. The B2 dry-run failure message read like it could be about missing
input even when every field was non-empty — confirmed there's no
code path where non-blank-but-wrong values actually trigger the
separate "Left blank" message (the two are on disjoint branches), but
the dry-run failure text itself didn't rule that out or point at the
actual likely cause. Now echoes back what was entered (bucket,
endpoint, Key ID — never the secret) so it's easy to eyeball against
B2's own confirmation screen, states plainly that this is a rejection
of non-blank input, and names the most likely cause directly: pairing
the Key ID from one Application Key with the Secret from a different
one, which is easy to do after creating more than one while
troubleshooting.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Confirmed live: a real failure ("WARNING: spare sync failed — error —
see system logs on ubuntu") didn't match any of categorize_error()'s
known patterns, fell into its generic catch-all bucket, and the actual
stderr text that would have explained it was sitting in a mktemp'd file
this script deletes on exit (trap ... EXIT) — so there was nothing in
"system logs" to actually go check. The categorization was silently
discarding the one piece of information that would have diagnosed the
problem.
Added log_raw_error(), called right after every categorize_error() site
(5 of them: two snapshot-failure paths, the primary REMOTE_TYPE mirror,
the new EXTRA_MIRROR_NAMES loop, and the DR-spare sync) — logs the raw
stderr text (truncated to 500 chars) into the same log stream as
everything else, so it survives past the run that produced it instead
of being deleted with the temp file. categorize_error()'s short bucket
label is untouched and still used for FAILED_SVCS/notification text,
which should stay concise — this adds the detail alongside it, not
instead of it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Confirmed: Kopia's sync-to sftp has its own SFTP client and doesn't read
~/.ssh/config the way the system ssh/scp binaries do — so an alias set
up via wg-easy's sync-ssh-aliases.sh (or any ~/.ssh/config Host entry)
worked fine for the DR-spare connectivity check (which shells out to
real ssh) but silently failed for this mirror: a plain @-split on an
alias like "main" (no @ present) produced --host=main, a name that only
resolves inside ~/.ssh/config, not real DNS. The dry-run check correctly
rejected it and the mirror was never saved — no error surfaced beyond
that, so it looked like nothing happened.
Now resolves the destination through `ssh -G` before building the Kopia
flags — the same mechanism ssh itself uses to expand config aliases —
and falls back to the previous plain @-split only if that comes back
empty. Verified against three cases: a bare alias (resolves via a mock
~/.ssh/config Host block), an explicit user@ip (passes through
unchanged), and an unrecognized name (falls back to a sane literal
hostname rather than erroring).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Requested: mirror to Backblaze B2 AND directly to the IONOS spare box
over Tailscale, at the same time, not one or the other. REMOTE_TYPE/
REMOTE_ARGS was hardcoded to a single mirror target — extending it to a
list would have meant redesigning the one thing that already works and
was already verified against real B2 credentials, so this adds a
separate, additive mechanism instead: EXTRA_MIRROR_NAMES, a space-
separated list, with per-entry MIRROR_<name>_TYPE/_ARGS (same argument
shape as REMOTE_ARGS). An existing B2-only backup.conf keeps working
completely unchanged if this new section is skipped.
install_backup() gets a new "ADDITIONAL MIRROR" prompt after the
existing B2 section: offers a direct SFTP mirror (Kopia's sync-to sftp,
not the deprecated b2 provider — same reasoning as the S3/B2 choice
already made), defaults the destination to whatever was typed at the
DR-spare prompt above (same box, same purpose, no reason to ask twice),
checks passwordless SSH and an SSH key exist first, then verifies with a
--dry-run against the just-created 'default' repo before saving it —
same "don't save something broken" discipline as the B2 flow. Verified
against a mock backup.conf that install-side writes and worker-side
reads agree on the exact format, and that reusing an existing mirror
name reconfigures it instead of duplicating it in the name list.
One correction while researching sync-to sftp's flags: unlike plain ssh,
Kopia doesn't shell out to the system SSH client, so it needs an
explicit --keyfile and --known-hosts path rather than picking up
whatever `ssh` already trusts automatically — checked Kopia's own docs
for the exact flags before writing this, same as the earlier S3 case.
extras/backup_kopia.sh's worker loops through EXTRA_MIRROR_NAMES after
the existing REMOTE_TYPE mirror step, running sync-to for each
destination against each additional mirror and folding failures into
the same FAILED_SVCS/notification reporting the primary mirror already
uses. Verified end-to-end against a mock backup.conf and a stubbed
kp_for: both the B2 and the new SFTP mirror get called in sequence with
the correct arguments.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Requested after a live failure: pasting into the hidden Application Key
field silently captured nothing (terminal/SSH-client dependent), and the
only symptom was a generic "one or more fields left blank" warning after
all four prompts had already gone by — no way to tell which field, or
even that the paste itself was the problem rather than something else.
Each of the four fields now echoes its character count right after entry
(never the value for the hidden Application Key field, just its length),
so a failed paste is visible immediately instead of discovered several
prompts later. The blank-field warning now also names exactly which
field(s) were empty instead of a generic message.
Verified against the user's actual reported case: bucket/endpoint/key-ID
entered normally, Application Key came back empty — reproduces as
"(0 characters entered)" on that line and "Left blank: Application Key"
in the warning, both confirmed against a second case where all four
fields are present and it passes through cleanly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Requested: don't push the operator toward installing wg-easy if they
already have a different mesh VPN (Netbird or Tailscale) running —
detect any of the three first, and only offer a choice when none are
present.
Detection checks wg-easy's own directory (this repo's install marker),
then falls back to checking whether the netbird/tailscale binaries exist
AND their systemd services are actually active — not just installed,
since an installed-but-never-connected client isn't a usable path to the
spare box either. wg-easy takes priority if somehow more than one is
present, since it's this repo's own chain-installable option.
When none are detected, offers a numbered choice: wg-easy (chain-installs
via the existing declare -F guard), Netbird, or Tailscale (both via their
official curl-pipe-sh installers — verified the current URLs against
each vendor's own docs rather than guessing, since a wrong URL here would
be a bad thing to ship). Both third-party options still need a manual
follow-up step this script can't complete unattended (Netbird needs a
setup key from the operator's account, Tailscale needs an interactive
auth link) — the success message says so rather than implying the
install alone finishes the job.
Verified the detection branching against all the cases that matter:
nothing present, only wg-easy's directory, only Netbird active, only
Tailscale active, and multiple present at once (wg-easy correctly wins).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Requested improvement: the disaster-recovery spare prompt in backup.sh
already ran a live connectivity check and, on failure, printed manual
instructions (set up wg-easy separately if the spare isn't reachable,
run ssh-keygen/ssh-copy-id yourself) — but never offered to do any of it
right there, even though every piece is safe to automate inline.
Now, when the passwordless SSH check fails:
- If wg-easy isn't installed yet, offers to chain-install it (guarded
with declare -F install_wg-easy, same pattern asterisk.sh already uses
for security-dashboard/pstn-trunk) — covers the common case where the
spare is a home box with no port-forward and no path there at all yet,
not just a missing key.
- If root has no SSH key, offers to generate one (ssh-keygen -t ed25519).
- Offers to run ssh-copy-id against the spare interactively right there
— it prompts for the spare's login password itself, so this script
never touches or sees that password, just invokes the real command
inline instead of telling the operator to go run it themselves after.
- Re-runs the connectivity check after ssh-copy-id succeeds, so the
install flow reports the actual current state instead of the
pre-fix failure message.
Verified the has-a-key detection (the part most likely to have a subtle
&&/|| precedence bug) against all four cases — no key, only id_ed25519,
only id_rsa, both — behaves correctly in each.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Confirmed live: backup.sh has no update/fresh distinction and re-runs
every prompt on every invocation, including the repository password
prompt — which always minted a fresh (typed or auto-generated) password
regardless of whether a repo already existed at that destination's path.
Re-running the installer (to add a destination, configure the new B2
offsite mirror, or just by habit) then fails to connect to the real,
already-populated repo with "invalid repository password", because the
repo's actual password is permanently whatever was set the first time
and nothing read that back.
Each destination's password is now read back from the existing
backup.conf (if that destination name was already configured there)
before falling through to prompt/auto-generate — same pattern already
applied to REMOTE_TYPE/REMOTE_ARGS, EMBEDDED_COTURN_SLOT, and everywhere
else in this session that re-running a script with no update/fresh gate
turned out to silently regenerate something it shouldn't have. Verified
against a mock backup.conf: an existing destination's password is reused
verbatim, and a genuinely new destination name still falls through to
fresh generation correctly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Both branches independently solved the same Mattermost/Asterisk coturn
relay-port collision problem. Kept main's find_free_coturn_range-based
approach (documented in CLAUDE.md as the canonical pattern, and shared
across every coturn-owning service) over this branch's earlier
Mattermost-only port-slot scheme, and cleaned up the now-unused
_MM_COTURN_PORT/_MM_COTURN_MIN/_MM_COTURN_MAX/EMBEDDED_COTURN_SLOT
references that had auto-merged without conflict markers.
Answers a direct ask: offsite mirroring existed only as a REMOTE_TYPE/
REMOTE_ARGS placeholder in backup.conf with a comment pointing at
`kopia repository sync-to --help` — no interactive setup at all, B2 or
otherwise.
Checked before building anything: Kopia's dedicated `sync-to b2`
provider is marked [DEPRECATED] on kopia.io's own command reference.
B2 also offers an S3-compatible endpoint (s3.<region>.backblazeb2.com,
same application key works as the access/secret key pair), and Kopia's
`sync-to s3` provider isn't deprecated — so this targets that path
instead of building on a command on its way out.
What's now automated vs. guided, deliberately split:
- Bucket creation and the application key are walked through as console
steps, not automated. Object Lock specifically is a one-time,
bucket-creation-only decision with a real tradeoff (undeletable-by-
design vs. genuinely can't delete early) that shouldn't be silently
flipped either way by a script on someone's behalf.
- Once the operator has a bucket + endpoint + scoped application key
(B2 requires a key scoped to one bucket, not the account master key —
noted in the walkthrough), this becomes mechanical: run a
`sync-to s3 --dry-run` against the just-created 'default' repo to
verify the credentials actually work, and only then write
REMOTE_TYPE=s3 / REMOTE_ARGS into backup.conf. A bad bucket name or
key leaves REMOTE_TYPE at "none" with a clear error instead of saving
a broken config that fails silently at 2am.
- Encryption isn't a separate step — Kopia already encrypts client-side
with the repository password set earlier in this same flow; called
that out explicitly since it was asked about as if it needed its own
setup step.
Also fixed a regression the new prompt would otherwise have caused:
backup.sh has no update/fresh distinction and re-asks everything on
every run, so an already-configured offsite mirror is now read back
from the existing backup.conf and preserved by default — answering "no"
on a re-run no longer silently resets REMOTE_TYPE to "none".
Verified the control flow (not just bash -n) against a mock kopia
binary and stubbed prompts: good credentials wire up REMOTE_TYPE/
REMOTE_ARGS correctly, a rejected credential leaves REMOTE_TYPE at
"none" rather than saving something broken, an existing configured
value survives a "no" answer on re-run, and blank fields skip cleanly
without attempting a dry-run at all.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Answers a direct ask: the automated restore-verify test
(extras/test_backup_kopia.sh — verifies the latest snapshot, restores it
over a moved-aside copy, compares, rolls back, reports PASS/FAIL, sends
an ntfy notification) was already fully non-interactive and already
wired to a systemd timer/cron fallback by install_backup() — it just
had no schedule choice at all, hardcoded to weekly (Saturday 03:00).
Every service in this test stops briefly while its data gets moved
aside and restored back, same interruption profile as the main backup
job — so the schedule is a real tradeoff (more frequent verification vs.
more frequent blips), not a free "always pick the most frequent" choice.
Gave it the same Weekly/Monthly/Custom shape the main backup schedule
prompt above it already offers, instead of a single hardcoded option.
Also added an explicit "run the first test now?" prompt right after
scheduling it — otherwise choosing Monthly means waiting up to a month
before finding out whether the test even works, rather than getting
that initial confirmation immediately and then settling into the
chosen cadence.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Found while adding a live scanner to the coturn-slot code: WEB_PORT and
CALLS_UDP_PORT were scanned unconditionally, before the reinstall-mode
prompt even ran and before anything stopped the currently-running
container. On an "Update" run that meant find_free_port would see this
instance's OWN already-published port as occupied and silently shift it
to the next free one — every plain update could have moved the service's
port out from under already-configured Caddy routes, bookmarks, and the
Calls plugin's client config, without the operator asking for that.
services/asterisk.sh already gets this right for WEB_ADMIN_PORT: update
reads the existing port back from .env (no rescan), fresh scans from the
plain default only after stopping the old container. Brought Mattermost
in line with the same shape — the port resolution moved from before the
reinstall-mode block to after it, so MODE is known and, for a fresh
install/"Full reinstall", the old containers are already stopped by the
time it scans.
WEB_PORT/CALLS_UDP_PORT are now also written to .env directly (they
weren't before), with a fallback to parse them from the existing
MM_SERVICESETTINGS_LISTENADDRESS / docker-compose.yml port mapping for
installs made before this change — so an update on an already-running
instance doesn't regress just because its .env predates the new
variables. Verified the explicit-var, fallback-parse, and priority-order
(explicit wins over fallback) cases against a mock before shipping.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn