Confirmed live: install_frigate()'s fresh-install path overwrote a
working, hand-crafted docker-compose.yml (Frigate + mosquitto +
frigate-notify) with zero backup, because that file's shape didn't match
what frigate.sh's own "existing install" detection knew how to recognize.
Every service's own detection is a judgment call about what counts as
"already installed" and can miss a real setup built outside this repo's
conventions.
lib/common.sh gains backup_if_exists(FILE) — copies FILE to
FILE.bak.<timestamp> if it exists, no-ops otherwise (including DRY_RUN).
Applied before every service's own `cat > docker-compose.yml`/`cat > .env`
write across all 60 services that do one (115 call sites), plus a matching
standalone-mode stub added to every service's own bootstrap block, same
convention already used for port_in_use/find_free_port. This doesn't
replace a service's own update/fresh-reinstall detection — it's the safety
net underneath it, so a wrong detection costs a .bak file to restore from
instead of the original silently disappearing.
Also fixes the actual gap that surfaced this: services/frigate.sh's
Authelia offer only checked for Authelia installed locally on Frigate's
own box, which is never true for a dedicated NVR box with no local Caddy
either (the common shape — Caddy lives elsewhere, snippet-generation mode
already handles that). Now offers Authelia protection unconditionally and,
when Authelia isn't local, asks whether it lives on the same machine as
Caddy (still "import authelia", since that's local to wherever Caddy ends
up) or on a genuinely separate third machine (the explicit
header-pinned forward_auth form, per CLAUDE.md's "forward_auth to a remote
Authelia" note, needed because a bare authelia:9091 shortcut only works
one hop).
_authelia_bulk_assign_group() (menu option 17) picks several users and one
target group in a single step, repeatable for multiple batches in one
visit (e.g. "1 4 5 6" -> internal, then "2 3 7 8" -> external1) — the
missing third combination alongside the existing per-user (option 6) and
per-group (option 16) toggles, which only handle one user or one group at
a time respectively. "Internal" clears every outside-access group instead
of assigning one, since internal access is the absence of a group.
_authelia_describe_user_access() is a new shared one-line summary (admin /
internal / group names) used both here and in edit_authelia_user()'s own
listing, so current access is visible right where you're about to change
it instead of requiring a separate trip to option 15's report.
Verified end-to-end against a mock users.yml: batch 1 correctly cleared
an existing group from 4 users, batch 2 correctly added a brand-new group
to a different 4, with the listing reflecting each change before the next
batch starts.
Pure wording change, no behavior difference — internal already meant
exactly this (any registered Authelia user, no group) before the rename.
Also brings CLAUDE.md's description of the outside-access/admin-bypass
feature up to date; it still described the pre-generalization one-group-
per-service shape from earlier in this branch.
Printed "hub Settings -> Auth providers", which doesn't exist. The real
location is PocketBase's own admin panel underneath the hub
(/_/#/settings -> unhide collection edit controls -> edit the "users"
collection -> Options tab -> OAuth2), confirmed against beszel.dev's
OAuth guide directly. Fixed in both beszel.sh's own offer and authelia.sh's
generic OIDC menu preset.
Confirmed live: offering DISABLE_PASSWORD_AUTH/ALLOW_PASSWORD_LOGIN in the
same breath as printing the Authelia paste-in values lets an admin say yes
before actually pasting those values into the app's own settings and
testing the button — leaving neither login path working (password form
gone, OAuth provider never actually finished on the app's side).
Both are now their own function, only reachable on a later run (Beszel:
independently after the SSO offer; Mealie: from the "already configured,
not reconfiguring" branch), and gated behind an explicit "have you already
logged in successfully via the Authelia button?" confirmation before the
disable prompt is even offered.
Group membership was previously only editable per-user (option 4's user
menu, option 6 toggles that one user's groups) — no way to pick a group
and see/toggle its members directly. Adds the reverse entry point: pick
a group, then toggle which users are in it. Same _authelia_toggle_group()
underneath, just entered from the other direction.
_authelia_provision_oidc_client gains an optional PKCE flag (new 5th
positional arg; every existing caller updated to pass "n", producing an
identical client block to before) — Audiobookshelf and Beszel's own
Authelia integration docs both require require_pkce/pkce_challenge_method,
which Authelia doesn't turn on by default.
immich.sh: _immich_offer_authelia_oidc() is real server-side automation,
not just paste-in instructions — confirmed the exact system-config "oauth"
JSON field names against Immich's own config-file.md and source (not
guessed, closing out the "needs one more verification pass" note this
repo's own CLAUDE.md already had on file). GET/PUT exchange the whole
config object, so it round-trips everything else unchanged. Needs an
admin API key that doesn't exist until first web-UI visit, so it's wired
into both the fresh-install path and the "update" rerun path.
audiobookshelf.sh, beszel.sh: both apps' OIDC config is UI-only (checked
against audiobookshelf.org and beszel.dev directly — no config API or env
var for the provider fields), so their new offers automate the Authelia
side and print exact paste-in values. Beszel also gets a real, separate
DISABLE_PASSWORD_AUTH/USER_CREATION toggle to fully replace its login,
gated behind a warning to register a working account first.
Also adds Audiobookshelf and Beszel as presets in authelia.sh's own
generic "Register another app" menu, and updates CLAUDE.md's OIDC
verification table to match reality (Immich now wired, Audiobookshelf
was wrongly listed as "high-confidence no", Beszel added).
The old note pointed at an OPENAI_MODEL env var for Mealie's "import
recipe from photo" feature. Checked against docs.mealie.io directly:
Mealie moved AI provider config off env vars entirely — it's a live
Group Settings > AI Providers UI setting now (base_url/api_key/model,
with a separate toggle for which provider handles image recognition).
Also spells out how to actually reach this stack's Ollama from Mealie's
separate compose project (host-published port, not a shared network).
authelia.sh: menu option 15 lists every outside-access group with its
site membership (from access_control.rules, excluding each group's own
deny-elsewhere rule) and user membership (from users.yml) in one place —
previously only visible by grepping both files by hand.
mealie.sh: _mealie_offer_authelia_oidc now offers to set
ALLOW_PASSWORD_LOGIN=false (hides Mealie's own login form) and
OIDC_AUTO_REDIRECT=true (skip the login page, go straight to Authelia),
both confirmed against docs.mealie.io rather than assumed. Off by
default since it's a real access-control change, not just an additive
SSO button — anyone without an Authelia account loses their login path.
_authelia_scope_access previously derived a throwaway "<service>-only" group
every time it ran, so scoping two different sites to the same set of people
meant either duplicating membership by hand or hitting a false "already
scoped" early-return that silently skipped adding the second site's own
rule. Now it offers existing groups by number (any site can join one), lets
a new name be typed freely (e.g. "customer1"), and the already-scoped check
is keyed to the (domain, group) pair instead of the group name alone.
Reframes the access question as native (default, unrestricted) vs. outside
access (a named group) per the AD-style users/groups mental model, and adds
menu option 14 to rename an existing group everywhere it's referenced
(access_control.rules subjects + every member's users.yml entry). The
"-only" suffix stays internal only — every other function that already
keys off it (reporting, per-user group toggle, unprotect cleanup) is
untouched.
Adds a "subject: group:admins" rule ahead of every domain's other rules, so
admins always match first regardless of any per-service scoping (existing or
future) on that domain — a group's deny-elsewhere rule can no longer catch an
admin even if they're accidentally added to that group later.
- install_authelia and add_authelia_domain bake the rule in at creation time
- _authelia_scope_access retrofits it just-in-time before inserting its own
deny-elsewhere rule, and anchors that rule below it instead of at the top
- new menu option 13 (_authelia_ensure_admin_access_everywhere) backfills it
across every domain on an install that predates this
- remove_authelia_domain cleans the rule up too when a domain is removed,
and its domain picker dedupes since two rules now share one domain string
Lets accounts (users.yml, portable argon2id hashes included) and 2FA/session
state (data/db.sqlite3 + the storage_secret needed to decrypt it) round-trip
through a reinstall without resetting passwords or forcing everyone to
re-enroll their authenticator.
install_authelia()'s and add_authelia_domain()'s "subdomain for the
login portal" prompts concatenated whatever was typed directly with
the apex domain (AUTHELIA_PORTAL_SUBDOMAIN + "." + AUTHELIA_DOMAIN),
with no guard against someone typing the full portal domain they
actually want (e.g. "authelia.mydomain.com") instead of just the
subdomain label ("authelia"). That produces a silently broken,
doubled hostname like "authelia.mydomain.com.mydomain.com" -- which
never matches a real request, so Caddy falls through to some default
response instead of ever reaching real Authelia policy evaluation.
Confirmed live: this is exactly what happened on a real box, and
explains a much bigger symptom than the obviously-wrong hostname alone
would suggest -- every forward_auth-gated site on the instance
silently bypassed Authelia entirely, not just requests to the portal
itself, since the forward_auth subrequest to the (wrong) portal URL
never got a real answer either.
Both prompts now detect and strip an accidentally-included apex suffix
(with a one-line notice), and fall back to "auth" if someone enters
the bare apex domain itself (which can't work as the portal -- it
would collide with the wildcard rule protecting every other domain).
Verified against the exact doubled-domain input, a bare-apex input,
and two ordinary short-label inputs before shipping.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Two changes:
1. Fixed several places that still described the login portal as
literally "auth.<domain>" in user-facing text, even though the
actual subdomain has been prompt-configurable since the last
session's fix (AUTHELIA_PORTAL_SUBDOMAIN) -- the prose just never
caught up. add_authelia_domain()'s intro, install_authelia()'s
generated README, and remove_authelia_domain()'s note now describe
the portal as "you'll pick the subdomain" instead of asserting a
fixed prefix that was no longer true.
2. Every menu in this file now uses a consistent 0-to-exit/cancel
convention instead of each one doing its own thing (a numbered
"leave as-is" as the highest number, blank-to-cancel, no cancel
option at all, etc.):
- Top-level "Authelia already exists" menu: "Leave as-is" moved
from option 12 to 0 (still the default).
- _authelia_add_oidc_client's app-choice menu: added explicit
"0) Cancel" (previously a blank Enter silently defaulted to
"Other/custom app" -- surprising, now it cancels instead).
- _authelia_manage_one_user's per-user action menu: "Done" moved
from 8 to 0.
- edit_authelia_user's user-selection list and its service-group
toggle sub-list: "blank to cancel" became "0 (or blank) to
cancel", explicit and documented instead of implicit.
- _authelia_protect_site / _authelia_unprotect_site: added
explicit 0-to-cancel (previously a literal "0" typed would have
been treated as a domain name, not a cancel).
- _authelia_remove_oidc_client_menu: same explicit 0, default
changed from blank to "0".
- remove_authelia_domain: was free-text domain entry against an
unnumbered list; now a proper numbered list with 0 to cancel,
consistent with every other domain/site picker in this file.
- _authelia_scope_access: renumbered so "0) Any Authelia user"
(the safe no-op default) takes the 0 slot, "1) Specific users
only" is the one real choice.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Every invocation of setup.sh (bare, --list/--status, or with a service
name) now ensures /usr/local/bin/post-install exists and execs this
checkout's setup.sh by its real resolved path -- idempotent (only
writes when missing or pointing somewhere else) and silent except for
a one-line notice the first time it's actually created. Previously
this required cd'ing into the repo (or a manually-created wrapper) on
every box separately; now it's automatic on first run, no separate
setup step.
A plain symlink wouldn't have worked here: setup.sh finds its own
directory via ${BASH_SOURCE[0]}, which bash doesn't resolve through
symlinks, so a symlinked invocation would have set HERE to the
symlink's own directory instead of the repo's.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
_authelia_add_oidc_client()'s "what domain is this app on" prompt only
ever took typed text (with a guessed SITE_DOMAIN-based default) even
though _authelia_protect_site already offered a numbered pick-from-
Caddy-or-type-a-domain UX for the equivalent question elsewhere in
this same file -- an inconsistency a user flagged directly after
registering Mealie's OIDC client and getting a plain text prompt where
they expected the same numbered list.
Factored the shared part into _authelia_pick_domain(): lists this
box's local Caddy sites by number, or accepts a typed domain
(including one not on this box's Caddy at all). Echoes the chosen
domain on stdout with the listing itself on stderr, verified separable
under $(...) capture before wiring it in. Used now by the OIDC domain
prompt; _authelia_protect_site/_authelia_unprotect_site keep their own
inline listing since they additionally annotate each site's current
protection status, which this shared version doesn't need to know
about.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Three related fixes so undoing an Authelia change never requires
hand-editing configuration.yml or the Caddyfile:
- _authelia_add_oidc_client() had its own redundant duplicate-client-ID
check that dead-ended with "pick a different app, or edit that entry
by hand" -- even though _authelia_provision_oidc_client (called a
few lines later in the same function) already handles that exact
case safely by replacing the stale registration. Removed the
redundant check; the flow now always reaches the safe path. This was
the actual blocker in the reported "client with ID 'actualbudget' is
already registered" error -- re-registering the same app a second
time was never actually broken, just gated by dead code.
- New option 6, _authelia_remove_oidc_client_menu(): lists registered
OIDC clients by ID and name, removes one via the existing internal
_authelia_remove_oidc_client() helper (previously only reachable
from the replace-on-duplicate path, never exposed directly).
- New option 11, _authelia_unprotect_site(): reverse of option 10
(_authelia_protect_site). Removes a local site's "import authelia"
or forward_auth block from its own Caddy block and reloads Caddy;
for a domain on a different box's Caddy, cleans up its access-
scoping rules here (the actual gate needs removing on that box by
hand, same one-way limitation option 10 already has in reverse).
Also removes any _authelia_scope_access rules for the domain, found
by the same "<domain>-only" group name convention, verified against
a synthetic multi-domain configuration.yml before shipping so an
unrelated domain's rules sharing the same "*.<apex>" line are left
untouched. Both the local-block removal (import authelia one-liner
and multi-line forward_auth block shapes) and the access-rule
removal were tested against realistic fixtures first.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Every individual service so far offers its own "Protect X with
Authelia SSO?" prompt on install, but there was no way to gate an
arbitrary existing site from Authelia's own menu -- especially useful
for a site on a DIFFERENT box's Caddy than the one Authelia runs on,
this repo's own recurring case (a DigitalOcean droplet's site,
protected by an Authelia instance on a separate IONOS box).
New option 9, _authelia_protect_site(): lists this box's own local
Caddy sites by number (flagging ones already protected), or accepts a
typed domain that isn't on this box's Caddy at all. A local site gets
"import authelia" inserted as the first line of its existing block --
before reverse_proxy, same ordering rule as everywhere else in this
codebase, since Caddy runs directives in the order written and an auth
check after reverse_proxy never runs at all. A remote site can't be
edited from here, so it prints (and saves to caddy-snippets/) the
remote-hop-safe forward_auth block that box's own Caddyfile needs
instead, with the portal's actual domain read back from
configuration.yml rather than assumed. Either way finishes by calling
_authelia_scope_access for the domain, so protecting a site and
restricting who can reach it happen in one pass.
Verified the site-listing regex, insertion, idempotency detection, and
remote-domain/portal lookup against synthetic Caddyfile/configuration.yml
fixtures before shipping.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
_authelia_scope_access() (services/authelia.sh) already works for any
Authelia-protected service, forward_auth-gated or OIDC alike -- it's
just never been called from security-dashboard.sh on either the
domain-takeover path (_secdash_offer_asterisk_domain) or the plain
separate-subdomain path, so every domain this dashboard ever protected
defaulted to "any Authelia user", with no way to restrict it to
specific people. That's why the Authelia menu's "Promote to a
specific service's access group" reported no scoped groups existing
yet even after protecting this dashboard with Authelia.
_secdash_configure_caddy() now offers scoping right after the domain
is Authelia-protected (guarded on EXTRA_BLOCK being non-empty, so
Basic-Auth-only or no-auth setups aren't offered a scoping question
for a gate that doesn't exist), guarded by declare -F for standalone
runs where authelia.sh was never sourced. Runs whether the Caddy block
was just freshly written or already existed, so re-running the
installer on an already-configured domain still offers it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
install_authelia() and add_authelia_domain() both hardcoded "auth." as
the login portal's subdomain prefix everywhere -- configuration.yml's
authelia_url, the Caddy portal block/domain, generated README/OIDC
text. No prompt ever offered anything else, despite this repo
otherwise treating "auth.<domain>" as just this one instance's own
choice, not a protocol requirement.
Both now prompt for the portal subdomain (default "auth", so existing
behavior is unchanged for anyone who doesn't care) and use the actual
chosen value throughout. Every function that operates on an EXISTING
domain (remove_authelia_domain, _authelia_add_oidc_client) now reads
the real portal domain back from that domain's own session.cookies
authelia_url entry instead of assuming "auth.<domain>" -- matching the
same read-back pattern _authelia_add_oidc_client already used for the
apex domain itself, and _authelia_provision_oidc_client already used
for the portal URL. _authelia_remove_caddy_portal_block now takes the
portal's full domain directly rather than reconstructing it, so
removing a domain whose portal used a custom prefix actually finds and
removes the right Caddy block.
Also fixed a real, separate small bug found while in here: the
primary portal's Caddy log path was hardcoded to a generic auth.log
(collides across instances/domains) instead of following every other
site block's own <domain>.log convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
_asterisk_configure_caddy_public() used to reverse-proxy Asterisk's own
web admin at its public domain, optionally gated by local or remote
Authelia, flipping WEB_ADMIN_AUTH_DISABLED=true in .env to hand auth
off to it. That coupling was the root cause of a real live exposure:
a box where Authelia protection was accepted once, but the Authelia
import/forward_auth block itself later went missing from the Caddyfile
(e.g. lost on a restore), was left with the web admin's own login off
and nothing else gating it -- extension/device data reachable with no
password at all. A remote Authelia's forward_auth also proved fragile
in practice for something that only ever needed to keep a domain's
cert alive (DNS/routing/access-rule mismatches spanning two boxes,
hard to diagnose from either one alone).
This domain now just serves a minimal keep-alive page (a bare "OK" 200
response) so Caddy can still issue/renew the SIP TLS cert -- cert
issuance only needs Caddy to own the site block, it's unrelated to
what the block serves. Auth is now optional Basic Auth handled
entirely inside Caddy itself, no external subrequest, so it can't fail
this way. Asterisk's own web admin is no longer exposed publicly by
this function at all -- reachable only via the CLI:
docker exec -it <container> easy-asterisk
The Security Dashboard's own domain-takeover offer
(_secdash_offer_asterisk_domain in security-dashboard.sh) is the
supported way to put something meaningful on this domain instead.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
add_authelia_domain() (menu option 1) had no reverse operation -- once
a domain was added there was no way to undo it short of hand-editing
configuration.yml and the Caddyfile. remove_authelia_domain() (new
option 2) does the reverse cleanly: removes the access_control.rules
entry, the session.cookies entry, and the auth.<domain> Caddy portal
block for one domain, verified against a synthetic multi-domain
configuration.yml before shipping. Warns loudly that any service still
pointed at the removed domain will stop authenticating, and requires
confirmation before touching anything.
The menu's own text now also flags the likely real mistake this
surfaces: adding a domain that's actually just a SUBDOMAIN of an apex
already on the instance creates a *.subdomain.apex wildcard rule that
doesn't match the bare subdomain itself, plus a redundant separate
auth.subdomain.apex portal -- when the subdomain was already covered
by the existing apex's own wildcard rule and portal all along.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Asterisk's own web admin was only Caddy-fronted at its own domain so
Caddy could issue it a trusted TLS cert for SIP -- cert issuance only
needs Caddy to own that domain's site block, it's unrelated to what
reverse_proxy target the block forwards to. _asterisk_configure_caddy_public
also never rewrites an existing site block on a repeat run, so a box
where WEB_ADMIN_AUTH_DISABLED got set true (from an earlier "protect
with Authelia" answer) but the Authelia import itself never landed or
got lost on a restore was stuck silently unauthenticated with no
reconfigure path ever revisiting it -- confirmed live: a real box was
found exposing its extensions/device list with no login at all.
_secdash_offer_asterisk_domain() now offers, whenever Security
Dashboard's Caddy setup runs (fresh install or reconfigure) and
Asterisk already has a public domain, to serve the dashboard there
instead of a separate subdomain: removes Asterisk's old site block for
that domain, rebuilds it fresh under this dashboard's own (always-on)
Authelia gate, and re-enables Asterisk's own web admin login in .env
as defense-in-depth now that its port isn't published at all.
Declining falls through to the normal separate-domain prompt
unchanged.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
New installs now name the container "asterisk", matching every other
service's container_name == service name convention, instead of
reusing the vendored easy-asterisk CLI tool's own name (which stays
/usr/local/bin/easy-asterisk inside the container, unrelated and
unchanged).
An existing "easy-asterisk" install is never silently renamed: every
place that resolves the container name (_asterisk_resolve_layout in
asterisk.sh, plus the duplicated copies in security-dashboard.sh,
sms-inbound.sh, pstn-trunk.sh, and tools/pstn-test-check.sh's docker ps
detection) now reads it from the box's own docker-compose.yml instead
of assuming it, falling back to "asterisk" only when there's no
existing install to read. Migrating a live box to the new name is a
one-time manual action (edit docker-compose.yml's container_name for
Asterisk and its coturn sidecar, docker compose down + up -d); every
sibling service then picks it up automatically on its next run.
The DigitalOcean-droplet layout (asterisk-digital-ocean directory,
easy-asterisk-do container) is untouched by this - that naming stays
exactly as documented for pre-merge droplet installs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
keep_alive_interval is a type=global pjsip.conf option, not a
type=transport option -- it never existed on [transport-tls] on any
Asterisk version. The IONOS TLS-keepalive mitigation was inserting it
there, which made sorcery reject the whole transport-tls object
("Could not find option suitable for category 'transport-tls' named
'keep_alive_interval'"), silently killing TLS SIP entirely instead of
just adding a keepalive.
Both _asterisk_patch_keepalive_vendor_files (deployed vendor copies)
and _asterisk_ensure_live_keepalive (live pjsip.conf) now target
[global]/type=global, and both self-heal a box that already picked up
the bad placement by removing it from [transport-tls] first.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
The stack health check's Security Dashboard/sms-inbound/ntfy checks only
had visibility into this box's own local Caddyfile — but all three can
legitimately be fronted by a Caddy (and Authelia) on a completely
different box instead, the same remote-Caddy pattern sms-inbound.sh and
ntfy.sh's own installers already support via CADDY_MODE/CADDY_REMOTE_HOST.
A site explicitly configured for remote Caddy was getting a false "Caddy
has no site block for it" for each of them, with a fix offer that would
have been actively wrong: adding a redundant local Caddy block for
something deliberately fronted elsewhere.
Now resolves the same site-wide CADDY_MODE the affected services'
installers themselves use before treating "not found locally" as a real
issue — only counts it, and only offers a fix, when the site is actually
in local Caddy mode. Remote (or no-Caddy) mode gets a plain informational
line instead: not wrong, just not something this box can verify.
Verified: default/local mode still flags a genuinely unwired dashboard as
an issue with a fix prompt; CADDY_MODE=remote (even with a local Caddy
directory also present) correctly downgrades the same finding to
informational with no prompt and no issue counted.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
"update" mode's whole promise is leaving already-configured things
alone — but that assumption silently breaks when something was
configured but never fully wired up, and update never re-asks the
questions that would reveal it. This session hit three separate
instances of exactly that on one droplet revert: a domain set with no
Caddy block, a Caddy block with no synced TLS cert (transport-tls fails
to bind — "Unable to retrieve PJSIP transport 'transport-tls'", breaking
every call), and a baked-in external IP left over from before the box
moved. sms-inbound and (potentially) Security Dashboard/ntfy can have
the identical "domain set, nothing serving it" gap with no way to
discover it either, since their own update modes don't re-ask.
_asterisk_run_stack_health_check(), called every "update", replaces the
narrower Caddy-only check added last time:
- Compares pjsip.conf's baked external_signaling_address against this
box's actual current public IP; offers to rewrite it and restart.
- Checks Asterisk's own DOMAIN_NAME has both a Caddy site block and a
matching TLS cert in the container; offers to fix each independently.
- Checks Security Dashboard / sms-inbound / ntfy (whichever are
installed) for a matching Caddy site block, via a new lib/common.sh
helper (caddy_domain_for_upstream) that finds the block without
needing to already know the domain — none of these three services
persist it anywhere. Points at that service's own "Full reinstall"
(the only mode that re-asks) since fixing their config isn't this
file's to script.
The cert-sync fix needed a non-interactive hook into the vendored
easy-asterisk CLI, which only exposed it as an interactive menu item
(Server Settings -> Force re-sync Caddy certs). Added a --sync-caddy-cert
flag via _asterisk_patch_cert_sync_cli(), patching the deployed vendor
copy the same way _asterisk_patch_voicemail_vendor_files and friends
already do — never vendor/ in git.
Every check runs unconditionally (never opt-in, so a gap is never missed
by nobody thinking to ask); every fix is individually opt-in and named
as a real config change, unlike the rest of "update"'s no-side-effects
default.
Also factored the DO-metadata/ifconfig.me/hostname-I public-IP detection
chain (previously duplicated 3 times) into _asterisk_current_public_ip().
Verified: caddy_domain_for_upstream against a multi-block Caddyfile
(distinguishes same-prefix upstreams correctly); the full health check
against fake docker/curl across every combination (all wired, IP
mismatch declined/accepted, cert mismatch declined/accepted, dashboard
unwired, sms-inbound wired vs. placeholder-domain, multi-instance ntfy
with one wired and one not); and _asterisk_patch_cert_sync_cli's
idempotency + resulting syntax against a real copy of the vendor script.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
"update" mode never re-asks the domain/networking/Caddy questions, on the
assumption there's already Caddy/Authelia config in place to leave alone.
That assumption breaks for an install where a domain was set at some
point (DOMAIN_NAME in .env) but Caddy never actually got a site block for
it — declined at install time, DNS wasn't ready yet, or Caddy was
reinstalled/reset separately since. Previously the only way back was a
full reinstall, which re-generates a dedicated coturn container with new
TURN credentials (every already-configured phone needs its QR re-scanned)
— a lot of blast radius just to add one missing Caddy block, and enough
that reaching for it risks the extensions/voicemail data a "fresh"
reinstall can also wipe if the wrong prompt is answered.
"update" mode now detects this specific gap (domain set, no matching
Caddyfile block) and offers to run _asterisk_configure_caddy_public()
right there — the same function "fresh" installs use, but it only ever
touches the Caddyfile and .env's WEB_ADMIN_AUTH_DISABLED line, never
coturn/extensions/anything else "update" already promises not to touch.
Verified in isolation: offers and calls the fix when the domain is set
with no matching Caddyfile block, stays silent when a block already
exists.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Confirmed live: neither this repo nor the vendored easy-asterisk script
ever sets a PJSIP endpoint's `mailboxes=` field. add_device()'s own
device_config template never writes it, and write_voicemail() only ever
touched voicemail.conf — so recording a voicemail worked fine
(voicemail.conf + the dialplan's VoiceMail() call), but no phone ever
actually subscribed to be told about it, regardless of whether the
voicemail flag was on. Matches the exact symptom of "voicemail records
fine, but no notice comes up on the phone."
Add _ea_set_endpoint_mailboxes(), called from write_voicemail(): adds/
updates mailboxes=<ext>@default in that extension's PJSIP endpoint stanza
when voicemail is enabled, removes it when disabled, and reloads
res_pjsip so it takes effect immediately. Bounded to just the
type=endpoint stanza (pjsip.conf reuses the same [ext] bracket name for
type=endpoint/type=auth/type=aor) the same way lib/common.sh's
_remove_caddy_site_block is bounded for Caddy blocks — verified against a
two-device pjsip.conf that editing one extension's mailboxes= never
touches its own auth/aor stanzas or another extension's stanzas, that a
repeat enable doesn't duplicate the line, and that disabling removes it
cleanly.
Existing extensions with voicemail already enabled won't get this
retroactively — the Extensions tab's voicemail toggle has to actually run
again (off then back on) to apply it, since this only fires on the
enabled/disabled transition itself.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
When the "Public domain for the webhook" prompt was left blank (DNS not
ready yet, or just missed), the installer built the Forward-to-URL as
literal https://<your-domain>/sms/... and persisted that placeholder to
settings.env as if it were real. "Update" mode never re-prompts for the
domain (by design — it's meant to leave already-configured settings
alone), so every later re-run silently re-served the same unusable
placeholder, with nothing indicating anything was wrong. A DID provider
(Anveo) correctly rejects it — it isn't a resolvable hostname.
- Only build FORWARD_URL when a real domain was entered; leave it empty
otherwise instead of substituting the placeholder.
- Fresh-install summary and README now say plainly that setup isn't
complete and how to finish it, instead of printing an empty/bogus URL.
- Update-mode now detects a missing/placeholder domain and tells you to
re-run with "Full reinstall" to be asked again, instead of reporting
success with a broken URL.
Verified with a direct test of _sms_write_readme() and the FORWARD_URL
construction for both the blank- and real-domain cases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Follow-up to Frigate's Authelia integration: both of these can also skip
their own login entirely once Authelia is doing the gating, each with a
different trust model appropriate to what the app actually supports.
- gitea: new _gitea_offer_reverse_proxy_auth(), a second Authelia
integration alongside the existing OIDC "Sign in with Authelia" button.
Enables Gitea's own ENABLE_REVERSE_PROXY_AUTHENTICATION so it auto-logs
in from a trusted Remote-User header — no click, no separate Gitea
session to expire on its own. Trust is IP-range based
(REVERSE_PROXY_TRUSTED_PROXIES), computed from caddy_net's real subnet
the same way ufw_allow_from_caddy_net does; refuses to enable the
feature at all if that can't be determined rather than fall back to a
permissive default — Gitea's own Docker image has shipped an unscoped
default before (GHSA-f75j-4cw6-rmx4, any IP could impersonate any user).
Rewires Gitea onto caddy_net and re-points Caddy at gitea:3000, since it
previously only reached Caddy via its published host port. Gitea's own
login stays available as a fallback, so unlike Frigate there's no
"native login off with nothing gating it" state to guard against.
- uptimekuma: sets DISABLE_AUTH=true only once Caddy's "import authelia"
gate is confirmed in front of it. Uptime Kuma already joined caddy_net
unconditionally, so this only needed the env var plus moving the
Authelia-gated Caddy call earlier (before docker-compose.yml is
written); the existing unconditional call at the end now only runs as a
fallback when the Authelia path wasn't used or wasn't completed. Kuma's
DISABLE_AUTH has no IP-scoping or secret check left once set — the
strictest of the three to get the ordering right on, since a mistake
here means wide open, not just spoofable.
Verified with a local test harness (fake Authelia/Caddy/docker-network
state): both the happy path and the "Caddy declined" safety fallback
produce the expected docker-compose.yml/.env/Caddyfile output for each
service, and Gitea's subnet-detection refusal + idempotent-rerun guard
were exercised directly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
Frigate has its own built-in login separate from Authelia's session, so
just adding `import authelia` in front of it (the pattern used for
no-built-in-auth services) would leave two independent logins stacked,
defeating the point of Authelia's "remember me" on mobile. Frigate has a
`proxy` auth mode built for exactly this — trust Remote-User/Remote-Groups
from an upstream forward_auth proxy and disable its own login entirely.
- Extend configure_caddy_for_service() with an optional 5th arg for
sub-directives inside the reverse_proxy block itself (header_up), needed
to pin an X-Proxy-Secret header so Frigate's proxy-auth trust can't be
spoofed by a request reaching its published port directly, bypassing
Caddy/Authelia. Backward compatible — every other caller is unaffected.
- services/frigate.sh: prompt to protect with Authelia when installed;
wires import authelia + the X-Proxy-Secret header_up into Caddy, and
only writes config.yml's auth.enabled: False + proxy block once Caddy
actually confirms it's fronting the domain (never disables the native
login with nothing else gating access). Reuses the secret across
reinstalls instead of rotating it. Calls _authelia_scope_access() so
access can be restricted to specific users instead of every Authelia
account. Fixed a latent bug in the standalone-mode Caddy stub where the
auth block was placed after reverse_proxy instead of before it (dead
code — the same "Authelia never prompts" bug class CLAUDE.md documents
for the real helper).
- CLAUDE.md: document the new configure_caddy_for_service parameter and
Frigate's hybrid built-in-auth/forward_auth pattern.
Verified end-to-end against a local test harness (fake Authelia/Caddy
dirs): config.yml, .env, and the generated Caddyfile block all agree on
the shared secret and header names, auth is skipped cleanly when Caddy
isn't configured, and the secret is reused on a second run.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpKTLpwAgZNooTacWeQLuc
User feedback: wanted chromaDB/rag-server/mcp-server included in the
optional-services picker added last commit, not just Gitea/Portainer/
Kiwix/InvokeAI/ComfyUI/Aider. Confirmed against the compose file's
depends_on chain before adding: open-webui only depends_on ollama (its
OLLAMA_BASE_URL connection works standalone), so none of these three are
actually required for regular chat — only Open WebUI's separate RAG tab
(routed through rag-server) and MCP tool-calling need them. Mealie's own
Ollama usage never touches this stack at all.
Bundled chromadb+rag-server+mcp-server as one option (7), not three
separate numbers — mcp-server depends_on rag-server depends_on chromadb,
so stopping only one of the three would leave the others running against
a dead dependency instead of a clean stop. Also added a cascade for the
existing kiwix option: mcp-server depends_on kiwix too (not just
rag-server), so stopping kiwix without also stopping mcp-server has the
same problem — now handled automatically with a dedup pass in case both
the kiwix cascade and option 7 add mcp-server to the stop list.
Verified all four cases in isolation: kiwix-only correctly cascades to
mcp-server, option 7 alone stops the right three, choosing both dedupes
to one clean list, and unrelated choices (gitea/portainer) are unaffected.
User feedback: local-ai-setup.sh always brings up the entire stack
unconditionally (Gitea, Portainer, Kiwix, InvokeAI, ComfyUI, Aider,
alongside the core Ollama/Open WebUI/ChromaDB/RAG/MCP) with no way to opt
out — e.g. Gitea when you already run git elsewhere, or Portainer when you
manage Docker some other way.
Didn't touch local-ai-setup.sh's own compose generation for this (it's
vendored upstream code, and other services reference these by container
name/network in ways that would need individual auditing to make safely
conditional). Instead, added a post-install picker in the wrapper: after
the full stack starts, offer to `docker compose stop` whichever of the six
non-core services aren't wanted. Images are already pulled either way, so
anything stopped comes back later with a plain `docker compose up -d
<name>` — no reinstall needed.
Verified the choice-parsing loop in isolation: "1 4 9 3" correctly warns
on the invalid "9" and resolves to gitea/invokeai/kiwix.
Confirmed live: a bare hostname in DOMAIN= (typing "vault.example.com"
instead of "https://vault.example.com" at the install prompt — easy to do
despite the example text showing the scheme) crash-loops the container
with no clear startup error, and re-running the installer doesn't fix an
already-written .env since "update" mode deliberately never touches it.
Two changes, mirroring how the existing SMTP half-state bug is already
handled in this file:
- Normalize VW_DOMAIN at prompt time — missing scheme gets https://
prefixed automatically instead of writing it verbatim.
- New _vaultwarden_fix_domain_scheme() self-heal, called at the same two
sites as _vaultwarden_fix_smtp_halfstate() (the "update" path and the
fresh-install "start now" path), so a box that already has a scheme-less
DOMAIN self-heals on its next start instead of staying stuck.
Verified the self-heal function in isolation: vault.mydomain.com ->
https://vault.mydomain.com.
local-ai-setup.sh runs as whoever invoked this wrapper — root, since
setup.sh itself runs under sudo — so every file it generates
(docker-compose.yml, .env, requirements.txt, server.py, mcp_server.py,
pull-models.sh, start/stop/status.sh) came out root-owned. Nothing handed
that back to ACTUAL_USER unconditionally: the only existing
ensure_docker_dir_ownership call was inside the cloud-provider wiring
block, so it silently never ran at all for anyone who skipped cloud
providers.
Confirmed live: this repo's own "Skipped. Run later: cd $AS_DIR && bash
local-ai-setup.sh" message tells the user to re-run it directly later as
themselves (no sudo) — which then fails with "Permission denied" on any
file root created during the original sudo run, e.g. requirements.txt.
Same root cause class as a stray root-owned .git/FETCH_HEAD blocking a
plain `git pull` — a root-run leaving files a later unprivileged run can't
touch.
Fix: call ensure_docker_dir_ownership "$AS_DIR" unconditionally right
after the installer-run block, not only on the cloud-provider path.
Confirmed live: `docker compose pull` failed with "yaml: line 44, column
29: mapping values are not allowed in this context" during the "Starting
Stack" phase of local-ai-setup.sh. Root cause: two healthcheck blocks
(ollama, chromadb) crammed interval/timeout/retries onto one
semicolon-separated line —
interval: 30s; timeout: 10s; retries: 5
— which isn't valid YAML; a scalar value can't contain a second `key:`
token like that unless quoted. Split each into three separate properly
indented keys, matching how every other multi-key block in this same file
is written.
Verified by generating the actual docker-compose.yml via the real heredoc
(same one docker-stack.md's variables would produce) and parsing the
result with PyYAML — line 44 is exactly the fixed `interval: 30s` line,
and the full file now parses as valid YAML.
Pre-existing bug in the vendored source, unrelated to this session's
earlier ai-stack.sh/local-ai-setup.sh changes (those only touched the
pull-models.sh heredoc and the cloud-provider prompt, both well before
this point in the install) — first surfaced now because this is the first
run in this session to actually reach the "Starting Stack" step rather
than stopping earlier.
Blank already meant skip, but user feedback wanted a keystroke that says
so explicitly rather than just leaving the input empty. Added "0) Skip —
stay fully local" to the menu, updated the prompt to mention it, and
handled "0" as a silent no-op in the choice loop (previously it would
have fallen through to the "Ignoring unknown choice" warning).
The instruction was only in explanatory text a few lines above the actual
prompt (prompt_text "Cloud providers to add []:") — easy to miss once
that's scrolled past, especially since the bracketed default shows empty
but doesn't say what empty means. User feedback: the screen itself should
say it, not just text above it. Now reads "Cloud providers to add (blank =
skip, stay fully local):".
None of local-ai-setup.sh's tier-selected models (CHAT_MODEL/CODE_MODEL/
EMBED_MODEL) can read an image — there was no way to get vision support out
of this stack at all before now. Added a numbered pick-list to the
generated pull-models.sh, right after the existing DeepSeek-R1 optional
pull, matching that same read -rp pattern:
1) moondream ~1.7 GB by Moondream AI — tiny, built for
CPU-only or weak/old-GPU hardware
2) llava:7b ~4.7 GB general-purpose vision
3) qwen2.5vl:7b ~6 GB stronger accuracy, more RAM/VRAM
4) llama3.2-vision:11b ~7.9 GB heaviest of the four
moondream is the recommended default — sized for exactly the "6 vCPU, 8GB
RAM, no GPU" case this was asked for, unlike the other three which assume
real GPU/RAM headroom.
Verified by actually running the heredoc that generates pull-models.sh
(with EMBED_MODEL/CHAT_MODEL/CODE_MODEL stood in) and syntax-checking the
resulting output script, not just the source — the outer heredoc is
unquoted so $-escaping mistakes wouldn't show up as a bash -n failure on
local-ai-setup.sh itself, only on what it generates.
services/ai-stack.md gets a matching "Vision models" section (sizes, the
manual pull command, and how to point an app's OPENAI_MODEL at one).
laptop_full_setup.sh's separate, non-interactive pull-models.sh generator
is untouched — it's not invoked anywhere in this repo's own install flow
(only local-ai-setup.sh is, from install_ai-stack()), so it's out of
scope here.
Confirmed live: install_mealie() pre-computes BASE_URL as
recipes<suffix>.$SITE_DOMAIN before ever asking about Caddy, then
configure_caddy_for_service() separately prompts for a domain — which the
user can freely override (e.g. typing mealie.mydomain.com instead of
accepting the recipes.mydomain.com default). Nothing fed that choice back
into BASE_URL, so it stayed stale. Since BASE_URL is exactly what
_mealie_offer_authelia_oidc() registers as the OIDC redirect URI, this
produced Authelia's "redirect_uri does not match any of the OAuth 2.0
Client's pre-registered redirect_uris" — Caddy and DNS were both correctly
pointed at the new domain, but the client Authelia had on file still said
the old one.
Added CADDY_SERVICE_DOMAIN as a new configure_caddy_for_service() out-param
(lib/common.sh) — the same out-param convention as the existing
CADDY_SERVICE_CONFIGURED/CADDY_SERVICE_MODE, set right after the domain
prompt is accepted. install_mealie() now reconciles BASE_URL against it
immediately after the Caddy call, before the Authelia OIDC step reads
BASE_URL back out of .env. ActualBudget's equivalent OIDC offer asks for
its own domain fresh each time rather than reading a pre-computed BASE_URL,
so it isn't affected by this class of bug and needs no equivalent fix.
Verified the reconciliation logic in isolation against a synthetic .env.
Root cause of the recurring Mealie OIDC "unexpected character '/' in
variable name" failure, confirmed against the user's actual
configuration.yml byte content: this repo's own scripts write
authelia_url/domain unquoted, but YAML makes quoting optional, and a
hand-edited config can add single or double quotes around the value
(here: authelia_url: 'https://authelia.example.com.'). awk's
`print $2`/`print $3` is a naive whitespace-split token grab that doesn't
know about YAML quoting, so it captured the value WITH the literal quote
characters attached. The generated discovery URL then came out
`'https://authelia.example.com.'/.well-known/openid-configuration` —
Docker Compose's env parser closed the quoted value at that embedded
closing quote and choked on the trailing text as an invalid new token.
The earlier \r-stripping commit was a real but different fix (a
CRLF-tainted line fails to match these anchored awk patterns at all) —
it didn't cause and couldn't have fixed this. Both guards are needed and
now both apply, in both _authelia_provision_oidc_client() (domain and
portal URL) and the same latent bug in _authelia_add_oidc_client()'s
domain parse.
Verified end-to-end: reconstructed the user's exact reported byte
content (od -c dump) in a synthetic configuration.yml, ran the actual
_authelia_ensure_oidc_provider/_authelia_provision_oidc_client/
_mealie_offer_authelia_oidc functions against it (docker calls stubbed),
and confirmed the generated .env line is now a single clean line with no
embedded quotes or split.
The previous commit added OIDC_AUTHELIA_PORTAL_URL parsing but only
tr -d '\r'-sanitized the awk output, not the input. That's insufficient
for a CRLF-tainted file (confirmed live: a configuration.yml line
hand-edited by something that saves Windows line endings) — every
line-anchored awk pattern here fails to match at all against a line like
" cookies:\r", since $ anchors end-of-string and the \r is still part of
it, not just leaves a stray \r in the captured value. Symptom was Mealie's
generated OIDC_CONFIGURATION_URL line getting split mid-string, which
Docker Compose's env parser (bare \r treated as a line break too) reported
as "unexpected character '/' in variable name".
Fixed by piping the file through tr -d '\r' before awk sees it, for both
the domain and portal-URL parses. Verified against a synthetic CRLF config
that reproduces the exact failure — both fields now parse clean.
_authelia_provision_oidc_client() gained a new out-param,
OIDC_AUTHELIA_PORTAL_URL, read back from the instance's own
configuration.yml (session.cookies[].authelia_url) — the actual source of
truth for where the portal lives — instead of every caller separately
assuming "https://auth.$domain".
install_authelia() and add_authelia_domain() both still default new
instances to "auth." as before (unchanged), but that's just a default, not
a guarantee: it's plain text in configuration.yml and gets hand-edited on
some boxes (e.g. a dedicated instance renamed to "authelia." to avoid
colliding with another instance's "auth." on a different machine). Mealie,
ActualBudget, and Gitea's native-OIDC wiring all independently hardcoded
"auth." when building their discovery URL, so a renamed portal silently
produced a discovery URL pointing at a host that doesn't serve Authelia —
surfacing as an opaque 500 during the OIDC token exchange with no useful
client-side error.
Verified the new awk parse against both a default ("auth.") and a renamed
("authelia.") cookies block before trusting it.
edit_authelia_user() previously only ever let you select one user, act on
them, and then returned all the way out of install_authelia() (which calls
it with an immediate `return 0`) — deleting several users meant re-running
`sudo ./setup.sh authelia` and re-navigating to option 3 from scratch for
every single one.
Restructured: the per-user action menu (edit/reset-password/2FA/admin/
service-access/delete) is now _authelia_manage_one_user(), and
edit_authelia_user() drives it in a loop — numbered multi-select up front
("2 4" deletes/edits both), then "Manage more users?" to go again with a
freshly re-read user list instead of exiting. Guards against acting on a
user who was already deleted earlier in the same batch.
Verified against a synthetic users.yml: selecting two users by number and
deleting both in one pass removes exactly those two, leaves the others
untouched.
tr -cs 'a-z0-9_-' '-' only allowed lowercase letters, so any uppercase
leading character (e.g. "Bob") got converted to a dash and then stripped
by the paired leading-dash sed, silently truncating the username. Widened
to a-zA-Z0-9_- in both add_authelia_user() and _authelia_scope_access().
Also extends the "Manage an existing user" menu (still numbered-selection
throughout) with:
- option 6: toggle a user's membership in any existing "<service>-only"
scoped-access group, picked by number, via two new helpers
(_authelia_list_scoped_groups, and re-resolving the user's line range
before each toggle since a prior toggle in the same pass shifts it)
- option 7: delete a user outright (_authelia_delete_user_block), with confirmation
Verified against a synthetic users.yml (add/remove toggling across
multiple groups, block deletion, uppercase-username round-trip) before
touching the live file.
Confirmed live: re-running ActualBudget's Update path produced zero
output for the Authelia SSO step — no prompt, no message, straight back
to the shell. Root cause: the idempotency guards in
_actualbudget_offer_authelia_oidc / _mealie_offer_authelia_oidc /
_gitea_offer_actions_runner were plain `grep -q ... && return 0` — silent
by construction. Indistinguishable from the step not running at all,
which is exactly what it looked like.
ActualBudget and Mealie's OIDC offers now explain what they found and
ask whether to reconfigure (registers a fresh Authelia client + secret,
clearing the old env vars first) instead of silently bailing. Gitea's
Actions-runner offer explains what it found and how to check its status
(reconfiguring that one means editing a docker-compose service block,
not just a couple of env vars, so it just informs rather than offering
to redo it).
Also: _authelia_scope_access now shows existing Authelia usernames as a
numbered list before asking who should have access — picking by number
works alongside typing new names directly (mix freely, e.g. "1 3
newperson"), rather than requiring exact usernames typed from memory
with no reference and no protection against a typo silently creating a
duplicate account.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
_authelia_scope_access asked for usernames to grant access to a service
without ever showing who already exists — confirmed live, the prompt
just showed a blank "Usernames:" line with nothing to reference. A typo
against an existing name doesn't fail or warn, it silently creates a new,
separate account instead of matching the intended one.
Now lists existing Authelia users (reusing _authelia_list_usernames,
already used elsewhere in this file) right before the prompt, and warns
about the typo/duplicate-account risk explicitly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Fixes two things found while answering a question about staying logged
in across every Authelia-protected service:
1. CLAUDE.md's own "stay logged in" instructions referenced
remember_me_duration — renamed to remember_me in Authelia 4.38, this
repo pins 4.39.20. Authelia doesn't error on an unknown key, it just
silently ignores it, so following that guidance as written would have
done nothing. install_authelia() itself already uses the correct
`remember_me` key at install time (default 7d) and was never affected
— only the hand-edit instructions in the docs were stale.
2. There was no way to change it afterward without hand-editing the file,
contrary to this repo's own "no manual config editing" direction.
Added _authelia_set_remember_me() (new menu option 7): prompts for a
new duration (12h/7d/1M/1y/-1 to disable), writes it, restarts.
Tested the sed replacement against a synthetic session block before
trusting it on real config. Also documented clearly (both in the
function's own prompt and in CLAUDE.md) that this only controls
Authelia's own session — a native-OIDC app's own session/token expires
on its own separate schedule, which this setting doesn't touch.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Confirmed live: ActualBudget's new automated "Sign in with Authelia" offer
hit a client_id ("actualbudget") already registered from an earlier use of
the interactive "Register an app" menu — that older flow only registers
the client in Authelia and prints instructions to paste the secret into
the app's own settings manually; if that paste step never happened,
ActualBudget's .env never got the OIDC vars, but Authelia still considered
the client_id taken. _authelia_provision_oidc_client's duplicate check
just warned and returned 1, permanently blocking the automated offer with
no path forward — the stale registration's secret was shown once and
already gone, so there was nothing to recover, only reasons to replace it.
Added _authelia_remove_oidc_client() (tested against a synthetic
multi-client config, both mid-list and last-in-list removal) and changed
the duplicate-client_id check to remove-and-replace instead of failing.
Every automated caller (gitea/mealie/actualbudget's SSO offers) uses a
fixed, service-specific client_id, so a collision there means "this same
service was already registered," not a different app's ID being
clobbered. The interactive menu's own earlier duplicate check (a distinct
code path, one step before this one) is untouched — it still warns and
stops before prompting further, since a human-typed ID colliding with an
unrelated app is a different, more ambiguous situation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Researched which of the "has built-in auth" services actually support
native OIDC before wiring anything in (checked live docs, not assumed) —
two services turned out to contradict general assumption: Portainer's
OAuth/OIDC is Business Edition only (this repo installs portainer-ce, which
doesn't have it), and ntfy has no auth-oauth2-* support at all despite it
seeming like the kind of thing a modern self-hosted tool would have added
by now. Full findings recorded in CLAUDE.md so this doesn't need
re-researching.
Two real, verified wins wired up, both entirely env-var driven — no
manual config file editing, matching this repo's "no manual wizard"
philosophy and reusing the exact _authelia_provision_oidc_client /
_authelia_scope_access machinery already built for Gitea:
- mealie: OIDC_AUTH_ENABLED/OIDC_CLIENT_ID/OIDC_CLIENT_SECRET/
OIDC_CONFIGURATION_URL appended to the existing .env (env_file: .env is
already how mealie.sh's compose reads it). Also adds a
--forwarded-allow-ips entrypoint override when Caddy-fronted — confirmed
against Mealie's own issue tracker that without it, the generated OIDC
redirect URI comes out http:// even when actually served over https://,
which providers reject as a scheme mismatch.
- actualbudget: ACTUAL_OPENID_DISCOVERY_URL/CLIENT_ID/CLIENT_SECRET/
SERVER_HOSTNAME, same pattern. Redirect path (/openid/callback) matches
the preset already used by authelia.sh's own "Register an app" menu for
this same service.
Both offered on fresh installs and Update reruns, default no, and both
call _authelia_scope_access() afterward so access can be restricted to
specific users instead of every Authelia user, same as Gitea.
Immich has real OIDC + a system-config API but needs one more
verification pass on the exact request payload before automating — not
guessing that part. Jellyfin and Home Assistant only have third-party
plugin/HACS-based OIDC, a bigger lift than an env-var toggle — noted but
not attempted this pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Every domain with an access_control rule was reachable by any Authelia
user by default (the existing catch-all *.${AUTHELIA_DOMAIN} rule) — no
way to restrict a specific service to a subset of users without hand-
editing configuration.yml and users.yml directly.
_authelia_scope_access(SERVICE_ID, DOMAIN) is a new generic, reusable
helper: call it after any service finishes being protected by Authelia
(forward_auth gate or native OIDC alike — it only cares about the domain).
Offers universal vs. specific-users access; if scoped, creates a
"<service_id>-only" group, adds every listed username to it (creating
accounts on the fly for names that don't exist yet, via the new
_authelia_create_user_noninteractive — a non-interactive sibling to
add_authelia_user, same extraction pattern already used for
_authelia_provision_oidc_client), and inserts an allow+deny rule pair
above the general catch-all. Idempotent on rerun.
_authelia_report_access_scope() (new menu option 6) is the read side —
lists who has universal vs. service-scoped access, and offers to promote
a scoped user back to universal by removing their "-only" group
membership.
services/gitea.sh's _gitea_offer_authelia_sso() is the reference
integration, calling _authelia_scope_access after successfully wiring up
Gitea's OIDC login. The other services with a plain "Protect X with
Authelia?" prompt (magicmirror, wolf-pair, js99er, drum-rhythm-game,
iopaint, paintplus, stirling-pdf, wolf) are natural follow-ups once this
is confirmed working live — each just needs one added call.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Sipnetic clients on the user's home WiFi/VLAN lose SIP/TLS registration
every ~5s when Asterisk runs on IONOS, but never on DigitalOcean or over
mobile data — CrowdSec, OPNsense firewall/IDS, TURN-for-registration, and
raw packet loss have all been ruled out via live testing. Leading theory
is an idle-connection timeout inside IONOS's network virtualization layer.
keep_alive_interval sends a periodic double-CRLF over the TLS transport to
keep it from going idle, the standard mitigation for this failure class.
Follows the existing dual-patch pattern (vendor-template copy + live file)
since transport objects aren't picked up by `pjsip reload` and need a
container restart to apply, same as the live_dangerously fix.
Confirmed live: a repo's local bare mirror clone got stuck on a stale
commit indefinitely even though every sync run reported [ok] — fetch
never failed, it just wasn't updating refs/heads/* the way this script
assumes. GitHub had the real current commit; the bare clone (and
therefore what got pushed to Gitea) stayed frozen on an old one.
`fetch --all --prune` trusts remote.origin.fetch as stored in the bare
repo's own git config rather than asserting what that mapping actually
is — if it drifted from the +refs/heads/*:refs/heads/* convention a
fresh `git clone --bare` sets up (for whatever reason — this specific
repo's local clone directory's history is unclear), fetch would
"successfully" land new commits somewhere this script never reads
(refs/remotes/origin/*) while refs/heads/* — the ref that actually gets
mirrored — never moves. Deleting and re-cloning the affected repo's
local directory fixed it immediately, consistent with a refspec-drift
theory, though the exact original cause wasn't pinned down further.
Now pins `+refs/heads/*:refs/heads/*` explicitly on every fetch instead
of relying on `--all` plus whatever's configured, in both sync
directions. Also stopped redirecting stderr to /dev/null on every
clone/fetch/push call — a real auth or network failure now shows up in
the log instead of a bare "Failed to X" with no reason, which is what
made this bug take three rounds of manual ls-remote/rev-parse forensics
across two machines to actually pin down.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Gitea Actions is Gitea's own CI, largely GitHub-Actions-workflow-compatible
(.gitea/workflows/*.yml). Off by default; this Gitea install is otherwise
just a passive GitHub pull mirror, so the main value here is resilience —
.gitea/workflows/*.yml can still run something like a GitHub Actions build
if GitHub itself is ever unreachable.
_gitea_offer_actions_runner(), offered on fresh installs and Update reruns
(idempotent — no-ops if already set up):
- Enables GITEA__actions__ENABLED / DEFAULT_ACTIONS_URL in the compose
file's environment, restarts to apply
- Generates a runner registration token via `gitea actions
generate-runner-token`
- Appends an act_runner service to the same docker-compose.yml, using
the host's Docker socket to launch a fresh container per job — the
same pattern this repo already uses for portainer/watchtower/
uptimekuma/beszel/traccar's autoheal
- Falls back to printing manual setup instructions if token generation
fails, rather than losing the attempt silently
_gitea_fix_ownership()'s data/-exclusion (added when we fixed the earlier
SQLite readonly-database bug) now also skips runner-data/, so a future
reinstall/update doesn't clobber the runner's own state the same way.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Gitea has its own built-in login, so it was never wired into the
forward_auth/Caddy pattern the rest of this repo uses to gate apps with
no auth of their own — that's still correct and unchanged. But Gitea
also supports adding an OAuth2/OpenID Connect authentication source
natively, and Authelia can act as an OIDC provider — a genuinely
different, additive integration: an extra "Sign in with Authelia" button
on Gitea's own login page, alongside local login, not a Caddy-level gate.
Refactored services/authelia.sh's _authelia_add_oidc_client() to split
out its non-interactive core as _authelia_provision_oidc_client() — same
behavior for the existing ActualBudget/Vaultwarden/Immich/custom-app menu
flow, but now callable directly by other services with explicit args
instead of walking a human through the menu, returning the plaintext
secret and Authelia's domain via out-params.
services/gitea.sh's new _gitea_offer_authelia_sso() uses that to fully
automate both sides when accepted: registers Gitea as an OIDC client in
Authelia, then runs `gitea admin auth add-oauth` itself to add Authelia
as an authentication source — no manual web-UI copy-paste on either side,
matching how this installer already avoids manual wizards for the admin
account/token. Falls back to printing the values for a manual add if the
Gitea-side CLI call fails. Offered on fresh installs and on Update
reruns (default no, so a plain Update stays silent), so it can be added
later without a full reinstall.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Root cause of "attempt to write a readonly database (1544)" on repo
creation: Gitea's container always runs internally as UID 1000
(USER_UID/USER_GID are fixed in docker-compose.yml, independent of
whoever's running this installer) — the image chowns /data to that UID
itself at startup. install_gitea()'s three ensure_docker_dir_ownership
calls recursively chown the *entire* service directory, data/ included,
to $ACTUAL_USER. On a box where the installer runs as root directly
(ACTUAL_USER=root), that resets a live data/ back to UID 0. If the
container doesn't happen to restart right after — confirmed live: Update
mode against an already-running container just no-ops instead of
restarting — nothing ever re-fixes it, and every subsequent write to
Gitea's own SQLite DB fails.
Added _gitea_fix_ownership(), which chowns everything in the service
directory except data/, and swapped it in at all three call sites. The
container continues to own data/'s permissions exclusively, as it always
has on first boot; this installer no longer fights it on every rerun.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
((pull_count++)) evaluates to the PRE-increment value — 0 on the very
first successful pull/push — and under this script's `set -euo pipefail`,
an arithmetic command evaluating to 0 counts as a failing command and
kills the script immediately. Confirmed live: a real, fully successful
GitHub -> Gitea pull (visible in sync.log as "PULL ... OK") still made
the whole run exit non-zero and get reported as "Sync run failed", purely
because it was the first repo to sync (0 -> 1). Any subsequent repo in
the same run would have been fine, but most real installs only have a
handful of repos, so this could look like sync is just broken.
Switched all three counters (pull_count, push_count, fail_count) to
assignment form (`count=$((count + 1))`), which always exits 0 regardless
of the resulting value. page++ elsewhere in the file starts at 1, not 0,
so it isn't affected by this and was left as-is.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
"Failed to reach GitHub API. Check GITHUB_TOKEN." (and the equivalent
Gitea message) pointed at the token every time, even when the real cause
was something else entirely — confirmed live twice in one debugging
session: once a GitHub-side 503 outage, once a Gitea account locked
behind a must-change-password 403. Both times the fix was to run the
same curl by hand to see the actual status/response.
Fold that same probe into the script itself: on failure, re-request with
-i and print the HTTP status and response body directly, so the failure
mode (bad token vs. remote outage vs. account lock vs. network/DNS) is
visible immediately instead of requiring a manual curl round-trip.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Root cause of the "Failed to reach Gitea API" / 403 errors on every retry:
`gitea admin user change-password` (used in the already-exists branch to
sync the account's password to what the user just entered) defaults to
setting must_change_password=true, unlike `user create` which was already
pinned to --must-change-password=false. Once set, Gitea rejects every API
call — including the sync script's own token-authenticated calls — with
403 "You must change your password", even though the token itself and
GITEA_URL were both completely correct. Confirmed live via a direct curl
against /api/v1/user.
Pin the same flag on change-password that create already used.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Live symptom: pasting the GitHub PAT over SSH showed the token text
landing on the terminal *after* "No GitHub token entered" had already
printed — the prompt's read() returned empty a beat before the paste
actually arrived (a paste/Enter race that isn't specific to this box,
just common over higher-latency SSH sessions). A single empty answer
was treated as "user has no token" and the install moved on silently.
Both token prompts (GitHub token, and the Gitea-token manual fallback)
now retry up to 3 times interactively before giving up, and strip
whitespace from what was captured in case the paste carried a stray
leading/trailing newline. Unattended installs still take one shot, same
as before, since nobody's there to retry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
The sync direction step only ever set up the timer (or printed manual
instructions) — there was no way to actually confirm tokens/config work
without waiting for the first scheduled run or invoking the script by
hand afterward. Add a post-configure prompt: dry-run preview (--list),
run for real right now, or skip. Defaults to dry-run interactively;
defaults to skip under UNATTENDED so a headless install with no GitHub
token configured doesn't spam preflight errors.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
generate-access-token used a fixed --token-name "sync", which Gitea
rejects on a second call for the same user (e.g. a retry against an
already-existing admin account, now common after the readiness-wait
fix). The failure was silent: it fell through to a manually-labeled
"Paste the Gitea token here" prompt appearing immediately before the
real "GitHub token:" prompt, so a pasted GitHub PAT could land on the
wrong prompt and leave GITHUB_TOKEN empty with no clear reason why.
Token name now includes a timestamp so it's always unique, and the
fallback prompt is relabeled to make clear it wants a Gitea token, not
the GitHub one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
Previously the admin account was silently auto-generated (username =
ACTUAL_USER, random password) and only created after a fixed 60s
readiness probe — a slow first boot (SQLite init on a slower disk) timed
the whole install out with no account ever created, leaving the user to
create one by hand with a raw docker exec.
Now the install prompts for admin username/password up front, then folds
account creation into the same retry loop used to detect readiness (up to
2 minutes), so a slow-but-eventually-successful boot no longer dead-ends
the install. A retry against a partially-completed prior run (account
already exists) is treated as success and syncs the password instead of
failing outright.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YEQNc4NfBST1m9NtCZVYa8
CONTAINER_NAME was correctly re-detected from the current layout
(ASTERISK_KIND) at the top of install_pstn-trunk(), then immediately
clobbered by sourcing the saved .pstn-trunk.env settings file — and
_pstn_apply_settings writes CONTAINER_NAME straight back into that
same file every run, so a stale value self-perpetuates forever once
it's wrong.
Confirmed live: a box migrated from a DigitalOcean droplet (where the
settings file correctly saved CONTAINER_NAME=easy-asterisk-do) to a
plain/home install kept that droplet-era value indefinitely, since
only "update" runs happened after the migration. Every update
reloaded/restarted a container that no longer existed instead of the
one actually running Asterisk, so dialplan changes (including
pstn-trunk-inbound-dialplan.conf) never reached the live process even
though the config files themselves were written correctly.
Re-asserts the freshly-detected value after the source instead of
trusting whatever was saved, so it can't drift from the box's actual
current layout and self-heals the persisted file on the next update.
Gitea previously only existed bundled inside the full ai-stack service
(Ollama/ComfyUI/InvokeAI/etc. all together) — no way to get just a git
server without the rest of that heavy stack. This adds it as its own
lightweight service, reusing the vendored gitea-github-sync.sh but not
any of ai-stack's other components.
- Auto-creates a Gitea admin account and API token via the container's
own CLI (no manual web setup wizard).
- Asks GitHub token, sync direction (GitHub->Gitea / Gitea->GitHub /
both), and whether to install a systemd timer for automatic sync —
prints manual commands instead if declined.
- Own systemd unit for the timer rather than the vendor script's
built-in --install-timer, since that always runs both directions
with no way to pin a single direction.
pstn_personal_ring dialed the owner and hung up regardless of
DIALSTATUS, so no-answer/busy calls to a personal DID just dropped
silently instead of offering voicemail. Now falls to
VoiceMail(<owner>@default,u) on anything but ANSWER, gated by the
owner's existing voicemail=yes/no flag in pstn-permissions.conf.
The shared ring-group and group-owned personal DID inbound paths have
the same gap but no single owning extension to pick a mailbox for —
left as-is pending a decision on what that should do.
Adds live_dangerously = yes to asterisk.conf automatically on install/
update, fixing silent PSTN call denial (AST_CONFIG() returning empty
with no error when this option is off).
AST_CONFIG() silently returns an empty string instead of erroring when
asterisk.conf's [options] section lacks live_dangerously = yes, so a
correct tier_out=full in pstn-permissions.conf still evaluates as no
permission — every outbound/ring-group call gets denied with nothing
in the logs pointing at the real cause. Easy Asterisk's vendor default
ships without this set. Now applied automatically on fresh install and
on every "update" rebuild, restarting only when the file actually
changes.
Immich has full native OIDC support (its own docs list Authelia as a
supported provider), but needed more than the single-redirect-URI
model _authelia_add_oidc_client() previously supported: it requires
three redirect_uris at once (web login, account-linking page, and the
mobile app's app.immich:///oauth-callback custom-scheme redirect).
Generalized redirect-URI handling from a scalar REDIRECT_PATH to two
arrays (domain-relative REDIRECT_PATHS, plus already-complete
EXTRA_REDIRECT_URIS for non-domain-based ones like the mobile scheme)
and build the YAML redirect_uris list from however many are present.
ActualBudget/Vaultwarden/Other still resolve to a single-entry array,
so their generated config is unchanged. Verified the multi-entry YAML
generation against a python yaml parser before wiring it in, and the
case-statement/array logic in isolation against the real file's code.
Confirmed live (not from docs): Authelia's in-portal Settings -> Change
Password also emails a one-time code to confirm, same as Forgot
Password — it is not a no-SMTP path as earlier text here assumed.
Reworded all three spots in authelia.sh that claimed otherwise to
point at the admin-side "Edit an existing user" -> "Reset password"
action instead, which never touches email.
Also corrected CLAUDE.md's "No built-in auth — should be protected"
list per an actual grep of services/*.sh: it was missing
drum-rhythm-game, iopaint, paintplus, stirling-pdf, wolf, and the
unconditionally-protected security-dashboard/asterisk, and wrongly
included sky-cam (a non-Docker batch script with no web UI or Caddy
integration at all, nothing for Authelia to protect).
New menu option lists existing users by number; picking one opens a
submenu to edit email/display name, force a password reset, reset a
2FA device (authelia storage user totp delete), toggle a one_factor
exemption for that user via a subject-scoped access_control rule, and
promote/demote admin group membership.
All the YAML-editing helpers (line-range lookup, scoped field/group
edits, and the access_control rule insertion/removal used by the 2FA
exemption toggle) are line-range-scoped to the target user only, and
were verified against single- and multi-domain/multi-user fixtures
before wiring them into the interactive flow — a bad edit to
access_control here would break every protected domain, not just one
user's account.
30 characters, guaranteed at least 5 uppercase, 5 digits, and 5 special
characters, shuffled. Scoped to add_authelia_user() only via a small
local generator — deliberately not routed through lib/common.sh's
shared generate_password, since that one is alphanumeric-only by
design (its paired validate_password rejects special characters) and
plenty of other services embed its output unescaped into .env/YAML/URLs.
ActualBudget requires inviting additional OpenID users from its own
"Server Online" screen before their login is accepted, separate from
Authelia authenticating them successfully. Companion doc gets appended
to the generated README automatically (write_readme convention).
Adding a user previously required hand-editing users.yml and generating
the argon2 hash manually. New menu option (2) on an existing Authelia
install prompts for username/email/display name/admin group, generates
the hash and temp password, inserts the users.yml block, and restarts
Authelia — mirroring the existing add_authelia_domain/OIDC-client flows.
Asterisk DO->IONOS migration (resolved), web-based extension messaging
(not started), Pi-hole (done), VPS-as-VPN-endpoint with encrypted DNS
(idea stage). Temporary -- delete once these are finished or turned into
real issues/PRs.
Prompts on fresh/new installs and writes MM_SERVICESETTINGS_COLLAPSEDTHREADS
into .env; update reruns read the existing value back instead of
re-prompting, since it may have been changed later via System Console.
Documents the setting and how to flip it later in the generated README.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GKXTDG1ivYASN9fvDZ9Rfv
restore replaces every file under the Asterisk directory with a fresh
extraction from the archive (rm -rf + mv in the new tree) -- every file is
a new inode, so any POSIX ACL grants the Security Dashboard holds on them
(_secdash_grant_asterisk_access's setfacl access to .env/config/logs/spool)
are gone, and were never on the new files to begin with. Confirmed live:
surfaces as "No permission to read .../.env" in the Extensions tab right
after a restore, on a box where the dashboard was already working fine.
The restore script can't fix this itself (runs standalone, no access to
services/*.sh's functions or the dashboard's service-user name), so it now
prints a reminder to re-run `sudo ./setup.sh security-dashboard` at the end
of a restore when the dashboard's systemd unit is present, plus a matching
note in the generated README.
Filters by IP/range, scenario, ASN, carrier name, or country in a single
free-text field -- asked for so a phone's current IP or its network/carrier
name can be searched directly instead of scanning the full ban list by eye.
DEVICE_MARKER_RE expected "; === Device: NAME [AA:marker] (category) ===",
but device_config's own template (further down this file) generates
"; === Device: NAME (category) [AA:marker] ===" -- category parens before
the AA tag, not after. The regex never matched a real device comment, so
list_extensions() silently returned [] for every device on every install,
and /api/pstn-permissions served {"extensions": []} regardless of what was
actually in pstn-permissions.conf. That's why the Extensions tab's
Messaging/Voicemail checkboxes always rendered unchecked after a save +
reload even though the file itself had messaging=yes/voicemail=yes written
correctly -- the JS falls back to an all-default row when the endpoint
returns nothing. ea_list_devices() and the rename-device code parse the
same comment via string-splitting/a differently-shaped regex and were
already correct; this was the one broken parser.
The prior fix (930233c) wrote noload lines for app_voicemail_imap.so/
app_voicemail_odbc.so into config/asterisk/modules.conf at install time, but
vendor's docker/entrypoint.sh regenerates /etc/asterisk/modules.conf
unconditionally on every container start (same bind-mounted file) and
clobbers it within seconds — confirmed live on a fresh install with the
prior fix in place. Move the noload patch into
_asterisk_refresh_vendor_files()'s sed pass over the vendor template
instead, alongside the existing logger.conf/cert-regen patches, so it
survives the container's own regeneration. Drop the now-dead host-side
_asterisk_write_modules_conf() and its call sites.
Live-discovered bug, present on every install using this script, not
specific to any one box or extension: the easy-asterisk image ships
app_voicemail.so, app_voicemail_imap.so, and app_voicemail_odbc.so all
autoloading by default -- three alternative storage backends for the
SAME application (VoiceMail, VoiceMailMain, VMAuthenticate,
VoiceMailPlayMsg, VMSayName, the VM_INFO function, several AMI
actions), which collide registering those names against each other on
every single Asterisk start. This box's own container log showed the
exact signature on every restart: "Already have an application
'VoiceMail'" (and every sibling) followed by "app_voicemail.c:15897
load_module: Failure registering applications, functions or tests" --
app_voicemail never actually finished loading. Confirmed against
Asterisk's own documentation this session rather than assumed: this
is a known multi-backend conflict ("administrators should enable only
one module at a time"), not something specific to this repo's config.
Fix: new _asterisk_write_modules_conf, called from both the fresh-
install and update paths (matching voicemail-dialplan.conf's own
call-site pattern) alongside the other config/asterisk files, all
sharing the already-bind-mounted ./config/asterisk:/etc/asterisk
volume -- no new mount needed. noloads the two backends this repo
never configures (no IMAP/ODBC settings are ever written anywhere in
this script), leaving only the plain file-based app_voicemail.so
(the one voicemail.conf's [default] mailboxes actually target) to
load cleanly. Regenerated on every install/update, unlike
voicemail.conf, since modules.conf carries no per-install state of
its own -- consistent with how messaging-dialplan.conf/voicemail-
dialplan.conf are already handled, and added to their same chmod 644
line.
Requires a container restart to take effect on an existing install
(re-run `sudo ./setup.sh asterisk` -> Update, which regenerates this
file, then restart the container once).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Live blocker: user's actual migration is DigitalOcean droplet
(asterisk-digital-ocean/, container easy-asterisk-do) -> IONOS (plain
asterisk/, container easy-asterisk) -- exactly the case this script
didn't handle. The archive's own top-level directory name and
docker-compose.yml reflect whichever layout produced it
(_asterisk_resolve_layout's two known layouts). The previous restore
extracted straight into $PARENT_DIR, which recreates whatever name is
baked into the archive -- restoring a droplet archive onto a fresh
non-droplet install would land the data at a *second*,
wrongly-named directory (asterisk-digital-ocean) alongside the
freshly-installed one it was meant to replace, with docker-compose.yml
still naming the old project/container(s). Every service that resolves
Asterisk's layout by directory/container name (security-dashboard.sh,
pstn-trunk.sh, CrowdSec's Asterisk acquisition, Caddy) would get
confused by having two candidate layouts on disk, one of them stale
and half-wired.
Fix: extract into a scratch staging directory first. If the archived
docker-compose.yml's container_name differs from this run's own
$CONTAINER (baked in at generation time, so always correct for
whichever layout THIS box's install actually uses), rewrite the
project name, container name, and coturn container name in place
(coturn's is always "$CONTAINER-coturn" on both known layouts, so no
lookup table needed) before the data ever lands at $HERE -- never
lets the archive's own naming leak through. A same-layout restore
(most common case, or two droplet boxes, or two plain boxes) detects
no mismatch and skips the rewrite entirely, unchanged from before.
Verified against the real generated script (extracted from the
heredoc, not reimplemented): a droplet-flavored archive restored onto
a fresh plain-layout box lands at the correct single directory with
no stray second directory, and docker-compose.yml's name/container_name/
coturn container_name all correctly rewritten to the plain layout
(confirmed by diffing the actual restored file, not just checking for
absence of errors); a same-layout restore (droplet archive onto a
droplet box) confirmed to skip the rewrite entirely; the pre-existing
external-IP patch (previous commit) still fires correctly stacked on
top of the layout fix; and the extraction-failure rollback path (a
corrupt/unreadable archive) still restores the pre-restore install
untouched, verified via a marker file surviving the rollback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
User confirmed via a live pjsip.conf on their actual DO box: both
external_media_address and external_signaling_address are literal
IPs, written by easy-asterisk at first container start (not something
this repo's install script controls directly). A straight restore of
a backup archive from a DIFFERENT box onto the new IONOS box would
leave the OLD DigitalOcean IP baked into pjsip.conf — dialplan and
PJSIP device credentials would come back fine, but RTP media (and
likely SIP signaling/registration) would stay broken, silently, since
nothing in the restore path previously touched these values.
Fix: after extracting the archive, `restore` reads the archive's own
external_signaling_address as "old IP", detects this host's actual
current public IP (same DO-metadata -> ifconfig.me -> hostname -I
fallback chain services/asterisk.sh's own install already uses), and
if they differ, rewrites every occurrence across config/ and .env
(fixed-string match, not a regex, so the IP's dots can't be
misinterpreted). Deliberately does NOT touch spool/, logs/, or lib/ —
those hold voicemail messages and call recordings, and a blind text
substitution across binary audio would corrupt it. A restore onto the
same host (e.g. rolling back a bad config change, no IP change)
leaves every file untouched — the check only fires on an actual
mismatch.
Verified against the real generated script (extracted verbatim from
the heredoc, not a reimplementation) with a full mock backup/restore
cycle: built a fixture archive with pjsip.conf's three transport
blocks (udp/tcp/tls) and .env's TURN_SERVER all hardcoded to a fake
"old box" IP, plus a fake binary voicemail file; restored it onto a
mocked "new box" with a different detected IP via a stubbed curl.
Confirmed every occurrence in both pjsip.conf and .env was correctly
rewritten to the new IP, and confirmed via byte-for-byte comparison
that the binary voicemail file was completely untouched. Separately
verified the same-IP case (mocked curl returning the archive's own
IP) makes no changes at all, matching a same-host config rollback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
This is what's actually needed for a DigitalOcean -> IONOS Asterisk
migration (item #1 of the user's original 4-item list): move the
whole stack to different hardware entirely, not restore onto the
box the offsite mirror already targets.
dr_bringup_kopia.sh only ever scanned DEST_NAMES for restorable
snapshots — those are always local filesystem Kopia repos
(services/backup.sh creates them with `repository create filesystem
--path=...`), meaning they only exist on whichever box originally ran
the backup. On a genuinely new box, every one of them fails to
connect and there's nothing left to restore from — the script's own
header comment only covered the case where "the spare box IS the box
the primary's mirror targets" (i.e. already holds a copy of the repo
data), not a fresh, unrelated box.
Fix: also try REMOTE_TYPE/REMOTE_ARGS (the offsite Backblaze/S3
mirror, if configured) as a same-shaped destination named "offsite",
reusing DEST_default_PASSWORD since sync-to always mirrors that exact
same encrypted repo. Connects once into a fresh local config file
scoped to this DR run, then folds into the existing per-destination
scan/restore loop unchanged — "offsite" just becomes another entry in
_DEST_ARR. Documented the actual migration workflow in the header
comment, including the BACKUP_CONF override so copying the old box's
backup.conf over doesn't clobber the new box's own freshly-configured
one.
Verified against the real script (not a reimplementation) with mocked
kopia/docker binaries and a crafted backup.conf, covering: local dest
unreachable + offsite connects successfully (the actual migration
shape) with correct service/path discovery; a real (non---list)
restore run confirmed it selects the latest of multiple snapshots by
startTime and issues the correct `kopia restore <snapshot> <path>`
call; offsite connect failing (bad REMOTE_ARGS) warns and degrades to
"no restorable sources found" instead of crashing; REMOTE_TYPE=none
skips the new code path entirely with no behavior change (regression
check against the pre-existing local-only case).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
User's actual question: Backblaze B2's web console lets them browse a
bucket as folders/files; Garage has no equivalent by default, so after
switching an additional backup mirror from Backblaze-only to also
target local Garage, they had no way to visually confirm data landed
there the way they could on Backblaze. "S3 storage is opaque, you
can't browse it" was true of Garage's *own* CLI, but wrong as a
blanket statement — Backblaze's browsability comes from a client (its
web console) layered on top of the same kind of object storage, and
Garage has an actively-maintained equivalent (khairul169/garage-webui,
1.1k stars, "integrated objects/bucket browser") that gives the same
experience against Garage's S3 API.
services/garage-webui.sh (new): standard service-template Docker
service. Requires an existing services/garage.sh install (checks for
$DOCKER_DIR/garage/.env, errors with instructions if missing — this
is a browser for an existing instance, not a replacement). Reaches
Garage over host.docker.internal (both containers' ports are already
published to the host — simpler and more robust than trying to join
garage's own Compose-project-scoped default network by name). Has its
own login (AUTH_USER_PASS, bcrypt via a throwaway `docker run --rm
httpd:alpine htpasswd` — same $ -> $$ escaping services/wg-easy.sh
already uses for its own bcrypt PASSWORD_HASH, verified here against a
real docker compose config run: unescaped, Compose tries to interpolate
$2y$05... as variable references and silently corrupts the value with
a "not set" warning; escaped, it passes through intact with no
warning), so it doesn't need Authelia gating by default.
Prerequisite fix in services/garage.sh: its admin API (bucket/key
management, object listing — the thing garage-webui talks to) has
been running with zero authentication since this service was first
built, because admin_token was never set in garage.toml. Nothing in
this repo called that API before now, so it went unnoticed; adding a
real consumer is what surfaced it. Fixed: generate admin_token
(openssl rand -base64 32) alongside the existing rpc_secret, persist
GARAGE_ADMIN_TOKEN/GARAGE_ADMIN_PORT to .env for garage-webui to read
locally (never sent over SSH, unlike the S3 credentials backup.sh
reads remotely). Update mode backfills admin_token into an existing
garage.toml (+ restarts just the garage container to apply it) for
anyone who installed before this change, same backfill-not-break
approach as the GARAGE_S3_API_PORT fix from the previous commit.
Verified: bash -n on both files; docker compose config against real
Docker Compose for both the primary garage.toml/.env generation (with
the new admin_token/GARAGE_ADMIN_PORT fields) and the new
garage-webui docker-compose.yml; the bcrypt-escaping behavior
specifically (proved via a minimal repro that unescaped $ corrupts
the value with a warning, escaped does not); the admin_token/
GARAGE_ADMIN_PORT Update-mode backfill logic against old- and
new-style .env/garage.toml fixtures, including idempotency (running
it twice adds nothing a second time); and the credential-parsing
regexes in garage-webui.sh against both a complete .env fixture and
an old one missing the new fields (confirms the "run garage's Update
first" error path actually triggers rather than proceeding with
blanks).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Live failure: `sudo ./setup.sh garage` → Full reinstall on a box that
already had a working Garage install crashed with:
Error: ApplyClusterLayout returned InternalError (500): Internal
error: Invalid new layout version
Root cause: a "Full reinstall" deliberately never wipes ./data or
./meta (that's real backup-mirror data — Kopia's sync-to s3 target —
and losing it silently on reinstall would be far worse than the
alternative), but the cluster-init step unconditionally re-ran `garage
layout assign` + `layout apply --version 1` every time it was reached.
Garage requires each apply to be exactly previous_version + 1; a node
that already has a committed layout (from the earlier install, still
sitting in the preserved ./meta) rejects a second "1". Fix: check
`garage status` for "NO ROLE ASSIGNED" first and only run the
assign/apply once, matching what the surrounding comment already
claimed happened ("Only ever run once") but the code didn't enforce.
Second, related issue this would have hit immediately after: the same
reused-./meta state almost always means an existing bucket + key from
the earlier install are still sitting in Garage's storage. The fresh
flow was about to silently create a brand-new bucket/key and overwrite
.env to point at those instead — orphaning any real data already in
the old bucket (nothing left on disk pointing at it, even though it's
still physically stored). Now: when the layout is already applied,
list existing buckets and require an explicit y/n (default n) before
creating new ones, with recovery instructions for reconnecting to an
existing bucket by hand instead.
Third, the actual reason a full reinstall was reached at all: Update
mode never backfills .env fields added to this script after someone's
initial install (GARAGE_S3_API_PORT, needed by services/backup.sh to
read an instance remotely) since Update deliberately never touches
.env otherwise — the only other path was the now-unsafe fresh
reinstall. Update now backfills just that missing key by reading the
real port back out of the already-written docker-compose.yml, so a
future .env schema addition doesn't force this tradeoff again.
Verified with standalone harnesses (not the live install, mocked
`garage status`/bucket-list output and .env/docker-compose.yml
fixtures): all four layout-state branches (fresh node, existing
buckets + decline, existing buckets + confirm, existing role but no
buckets), and both backfill cases (missing key added, existing key
left alone). Caught and fixed a real bug in the first draft of the
port-extraction regex during this testing — grep -oE '^[0-9]+' never
matched because the captured group still had its surrounding quotes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Previously, if the additional-mirror S3/Garage check couldn't find
~/docker/garage/.env on the remote box, it just warned and silently
dropped the mirror — forcing a full re-run (and re-entering every
already-answered prompt: destinations, passwords, schedule, B2,
DR-spare, etc.) once Garage was actually installed.
Wrap the S3/SFTP branch in a loop so the "Garage isn't installed yet"
case now offers a real 3-way choice:
1) install Garage in another session, then retry the same .env check
without leaving this script
2) fall back to SFTP for this one mirror, reusing the already-resolved
destination host/port/user/mirror-name with no re-prompting
3) skip just this mirror (default — safe for UNATTENDED, which
resolves to this automatically since prompt_text returns its
default without blocking)
Everything else install_backup() has already collected lives outside
this loop, so none of it is at risk regardless of which of the three
exits it via.
Verified against a standalone harness reproducing the state machine
with a mocked ssh (empty .env vs. populated .env after a simulated
install) and prompt_text, covering all three interactive choices, the
blank/Enter default, and UNATTENDED mode (confirms the blocking
"press Enter to retry" read is unreachable there since prompt_text
resolves choice 1's prompt to default "3" without waiting on stdin).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Extends the "ADDITIONAL MIRROR" section (previously SFTP-only) with a
type choice: SFTP, or S3 against a Garage instance already running on
that box. For the S3 path, this script never asks the operator to retype
a bucket name or key — it SSHes to the destination, reads
~/docker/garage/.env directly (the real, currently-configured values,
generated once by services/garage.sh and never touched again on its own
Update runs), and uses those for the dry-run verification and the
persisted mirror args. If Garage isn't installed there yet, it says so
plainly with the exact install command instead of failing cryptically or
silently skipping.
Also removed the last hardcoded suggestions from services/garage.sh
itself ("kopia-backup" / "kopia" as fixed prompt defaults) — replaced
with a freshly-generated suggestion each run (timestamp-suffixed), so
nothing about the bucket/key name is a fixed string baked into this repo
at any point in the chain; it's always the operator's actual choice, read
back live wherever it's needed.
Verified end-to-end against a mocked ssh (returning realistic
~/docker/garage/.env content) covering both outcomes: Garage installed
with a real bucket/key correctly parsed, dry-run run, and persisted; and
Garage missing, correctly warning with the install command and leaving
backup.conf untouched either way.
Garage's real CLI output pads labels with extra spaces for column
alignment ("Key ID: GKxxxx"), not a single space like the
mocked test used ("Key ID: GKxxxx") — the fixed ": " field separator left
that padding stuck to the parsed value, so .env ended up with access
key/secret strings carrying leading whitespace inside the quotes.
Confirmed live by the user right after install. This would have broken S3
auth outright once actually used, since access keys have to match exactly.
Switched to ':[[:space:]]+' as a regex field separator, which consumes
however many spaces are actually there instead of assuming exactly one.
Verified against both the single-space and padded/aligned formats — both
now produce the identical clean value with no leading whitespace.
MinIO's open-source community edition is dead: console GUI stripped May
2025, Docker images stopped publishing October 2025, repo formally
archived April 2026, with MinIO redirecting everyone to their paid AIStor
product. Verified this directly before building anything, since recommending
a since-abandoned image would have been worse than the SFTP problem this
was meant to solve.
Garage (Deuxfleurs) is the actively-maintained small-scale self-hosted
replacement — single Rust binary, purpose-built for exactly this "one
lightweight node" use case (as opposed to SeaweedFS, which targets large
object counts / large-scale deployments, more machinery than a single
backup-mirror target needs).
services/garage.sh follows this repo's standard service template: port
scanning for the S3 API/RPC/admin ports, an RPC secret generated once and
never touched again on Update, and a one-time cluster init sequence
(layout assign/apply, bucket create, key create, bucket allow) gated on
whether .env already has a saved access key — Update reruns skip all of it
and just refresh the image.
Primary intended use: a local S3-compatible target for services/backup.sh's
additional-mirror Kopia sync, so a local mirror can reuse the exact same
sync-to s3 code path already proven reliable for the Backblaze B2 mirror,
instead of Kopia's separate, less-exercised SFTP backend that's been the
source of today's connection troubleshooting.
Verified end-to-end against a mocked environment (fake docker exec
returning realistic `garage status`/`garage key create` output) — caught
and fixed a real off-by-one in the status-output parsing this way (grabbed
the column-header row's literal "ID" instead of the actual node ID; output
has a title line, then a header line, then the data row). Also validated
the generated docker-compose.yml with real `docker compose config` in both
the no-network and network-created cases.
The additional-mirror "Remote path for the repo" prompt always suggested
a generic ~/backups/kopia-mirror default, unrelated to wherever the
operator already pointed the DR-spare sync. Requested directly: default
to that same location instead, in its own /kopia-data subdirectory so
Kopia's repository files don't end up visually mixed in with the two
plain config files (backup.conf, README.md) the DR-spare sync writes
straight into DR_SYNC_PATH itself.
Falls back to the original generic default when DR_SYNC_PATH isn't set
(no DR-spare configured yet). Verified the path computation handles a
DR_SYNC_PATH with or without a trailing slash correctly (no double slash),
and the unset case still falls back as before.
The additional-mirror setup already resolves user/hostname through ssh -G
so a ~/.ssh/config alias works, but never extracted port — Kopia's sftp
storage backend doesn't read ~/.ssh/config at all and defaults to 22
regardless of what the alias actually configures. Confirmed live: this
produced "server unexpectedly closed connection: unexpected EOF" on the
dry-run verification — Kopia connecting to the right host on the wrong
port, not a credentials or host-key issue, which is exactly why plain
`ssh main` kept working the entire time this was being debugged (it reads
the alias's Port line correctly).
Now parses `port` out of the same ssh -G output, defaults to 22 if absent
(matching ssh's own default), and passes --port= through to both the
dry-run check and the persisted EXTRA_MIRROR_ARGS string — the latter
matters as much as the former, since that's what every actual scheduled
sync reuses afterward, not just the one-time verification.
Verified the parsing against three cases: a custom-port alias, a
default-port alias, and an unresolvable alias — all three resolve to the
correct port with no manual intervention needed.
extras/fix_pikapods_dump.py patches two confirmed Adminer PostgreSQL-export
bugs that otherwise make a PikaPods Mattermost migration fail outright:
unquoted enum-label DEFAULT values (Postgres reads the bare label as a
column reference and rejects the CREATE TABLE) and boolean columns
serialized as bare 0/1 instead of true/false (Postgres doesn't implicitly
cast integers to boolean). Boolean columns are discovered by actually
parsing each CREATE TABLE in the dump rather than working from a
hand-curated list — Postgres only reports the first bad column per failed
row, so a list built from error output alone would likely be incomplete.
Already verified earlier this session against a real local Postgres 16
instance; reviewed now for anything needing redaction before committing —
it's a generic text-processing tool with no hostnames, credentials, file
paths, or personal data in it, so nothing needed changing.
Cross-referenced from the generated migrate-from-pikapods.sh's header
comment (services/mattermost.sh) so anyone hitting a CREATE TYPE/CREATE
TABLE or boolean-column import error is pointed at the fix instead of
having to rediscover it.