The user found ToS language stating the account "may run on a negative
balance" and describing a 30-day-notice-then-suspend / 30-consecutive-
days-then-close process - worth reading against the wiki's "balance must
be over $0 to call" claim the whole toll-fraud design leans on.
Reconciled: these describe two different things, not a contradiction.
New call attempts should still be blocked in real time at $0 (the wiki's
claim, and the core assumption this design needs). The ToS's negative-
balance language most plausibly covers recurring fees (DID/E911) landing
when balance is already near zero, and in-progress-call settlement edge
cases - not a window where fraud keeps dialing while negative. Still not
verified against a live account either way.
Adds an inbound concurrent-call cap mirroring the existing outbound one -
outbound alone didn't protect against an inbound call-flood, which also
costs money per-minute on VoIP.ms. Both defaults bumped from 3 to 10.
Moves the cap numbers themselves out of static dialplan text and into a new
pstn-limits.conf, read live via AST_CONFIG() the same way permission tiers
already are - changing either cap takes effect on the next call, no
Asterisk restart, no reinstall. "update in place" never touches this file,
matching the existing protection for pstn-permissions.conf/.env/firewall/
Caddy config.
Adds a concurrency-caps card to security-dashboard.sh's "PSTN Trunk" tab,
above the existing permissions table, so both caps are visible and editable
from the same web page. Tested against a real running instance of the
Python app: default fallback when the file doesn't exist yet, save/persist,
invalid-input rejection, and a bash-to-Python round trip on the generated
file format.
Inbound dialplan ordering mirrors outbound's existing pattern: permission
check (is any ring-group member authorized for this caller) before the
concurrency check, consistent with outbound's tier-check-then-busy-check
order.
Reworks the outbound permission model from a flat allow-list into three
per-extension tiers (internal / restricted / full), addressing the ask for
extensions that can only reach pre-approved numbers plus extensions with
full US calling, while internal extension-to-extension dialing and ring
groups stay ungated for everyone regardless of tier.
Permissions now live in pstn-permissions.conf, read by the dialplan via
Asterisk's AST_CONFIG() on every call instead of being baked into static
dialplan text - editing that file takes effect on the next call, no
Asterisk restart and no re-running the installer. "update in place" mode
never touches this file (same protection this repo's update-mode
convention already gives .env/firewall/Caddy config); only a "fresh"
reinstall (with confirmation) or the web UI change it.
Adds a "PSTN Trunk" tab to services/security-dashboard.sh: lists every
extension (parsed from pjsip.conf) with its live tier and approved numbers,
editable with no restart - this is what makes the tier model actually
manageable day to day. Extracted the dashboard's systemd-unit writing into
its own function so "update" mode refreshes it too (previously only fresh
installs did), and generalized both the dashboard and the trunk service to
detect either asterisk-digital-ocean or the home/LAN asterisk install.
Inbound ring-group membership now checks each member's tier live per call
via an unrolled per-member dialplan block (full always rings, restricted
only if the caller's number is approved, internal never rings) rather than
a single static Dial() string.
Caught and fixed two real bugs during testing against a sandboxed vendor
copy and a live instance of the (stdlib-only) Python dashboard app:
- Asterisk Goto/GotoIf argument parsing: ring<ext>/skip<ext> are named
priorities within the same extension (declared via "same => n(label),..."),
not separate exten => entries, so jumping to them needs the single-argument
Goto(label) form - the two-argument Goto(label,1) form used initially
addresses a different, nonexistent extension named "label" instead.
- A security-relevant REGEX() direction issue: the inbound Caller-ID check
initially interpolated attacker-influenced call data into the PATTERN side
of a REGEX() match rather than the tested-string side, which would let a
crafted Caller-ID forge a match against an unrelated approved-numbers
entry. Fixed by keeping the admin-controlled approved-list as the pattern
and the live call data as the string being tested, consistently on both
the outbound and inbound checks.
Verified end-to-end: dialplan/pjsip generation and vendor-file patching
(idempotent, syntax-checked) as before, plus the new permission-file
round-trip between bash and Python, and the dashboard's new API endpoints
exercised against a real running Python server (extension parsing, tier
changes, number normalization, invalid-input rejection, atomic file writes).
Generalizes services/pstn-trunk.sh (renamed from voipms-trunk.sh in the
prior commit) away from VoIP.ms specifics - any IP-authenticated SIP
provider works, VoIP.ms is just the suggested default. Adds:
- Role-based outbound permission: a configurable allow-list of extensions
that may dial PSTN numbers (regex-gated on CHANNEL(peername)), separate
from internal extension-to-extension dialing which stays open to everyone
regardless. Blank list preserves the original "everyone can dial out"
behavior.
- Inbound ring-group: rings a configurable list of extensions instead of a
single hardcoded one.
- ntfy alerts: immediate on denied (unauthorized extension) or rejected
(concurrency cap hit) calls, plus an hourly cron-driven check that alerts
once per month when estimated spend crosses a threshold and every hour
call volume looks like a burst. Uses a self-contained pipe-delimited call
log rather than Asterisk's CDR, to avoid depending on CDR module
availability and CSV comma-quoting.
- Settings persisted to .pstn-trunk.env so "update in place" reapplies
everything from that file instead of fragile re-parsing out of generated
Asterisk config (which had a real bug: update mode was extracting the
wrong Dial(PJSIP/...) line).
Tested end-to-end against a sandboxed copy of the real vendor files:
permission-gate regex, ring-group dial-string construction, ntfy line
injection/removal, and the usage-alert script's threshold/burst/monthly-
dedup logic all verified with synthetic data. Caught and fixed a sed `&`
escaping bug in the ring-group substitution before it shipped (RING_DIAL
contains literal `&` join characters, which sed's replacement syntax
otherwise treats as "insert the match").
Renames services/voipms-trunk.sh to services/pstn-trunk.sh and generalizes
it away from VoIP.ms specifics - any IP-authenticated SIP provider works,
VoIP.ms is just the suggested default. Adds:
- Role-based outbound permission: a configurable allow-list of extensions
that may dial PSTN numbers (regex-gated on CHANNEL(peername)), separate
from internal extension-to-extension dialing which stays open to everyone
regardless. Blank list preserves the original "everyone can dial out"
behavior.
- Inbound ring-group: rings a configurable list of extensions instead of a
single hardcoded one.
- ntfy alerts: immediate on denied (unauthorized extension) or rejected
(concurrency cap hit) calls, plus an hourly cron-driven check that alerts
once per month when estimated spend crosses a threshold and every hour
call volume looks like a burst. Uses a self-contained pipe-delimited call
log rather than Asterisk's CDR, to avoid depending on CDR module
availability and CSV comma-quoting.
- Settings persisted to .pstn-trunk.env so "update in place" reapplies
everything from that file instead of fragile re-parsing out of generated
Asterisk config (which had a real bug: update mode was extracting the
wrong Dial(PJSIP/...) line).
Tested end-to-end against a sandboxed copy of the real vendor files:
permission-gate regex, ring-group dial-string construction, ntfy line
injection/removal, and the usage-alert script's threshold/burst/monthly-
dedup logic all verified with synthetic data. Caught and fixed a sed `&`
escaping bug in the ring-group substitution before it shipped (RING_DIAL
contains literal `&` join characters, which sed's replacement syntax
otherwise treats as "insert the match").
Adds a VoIP.ms SIP trunk on top of asterisk-digital-ocean: IP-authenticated
trunk (no password stored), NANP-only outbound dialplan, a global 3-call
concurrent cap via GROUP()/GROUP_COUNT(), and inbound routing to one
extension. Config lives in its own include files rather than being
appended directly to pjsip.conf/extensions.conf, since Easy Asterisk fully
regenerates both from its own internal state — the includes are patched
into the vendor's generator functions so they survive that regeneration.
Wires the new service into setup.sh's is_installed() and README's services
table, and updates docs/pstn-calling-voipms-plan.md to reflect what's now
implemented vs. still open (spend/volume alerting, live-account
verification).
Inbound (DID) is now decided as wanted, not outbound-only. Adds a cost
estimate table for 100 min/month each direction, and clarifies that
NANP-only restriction bounds cost-per-minute but not burn speed, so the
concurrent-call cap and spend alert are required before funding a live
trunk, not optional hardening.
Captures the toll-fraud / prepaid-cap research and decisions from a design
discussion (provider: VoIP.ms, US-only calling for now) so implementation
can be picked up in a future session without re-deriving the background.
Nothing implemented yet.
- CrowdSec tab: per-ASN "Unwhitelist" (drop from the exempt filter, future
traffic evaluated normally) and "Unwhitelist + Ban" (also immediately
bans every IP CrowdSec has on record for that ASN, for accidental-
whitelist cases) buttons. set_asn_exempt now allows clearing the list
down to zero ASNs, needed to unwhitelist the last remaining entry.
- New sudoers permission (cscli decisions add --ip * --duration * --type
ban --reason *) scoped narrowly, list-form subprocess args only.
- Caddy/Authelia config factored into _secdash_configure_caddy() and a new
_secdash_remove_caddy_block() (whole-block delete-and-regenerate, not
in-place patching) so "update" mode can now offer to reconfigure it.
- Installer offers an independent HTTP Basic Auth layer in front of
Authelia (Caddy basicauth, generated via `caddy hash-password`) so a
future Authelia bug/misconfig alone isn't enough to expose a page that
can delete active security bans.
Currently-exempt ASNs with no active ban (e.g. T-Mobile once its bans
stop firing) had no carrier name to show, since the name lookup only
looked at cscli decisions list (active bans only). Add a
cscli alerts list-based fallback (includes expired/resolved alerts)
and merge it into the name lookup used by /api/asn-exempt.
Confirmed the real cscli decisions list -o json structure live rather
than guessing again: AS number/name and country live on each alert's
"source" object (source.as_number, source.as_name, source.cn), not on
the individual decision. Surfaces this as a "Network / Carrier" and
"Country" column in the bans table, adds a per-row "Exempt ASN" button
that appends straight to the Asterisk brute-force ASN exemption list,
and labels the exempt list's own entries with carrier names (pulled
from current ban data where available) instead of showing bare numbers.
Confirmed live: Caddy (in a container) reaches this via
host.docker.internal, a Docker bridge gateway IP, not localhost — a
loopback-only bind refuses that connection outright ("dial tcp
172.17.0.1:8092: connect: connection refused"), even though curl from
the host itself worked fine on 127.0.0.1. Bind to 0.0.0.0 and rely on
UFW for the actual access scoping instead, matching every other
host-network service in this repo (e.g. the Asterisk web admin, which
already binds this way successfully with the same
ufw_allow_from_caddy_net pattern).
Confirmed live: setup.sh's run_service() calls install_${name} with
no hyphen-to-underscore conversion, so a hyphenated service name needs
a literally-hyphenated function name (install_security-dashboard, not
install_security_dashboard) to be found at all — got this wrong on
first pass by following CLAUDE.md's own (incorrect) guidance, which
said to convert hyphens to underscores. Every other hyphenated service
in the repo (asterisk-digital-ocean, wolf-pair, mail-archiver,
drum-rhythm-game) already keeps hyphens literal; corrected CLAUDE.md
to match actual practice instead of the other way around.
New service, native on the host (not Docker) so it can call cscli and
read Asterisk's security log directly without bridging the
container/host boundary or exposing CrowdSec LAPI credentials to a
containerized frontend.
- Security Log tab: parses ~/docker/asterisk-digital-ocean/logs/full
for SIP auth failures (wrong password, unknown extension, etc.) with
timestamp/account/remote IP, classified by severity.
- CrowdSec tab: current bans via cscli, a delete/unban button per
entry, and ASN-exempt management for the Asterisk brute-force
scenarios (services/crowdsec.sh) without SSHing in.
- Link out to the existing Asterisk web admin (reads its domain from
asterisk-digital-ocean's own .env, doesn't hardcode or embed it).
Runs as a dedicated unprivileged system user (secdash), with sudo
scoped to exactly three commands via /etc/sudoers.d/security-dashboard
(cscli decisions delete --id <digits>, cscli decisions list -o json,
systemctl restart crowdsec) — validated with visudo -c. Listens on
127.0.0.1 only, reachable through Caddy, and refuses to proceed without
explicit confirmation if no Authelia (local or remote) is configured,
since this page can delete active security bans.
Stdlib-only Python (no framework), matching the RAM-conscious pattern
already used for Easy Asterisk's own web admin. All embedded code
(bash, Python, JS) syntax-checked; the generated sudoers rule
validated with visudo -c -f.
Confirmed live: unquoted integer literals in the ASN exclusion filter
(evt.Enriched.ASNNumber in [21928, 14593]) made CrowdSec fatal-crash-loop
at startup with "cannot use string as type int in array" — ASNNumber is
a string field internally despite printing as a bare number in cscli
output, same as IsoCode in the geo-allowlist scenario. Quote each ASN
as a string to match, exactly like the working geo-allowlist pattern.
Confirmed live: a phone roaming WiFi<->mobile on a CGNAT carrier
(Starlink, T-Mobile home internet) got banned by crowdsecurity/asterisk_bf,
either from its own re-registration burst or collaterally from another
customer sharing the same rotating public IP. Forks asterisk_bf and
asterisk_user_enum locally with an ASN exclusion added to their filter,
disabling the hub originals so events aren't double-processed. Scoped
narrowly to Asterisk auth-failure detection only — SSH, web scanning,
and the geo-allowlist scenario are all unaffected, so this doesn't
broadly exempt the carrier from every protection on the box.
Confirmed live: a legitimate SIP device on a CGNAT ISP (Starlink,
T-Mobile home internet) got collaterally banned by
crowdsecurity/asterisk_bf alongside actual bad actors sharing the same
carrier IP. The alert now includes the exact commands (with the banned
IP substituted in) instead of just naming the ban, so recovering from
this doesn't require remembering or looking up cscli syntax.
New "Add another protected domain to this instance" option on re-run,
via add_authelia_domain(): appends a session.cookies entry and an
access_control.rules entry (both YAML lists Authelia natively supports)
plus a Caddy auth.<domain> portal block for the new domain, all on the
same Authelia + Redis container instead of standing up a second full
stack. Each domain gets its own login/session, sharing one user
database — the right fit when a single (possibly upsized) droplet ends
up fronting more than one domain, without doubling the RAM cost of a
second Authelia+Redis instance. Documents both this and the
already-working separate-instance path in CLAUDE.md, with the
per-approach tradeoffs.
Confirms services/authelia.sh's standalone pattern and
asterisk-digital-ocean.sh's local-vs-remote auto-detection already
support a second, fully independent instance on another machine with
no code changes needed. Documents the one real constraint: two
instances must not share the same AUTHELIA_DOMAIN, since the session
cookie scope and the auth.<domain> portal hostname would collide.
Misread the previous request as "add Russia to the allowed list" —
it meant the opposite: Russia should stay excluded, same as the other
high-risk/Eastern Europe entries already left out.
Excludes Bulgaria, Czechia, Hungary, Moldova, Poland, Romania, Slovakia,
and Ukraine per user request, while keeping the Balkans and Baltics
(several of which, e.g. Estonia, don't fit the same risk profile despite
the old Cold-War grouping). Russia added back to the allowed list per
explicit user request.
Opt-in prompt that bans any Caddy-fronted web request from outside an
editable North America + Europe country list, via a local CrowdSec
trigger scenario scoped to the existing "type: caddy" acquisition label.
Uses CrowdSec's bundled GeoLite2 enrichment data (already active with no
extra setup) rather than the firewall bouncer's separate MaxMind-key
country-CIDR feature, so no account signup is needed. SSH is untouched
so a bad edit can't lock out the session running the installer.
The generated auth.<domain> block's bare "reverse_proxy authelia:9091"
let Caddy recompute X-Forwarded-Host from its own incoming request
(always auth.<domain> itself) on every hop through it, overwriting
whatever a forward_auth caller elsewhere had already set for its own
domain. Confirmed live: a remote site's forward_auth check always
evaluated as if it were for the Authelia portal itself (bypass policy),
so 2FA silently never triggered for any domain going through it.
The previous fix used the {host} Caddy placeholder for X-Forwarded-Host,
but confirmed live it still evaluated to the upstream Authelia's own
hostname rather than the original site's — Caddy appears to rewrite the
outgoing request's Host to the upstream target before header_up
placeholders resolve for a scheme-qualified remote upstream, so {host}
echoed back the already-rewritten value. Since this site block only ever
serves one domain, hardcode it instead of depending on placeholder timing.
The remote-Authelia forward_auth block dialed a scheme-qualified upstream
(https://auth.example.com), which is a second Caddy hop. Caddy rewrites the
outgoing Host header to the upstream host for routing, and without an
explicit override X-Forwarded-Host picked up that rewritten value instead
of the original site's host. Authelia was evaluating every protected
domain as auth.example.com itself (bypass policy), so 2FA never triggered
for any domain behind the remote instance. Pin the forwarded headers to
the original request explicitly to fix it.
Both Caddy site block generators (the shared configure_caddy_for_service
helper, and asterisk-digital-ocean.sh's own inline template) wrote
reverse_proxy before the forward_auth/import authelia block. Caddy
doesn't reorder repeats of the same directive within a block — forward_auth
and reverse_proxy are the same directive family internally, so they run in
the order written. With reverse_proxy first, it handled and terminated
every request immediately; the auth check written after it never ran at
all. Full bypass on every domain using either generator with Authelia
protection, regardless of how correct the Authelia access_control rules
themselves were — confirmed live against a config that was otherwise
completely correct (default_policy: deny, explicit wildcard rule covering
the affected domain).
Affects every service that's ever passed `import authelia` or a
forward_auth block through configure_caddy_for_service (asterisk.sh,
wolf-pair.sh, and any future caller), plus asterisk-digital-ocean.sh's
own site block.
Moved the auth block before reverse_proxy in both generators.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Live-confirmed the retry fix from the previous commit wasn't enough:
still "Address already in use" after 20s of retries (10 attempts,
2s backoff), only to succeed on its own sometime after that. That
delay pattern is TIME_WAIT, not a process-death race — and this code
was never going to avoid it, because socketserver.TCPServer defaults
allow_reuse_address to False. (http.server.HTTPServer sets this for
you; the plain base class used here does not.) Without SO_REUSEADDR,
the kernel can refuse to rebind a port with a lingering TIME_WAIT
socket from the previous instance for up to 60s, regardless of
whether that old process is even still alive — which is also why the
entrypoint.sh fix waiting for the process to exit didn't help either.
Set socketserver.TCPServer.allow_reuse_address = True before binding.
This is the standard fix for exactly this symptom. Keeping the retry
loop from the previous commit too, for the (now much smaller) window
where network_mode: host still has no Docker-managed port mapping to
instantly free.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Complements the retry fix on the bind side: pkill only sends SIGTERM
and returns immediately, it doesn't wait for the process to exit and
release its socket. Under network_mode: host there's no Docker-
managed port mapping to tear down, so the next container's bind
attempt was racing however long this process actually took to die —
sometimes still holding the port when the next container started.
Poll for it to actually exit (up to 2s), falling back to SIGKILL if
it's still lingering, before proceeding with the rest of shutdown.
With a clean handoff here, the web admin's own bind-retry (previous
commit) should rarely even need to kick in.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Live-confirmed: OSError: [Errno 98] Address already in use on the
web admin's TCPServer bind, right after a container recreate under
network_mode: host. Unlike bridge-mode port publishing, there's no
Docker-managed mapping to instantly free on teardown — the previous
container's own web admin process has to actually die first, and a
fast recreate-right-after-recreate can race that. The process crashed
immediately instead of retrying, so the web admin silently never came
up despite entrypoint.sh correctly launching it.
Retry the bind up to 10 times with a 2s backoff before giving up.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
"pjsip reload" was never a valid Asterisk CLI command — confirmed
live: `asterisk -rx "pjsip reload"` returns "No such command 'pjsip
reload'". The real command is "module reload res_pjsip.so".
Every reload-after-change call in the script used the invalid form,
both in the interactive CLI (add/edit/delete device, transport setup,
TLS cert sync) and in every web admin mutation (add_device,
delete_device, rename_device, change_device_category) — all silently
no-op'd, since `asterisk -rx` just prints its own "no such command"
error to a discarded/redirected output and returns normally either
way. Endpoints only ever picked up new pjsip.conf entries after a
full container restart (which re-reads config from scratch at
startup) or a manual `module reload res_pjsip.so` — never from the
web admin's own reload call, live-confirmed: a device stayed
Unavailable with zero registration attempts logged until a manual
reload picked it up immediately.
Global replace across all 11 occurrences.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
wolf-pair has no login of its own — Authelia via Caddy is the only
protection option offered for it — but UFW opened its port to the
whole internet unconditionally, before the Caddy/Authelia prompt even
ran. Same gap just fixed for the Asterisk web admin: reachable
straight over the bare port regardless of Authelia.
Reordered so the Caddy decision happens first, and scope the port to
caddy_net's subnet via ufw_allow_from_caddy_net() instead of leaving
it open to 0.0.0.0/0 when Caddy fronts it locally. Also enables UFW
via ensure_ufw_enabled() like the Asterisk services.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Confirmed live: a bare `ufw delete allow <port>` closes it on every
interface, including the caddy_net bridge — Caddy's own request to
host.docker.internal:PORT is ordinary INPUT-chain traffic as far as
UFW is concerned, not something that bypasses it just because the
source is a local container. Closing the port outright silently took
Caddy's reverse-proxy path down with it.
Added ufw_allow_from_caddy_net() to scope the port to caddy_net's own
subnet instead of leaving it fully closed — reachable from Caddy,
still closed to the public internet. Wired into both
asterisk-digital-ocean.sh and asterisk.sh in place of the plain
delete.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Both defaulted to whatever sorted/listed first (kiosks category,
LAN/VPN UDP transport) — reasonable for a fixed intercom install, but
the common case here is adding a phone over the internet. Default the
category select to "mobile" specifically (not just first-in-list, so
it survives category reordering) and make FQDN/Internet (TLS) the
default transport option instead of LAN/VPN (UDP).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Mirrors the fixes just made in asterisk-digital-ocean.sh:
- Reordered so the Caddy reverse-proxy decision happens before the
UFW rules are built, using the new CADDY_SERVICE_CONFIGURED/
CADDY_SERVICE_MODE signal from configure_caddy_for_service() to
skip opening the web admin port on the LAN when a local Caddy is
already fronting it (still opens it for a remote Caddy machine,
which needs LAN access to reach this host directly).
- Calls the new ensure_ufw_enabled() so UFW actually enforces the
rules this script adds, instead of leaving them queued but inert.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
configure_caddy_for_service() previously gave callers no way to know
whether Caddy actually ended up fronting the service, or whether that
was local (reachable only over host.docker.internal) vs remote
(needs network access to this host). Services that also open a host
firewall port for the same thing had no way to correctly skip that
when Caddy is the only intended way in. Now sets
CADDY_SERVICE_CONFIGURED/CADDY_SERVICE_MODE out-params after each
exit point.
Added ensure_ufw_enabled(): flips UFW from inactive to active (no
service in this repo has ever done this — ufw allow rules just sat
unenforced). Always allows SSH first, reading the real port from
sshd_config in case it's non-default, so this can't lock out the
session running the installer. No-ops if UFW is already active.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
UFW and the DO Cloud Firewall both opened the web admin port to
0.0.0.0/0 unconditionally, even when Caddy+Authelia was configured to
protect it on the actual domain. Caddy reaches the container over the
host's internal network (host.docker.internal), not the public
internet, so that direct port was pure attack surface: anyone could
hit http://<droplet-ip>:<port>/clients directly, fully bypassing
Authelia and the built-in web admin auth (which gets disabled
whenever Authelia is handling it instead).
Reordered the install flow so the Caddy reverse-proxy decision is
made before the firewall rules are built, and only open the web
admin port publicly when there's no local Caddy actually fronting
it (no domain, Caddy not installed, proxy declined, or a remote
Caddy machine that needs to reach it over the public IP instead).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
The web admin script (/usr/local/bin/easy-asterisk-webadmin) is
generated on demand by the interactive CLI, but only ever lived in
the container's writable layer — not baked into the image, not
bind-mounted. Every docker compose down/up wiped it, and the
entrypoint's start logic only ran "if the file already exists", so
it silently never started again until someone manually ran the CLI's
Web Admin menu once per recreate.
Added a --write-web-admin-script non-interactive entry point
(same pattern as --rebuild-dialplan) and call it unconditionally
before the existence check, so the web admin comes back on its own
every time the container starts.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
The "device created" modal only ever showed extension/password/name,
so setting up a client (e.g. Sipnetic) meant hunting down the server
domain, port, and transport separately — and the password is only
ever shown this once, so re-checking it later isn't an option.
Now shows everything a SIP client needs in one place: display name,
server, port, transport, username, password, plus TURN/STUN details
when enabled. The backend reports the actual transport/port used
(the container always forces TLS/FQDN mode regardless of what's
selected in the form, so the frontend no longer has to guess).
Added "Copy All" and "Copy Password" buttons, with a document.
execCommand fallback for contexts where the Clipboard API isn't
available (e.g. plain-HTTP self-signed-cert access).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Devices/rooms trigger a dialplan rebuild themselves via the web admin
now, but that only fixes the problem going forward — endpoints added
before that fix (or by any future path that misses the call) stay
registrable-but-uncallable with no obvious cause until someone thinks
to run --rebuild-dialplan by hand.
Call it unconditionally once Asterisk is up, before the PJSIP
transport check. Cheap and idempotent — it just regenerates
extensions.conf from the current pjsip.conf/rooms.conf state.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
The web admin's device functions (add_device, delete_device,
rename_device, change_device_category) only ever called
"asterisk -rx pjsip reload" — they never regenerated
extensions.conf's [intercom] context, so newly added SIP endpoints
could register but could never call each other or dial into rooms
("extension not found in context 'intercom'").
The room functions (create_room, delete_room, rename_room,
update_room_members) already tried to fix this correctly by shelling
out to `easy-asterisk --rebuild-dialplan`, but that flag was never
actually wired up — main() at the bottom of the script ignores all
arguments and always launches the interactive menu, so every one of
those calls was a silent no-op too.
Fixed both: added real --rebuild-dialplan argument handling that
calls the existing rebuild_dialplan() bash function non-interactively,
and added the same subprocess call to the four device functions that
were missing it entirely.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
pjsip show transports was checked once, immediately after "core show
version" first responded — but res_pjsip can take a moment longer to
finish binding its transports, so the check would sometimes read an
empty transport list and print "NOT LOADED" even though transport-tls
came up correctly a second later (confirmed live: TLS SIP traffic on
5061 in the container logs right after the misleading banner).
Poll for up to 10s instead of checking once, matching the existing
core-show-version wait pattern further up the same script.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Renamed services/asterisk-do.sh -> services/asterisk-digital-ocean.sh
(register_service name, install function, install dir, and all prose/
comments) so the whiptail menu shows a clearer, more discoverable name.
Updated the functional cross-references that depend on the old name:
crowdsec.sh's SIP-log auto-detection path and acquisition filename,
caddy.sh's host.docker.internal comment, and the CLAUDE.md/README.md
docs (services table, directory listing, network-wiring example).
Container names, the Docker Compose project name, and the internal
_asterisk_do_* helper function identifiers are left unchanged since
they aren't user-facing and renaming them would add risk for no
benefit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015X1jRGHwrvovz2qkhKfDZi
Same pattern already used for NetBird: a simple prompt (default yes,
since these are the two explicitly called out as recommended) right
after the mandatory package/Docker/SSH setup, before the whiptail
menu. Both stay fully optional and available later from the menu
either way — this just surfaces them earlier as a nudge, matching
how most other services in this repo end up wanting a reverse proxy
and something watching for brute-force/scan traffic.
asterisk-do previously offered to auto-install base, Caddy, CrowdSec,
and a numbered extras menu (authelia/ntfy/watchtower/wg-easy/netbird/
backup) on top of its own setup, layering a second install flow on
top of the whiptail menu setup.sh already provides. Strips all of
that back out — asterisk-do now only installs Asterisk + coturn, same
scope as any other service. Caddy/Authelia integration (reverse
proxy, cert sync, SSO) is kept, since it only activates when those
are already installed — no auto-install behind it. CrowdSec SIP
protection still wires up automatically via crowdsec.sh's own
asterisk-do detection, regardless of which one installs first.
Also fixes a real regression from ensure_caddy_network (added
earlier): it created caddy_net via a bare `docker network create`,
which doesn't carry Compose's ownership labels, so Caddy's own
non-external network declaration conflicted with it and failed to
start ("network exists but was not created by compose"). Caddy's
compose file now declares caddy_net as external: true like every
other service, since ensure_caddy_network is the single creator for
all of them, Caddy included.
43 services declare caddy_net as "external: true" in their compose
file, meaning they require it to already exist — but only Caddy's own
compose file actually creates it (authelia.sh was the sole exception,
with its own inline check-and-create). Installing any of the other 42
before Caddy fails outright with "network caddy_net declared as
external, but could not be found."
Adds ensure_caddy_network to lib/common.sh, called from require_docker
(which every install_* function already calls first), so the network
exists regardless of install order without touching each service file.
Removes authelia.sh's now-redundant duplicate of the same check.
Also documents in CLAUDE.md that network_mode: host services (asterisk/
asterisk-do) need host.docker.internal, not localhost, when Caddy
reverse-proxies to them — the fix from the previous commit.
Caddy runs in its own container on the caddy_net bridge network, so
"localhost" in a Caddyfile site block resolves to Caddy's own
container — never the host, and never a sibling container. That
broke every reverse proxy pointed at a network_mode: host service
(confirmed live with asterisk-do's web admin): once nothing else
(like a forward_auth redirect) intercepted the request first, Caddy
couldn't actually reach the upstream.
- services/caddy.sh: add extra_hosts so host.docker.internal resolves
inside the Caddy container (Linux Docker needs this explicitly —
it's automatic only on Docker Desktop).
- lib/common.sh's configure_caddy_for_service: bare-port upstreams
(its documented "host-network service" case) now target
host.docker.internal instead of localhost.
- services/asterisk-do.sh: its self-contained Caddy block (doesn't go
through configure_caddy_for_service) gets the same fix for local
Caddy, and now correctly targets the droplet's public IP instead of
localhost for the remote-Caddy snippet case, which had the same bug.
services/asterisk.sh needs no direct change — it already goes through
configure_caddy_for_service, so it inherits the fix.
If an existing ~/ubuntu-post-install checkout has a broken/SSH-only
origin remote, `git pull --ff-only` fails and the script fell through
to "continuing with existing version" — even when that existing copy
is missing setup.sh entirely, guaranteeing a crash right after. Now
checks for setup.sh post-pull and wipes + re-clones over HTTPS (no SSH
key needed) if it's still missing.
Replaces the y/n "update in place?" prompt in asterisk.sh/asterisk-do.sh
with an explicit r/f/c choice — (r)einstall in place, (f)ull install,
(c)ancel — defaulting to cancel on a bare Enter (or Ctrl-D) instead of
falling through to a destructive full reinstall.
Adds prompt_reinstall_mode to lib/common.sh (plus matching standalone
stubs in both asterisk scripts for when they run without the full repo)
and documents the convention in CLAUDE.md: any service with a persistent
install directory should offer this choice on rerun instead of re-asking
every prompt just to pick up a script fix.
Re-running either installer on an existing install used to re-ask
every prompt (domain, extras, firewall, Authelia) just to pick up a
script fix like the exports mount. Both now detect an existing
docker-compose.yml + .env and offer to update in place instead: only
vendor files and docker-compose.yml are refreshed and the stack is
rebuilt, leaving .env, firewall rules, and Caddy/Authelia config
untouched.
The vendor-copy and docker-compose.yml generation blocks (previously
inline and duplicated between what would have been two near-identical
code paths) are factored into per-file helper functions
(_asterisk_do_refresh_vendor_files/_asterisk_do_write_compose and
_asterisk_refresh_vendor_files/_asterisk_write_compose) so fresh
installs and updates share one copy of the logic instead of drifting
apart — the same problem that caused the /root export path and the
vpn-diagnostics.sh COPY bug to slip through unevenly between the two
services in the first place. Names are per-file since setup.sh sources
every services/*.sh into one process.
The vendor easy-asterisk script hardcodes /root for both export output
and its import file listing, but nothing was mounted there — exports
were being written to the container's ephemeral filesystem and lost on
recreate. Bind-mount ./exports to /root in both asterisk.sh and
asterisk-do.sh so exports/imports land under ~/docker/<service>/exports
on the host.
Confirmed on a real deployment: the template Caddyfile ships with
"admin off" (deliberate — no local API attack surface), which means
`caddy reload` can never work, since it depends on that same admin
endpoint. Every Caddyfile-editing code path was silently failing to
apply changes as a result — `docker logs caddy` showed
"admin endpoint disabled" and the reload command errored, but the
Caddyfile edit itself (which doesn't need the admin API) had already
succeeded, leaving the running config stale until something else
happened to restart the container.
Fixed in the two places that actually matter here: lib/common.sh's
configure_caddy_for_service (used by asterisk.sh and most other
Caddy-fronted services in the full repo) and asterisk-do.sh's own
self-contained Caddy block (both the standalone-bootstrap stub and the
main path). Each now tries the lightweight reload first — harmless,
and still works if a box ever has the admin API enabled — then falls
back to `docker restart caddy` if that fails, rather than leaving an
edited-but-unapplied Caddyfile.
Not fixed: the same duplicated pattern in ~35 other service files that
carry their own standalone-bootstrap copy of this logic. Those only
matter for the rare single-file standalone execution path for each of
those specific services and are unrelated to tonight's actual issue —
out of scope here.
Verified: full regression run on both asterisk.sh and asterisk-do.sh
still completes cleanly end to end.
The 8080->8081 fix from the last commit just moved the collision
risk, not removed it — any hardcoded port can eventually collide with
something else on a box running several services. Both services now
scan for the first genuinely free port starting at 8081 (ss -tlnH
"sport = :$PORT", capped at 100 ports checked) and use whatever they
find — .env, UFW, the DO Cloud Firewall rule, and the Caddy proxy
target all follow the actual chosen port, not a fixed number.
asterisk-do.sh's self-contained Caddy block (unquoted heredoc) reads
the port live. asterisk.sh's README heredoc is quoted (no expansion),
so its generated docs keep the static "8081" default with an added
note to check .env for the real value if it differed — the summary
echo outside that heredoc still reports the live value correctly.
Verified: normal case still lands on 8081; with 8081 deliberately
occupied by another process, both services correctly detect the
collision and fall through to 8082 instead, confirmed via the actual
generated .env in each case.
Real-world failure: CrowdSec's Local API listens on 127.0.0.1:8080 by
default (confirmed against its actual upstream config.yaml), and Easy
Asterisk's web admin also defaults to 8080. Both services in this repo
run with network_mode: host / directly on the host, so whichever one
starts second gets "OSError: [Errno 98] Address already in use" — in
this case CrowdSec (started earlier via the auto-install chain) had
already claimed the port before the web admin tried to start.
Moved the web admin's default to 8081 in both asterisk-do.sh and
asterisk.sh — WEB_ADMIN_PORT in .env, the UFW rule, the DO Cloud
Firewall rule, the Caddy reverse_proxy target, and every doc/summary
reference. 8081 doesn't collide with anything else in either stack
(5060/5061/8088/8089/3478/10000-20000/49152-49252) or with CrowdSec's
LAPI (8080) or Prometheus metrics (6060, localhost-only either way).
Left the vendor files' own internal fallback (WEB_ADMIN_PORT:-8080)
untouched — .env's explicit value overrides it at runtime regardless,
and vendor/ stays pristine per this repo's convention.
Verified: no stray 8080 in any generated .env/docker-compose.yml for
either service after a full install run; the vendor files' own
internal 8080 fallback (never applies here, since .env always sets it
explicitly) is the only remaining occurrence anywhere.
The Dockerfile COPYs scripts/vpn-diagnostics.sh and
scripts/dns-whitelist.sh into the image, but the vendor-file-copying
step in both asterisk.sh and asterisk-do.sh never copied (or
downloaded, in the GitHub-fallback branch) that scripts/ directory —
only Dockerfile, entrypoint.sh, coturn-entrypoint.sh, and the
management script. Every real install hit "docker compose up -d
--build" failing with:
failed to compute cache key: ... "/scripts/dns-whitelist.sh": not found
Confirmed live on a deployed droplet. vendor/easy-asterisk/scripts/
already has both files — this was purely a missed copy step, not a
vendoring gap. Fixed in both files identically (mkdir scripts/, copy
or curl both scripts, chmod +x alongside the existing executables).
Verified at the filesystem level: after a full install run, both
files land in the build context with correct executable permissions,
resolving the exact COPY instructions that were failing. Full
docker build verification wasn't possible in this sandbox (a separate,
unrelated network restriction blocks pulling the ubuntu:24.04 base
image here), but the missing-file root cause is directly fixed.
Real-world failure: configure_caddy_for_service's own domain prompt
defaults to "<subdomain>.${SITE_DOMAIN}", which only equals
$DOMAIN_NAME if SITE_DOMAIN happens to be set to match. In practice
SITE_DOMAIN is never set when this service is run by name (e.g.
`sudo ./setup.sh asterisk-do`), since that path skips setup.sh's own
site-defaults wizard — so the reconstructed default silently came out
wrong/blank, and a user had to guess whether to type the SIP domain or
something else at a bare "Domain [ ]:" prompt.
There's exactly one correct domain for this site block — $DOMAIN_NAME,
the same one already used for SIP — so it's no longer asked for at
all. This inlines the same Caddyfile-writing logic
configure_caddy_for_service uses (backup, dedup check, reload; local
and remote-Caddy modes both preserved) but targets $DOMAIN_NAME
directly. The only remaining question is a plain yes/no to proxy it.
Verified the block-generation logic directly against a real Caddy
directory + domain (produces the exact expected Caddyfile entry), plus
a full end-to-end regression run.
Typing out keyword names (e.g. "netbird backup") was more friction
than necessary. Now a numbered list (1-6), answered as comma-separated
digits with an example shown ("Example: 5,6"), translated internally
back to the same space-separated keyword string every existing
dispatch check (authelia/ntfy/watchtower/wg-easy/netbird/backup) was
already matching against — so none of those call sites needed to
change. Handles spaces after commas and silently ignores invalid
entries rather than erroring. Verified the number-to-keyword mapping
in isolation across normal input, spacing variants, invalid digits,
and blank, plus a full end-to-end regression run.
Step 7 (ntfy ban alerts) always defaulted straight to the public
ntfy.sh, regardless of whether the box (or a homelab) already had a
real ntfy instance. Confusing in practice: this step runs before
asterisk-do's own ntfy extra is dispatched, so even selecting it
wouldn't have helped at prompt time.
Now checks the local ntfy install's own config/server.yml for a
configured base-url (skipping it if it's still the ntfy.sh-written
placeholder) and uses <base-url>/crowdsec-alerts as the default. If
there's no configured local instance, it says so explicitly and
prompts toward a hosted instance elsewhere (e.g. a homelab) instead of
silently assuming the public service. Verified all three cases
(configured local, unconfigured placeholder, none) in isolation, plus
a full regression run.
Naming a service directly (sudo ./setup.sh asterisk-do) bypasses
setup.sh's own first-run base step entirely — essential packages, SSH
key import, disabling password auth. Docker still gets installed
either way (asterisk-do's own require_docker handles that), but the
SSH-hardening part of this setup's security story was silently
skipped on a genuinely fresh droplet unless the user knew to run
`base` separately first.
Checks the same marker setup.sh itself uses for "is base installed"
(command -v ncdu) and offers to run install_base directly if not —
same cross-service-call pattern already used for Caddy/CrowdSec/etc.
Verified end-to-end in a real sandbox run: base actually installed
packages, and execution correctly continued through the rest of the
asterisk-do flow afterward.
Both let the DO droplet lean on services already running on a
homelab instead of duplicating them locally, per the RAM-budget
discussion (Authelia+Redis and a second CrowdSec LAPI+DB add up).
crowdsec.sh: new step lets this agent register against a remote LAPI
(cscli lapi register -u <url>) and disables its own local API server
by removing the api.server block from config.yaml (backed up first;
verified the exact block boundaries against CrowdSec's actual default
config.yaml from upstream before writing the awk removal). Parsers,
scenarios, and the firewall bouncer still run locally regardless —
only banning decisions centralize, and only after the registration is
approved with `cscli machines validate` on the central machine, which
this script can't do since that's a different box. The final restart
step is skipped with an explanation when registration is pending,
instead of showing a misleading "failed to restart" for an expected
state.
asterisk-do.sh: when no local Authelia is installed, the web-admin
Caddy step now offers a remote Authelia option instead, building the
same forward_auth block inline (authelia.sh's shared Caddy snippet
only exists for local installs) targeting either a bare host:port
(e.g. a NetBird mesh IP) or a full https:// URL. Documents that this
couples web-admin availability to the remote instance's reachability,
while SIP/calling on the droplet stays unaffected either way.
Both changes verified: the config.yaml block-removal awk logic tested
against CrowdSec's real upstream default file structure, the remote
Authelia forward_auth block construction tested in isolation, and
full regression runs confirm the default (declined) path through both
new prompts is unchanged.
Adds a 'netbird' keyword to the existing extras prompt, dispatching
services/base.sh's _base_setup_netbird helper — a plain function like
any other once setup.sh sources every services/*.sh file, despite its
underscore-prefixed, not-independently-registered naming. Its own
prompt already defaults to enabling NetBird's built-in SSH server
(--allow-server-ssh), which is what makes the 'backup' extra usable
against a home machine without port-forwarding a router: install
NetBird here and on that machine, join both to the same network, and
Borg's SSH remote target becomes the home machine's mesh IP instead of
a public address. Skips cleanly if NetBird's already installed.
README's Optional extras section documents the pairing.
Extends the self-contained pattern from Caddy/CrowdSec to five more
services, offered through one consolidated "Install:" prompt instead
of five separate interruptions:
- authelia: only offered if Caddy is present (it's useless without
Caddy's forward-auth snippet); dispatched right where Caddy's state
is already known.
- wg-easy: installed alongside the other firewall rules so its port
lands with them. Only 51820/udp (the VPN handshake) goes on the
public firewall — the web UI (51821) is deliberately left closed,
documented as reachable via SSH tunnel instead, since exposing a
VPN's own admin panel publicly is a real foot-gun.
- ntfy, watchtower: independent, dispatched after CrowdSec. Watchtower
section is explicit that it only benefits coturn (a pulled image) —
Asterisk is a local Dockerfile build with no registry tag to check.
- backup (borg-backup): dispatched last. Documented clearly as a
config/data backup to a local machine or SSH remote, not a full
droplet image — the alternative to DO's paid Droplet Backups.
Every sub-install this calls does its own `cd` into ~/docker/<name>;
each call site restores `cd "$EA_DIR"` afterward so the later bare
`docker compose up -d --build` still targets the right directory.
Verified in isolation (mocked cd side effects) since driving five
real interactive sub-installs through piped stdin isn't practical.
README updated with an "Optional extras" section covering all five.
Self-contained by default now: if Caddy or CrowdSec aren't already on
the box, asterisk-do offers to install them itself (calling their
install_ functions directly — setup.sh sources every services/*.sh up
front, so they're already in-process during a wizard run). Standalone
single-file runs get a manual pointer instead, since those functions
don't exist outside the full repo checkout.
Also fixes the confusing "Configure Caddy reverse proxy for Asterisk
Web Admin" domain prompt: it used to ask for a second, independent
domain, which silently breaks the TLS cert sync if it doesn't match
the SIP FQDN exactly (Caddy only holds a cert for the domain it's
actually serving). It now always reuses the SIP FQDN automatically —
reconstructing configure_caddy_for_service's subdomain default so the
common case (SIP domain is a subdomain of SITE_DOMAIN) needs zero
extra input, with clear wording either way. FQDN prompt, README, and
final summary updated to match.
Vendor's logger.conf only sent Asterisk's security-level log lines
(auth failures, SIP registration scanning) to the console, i.e.
Docker's stdout — not a file CrowdSec could tail. asterisk-do.sh now
patches its copy of entrypoint.sh (vendor/ untouched) to also write
those events to /var/log/asterisk/full, which is bind-mounted to
~/docker/asterisk-do/logs/full on the host.
crowdsec.sh now detects that directory and, if present, installs the
crowdsecurity/asterisk collection (asterisk_bf + asterisk_user_enum
scenarios) with a matching log acquisition — mirroring the existing
Caddy detection pattern. Order-independent: asterisk-do's install
summary tells the user to rerun crowdsec if it's already installed,
since detection only runs during crowdsec's own install step.
DigitalOcean doesn't provision swap by default and the $4/mo (512MB)
droplet has little headroom once Docker + Asterisk + coturn are
running. The installer now detects RAM <=2GB with no existing swap and
offers to add a persistent 2GB swapfile before doing anything else, so
that tier is safe to use instead of risking an OOM kill under load.
README updated with the corrected sizing table.
Duplicates services/asterisk.sh (left untouched) into a DO-specific
variant: auto-detects the droplet's public IP/ID via the DO metadata
service, always assumes a public FQDN (no LAN/VLAN prompts), offers to
provision a matching DigitalOcean Cloud Firewall via doctl (never
touching one that's already attached), and documents droplet sizing,
firewall rules, and Sipnetic client setup in the generated README.
install.sh generates the sunrise, sunrise-upload, seasons, and moon jobs
as Type=oneshot with only OnFailure=notify - a transient ffmpeg/network
blip fails the whole day's job with just an alert, no retry.
Add systemd drop-in overrides (Restart=on-failure, RestartSec=60,
StartLimitBurst=3 within a 10 min window) for each of these units after
install.sh runs. Drop-ins live outside the files install.sh generates,
so they survive re-running install.sh (e.g. after editing
sky-cam.conf), unlike a direct edit to the generated unit which would
be silently overwritten next time. systemd only fires OnFailure once
retries are exhausted, so this doesn't add notification spam - just
one alert after 3 tries, 60s apart.
capture.sh/capture-watchdog.sh already have Restart=on-failure baked
into install.sh's own generation (Type=simple, long-running) and don't
need this.
Nothing in the repo actually installed the NVIDIA driver or
nvidia-container-toolkit — ai-gpu.sh, wolf.sh, etc. all assumed both were
already present. Adds _base_setup_nvidia_gpu, called during base install
right after Docker:
- No-ops silently on boxes without an NVIDIA GPU (lspci VGA/3D controller
check) so non-GPU installs are unaffected
- If a GPU is present but nvidia-smi isn't working, offers to run
'ubuntu-drivers devices' (shown to the operator) then
'ubuntu-drivers autoinstall', and warns a reboot is required
- If Docker is present and nvidia-container-cli is missing, offers to
install NVIDIA Container Toolkit and run
'nvidia-ctk runtime configure --runtime=docker' so GPU-accelerated
Docker services (ai-gpu, wolf, paintplus, iopaint) can request the GPU
- Offers to reboot immediately if a driver install requires it
Verified with a mocked-lspci/nvidia-smi/ubuntu-drivers test harness across
three scenarios: no GPU (silent no-op), GPU with no driver (full install +
toolkit + reboot prompt flow), and GPU with driver already active (skips
driver prompt, still offers toolkit).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Lets 'ssh <alias>' connect directly to user@host instead of retyping it —
especially useful once machines are reachable over NetBird/VPN and have
IPs that aren't worth memorizing.
- lib/common.sh: ssh_config_path/add_ssh_host_alias/list_ssh_host_aliases/
remove_ssh_host_alias helpers, operating on the invoking user's own
~/.ssh/config (not root's) with correct 700/600 permissions and ownership
- base.sh: after SSH key import, optionally add one or more Host aliases
interactively as part of the base install
- services/ssh-config.sh: new standalone service (sudo ./setup.sh ssh-config)
to list/add/remove aliases any time, independent of base install; follows
the existing non-Docker standalone-bootstrap pattern (see crowdsec.sh)
- setup.sh: ssh-config never shows [installed] since it's a repeatable
management tool, not a one-time install
- README: new 'SSH Host aliases' section, base row and wizard-flow step 1
updated, ssh-config added to the extras group and copiable service list
Verified end-to-end with a test harness: add with defaults, add with a
custom user/port, list (correct numbering), and remove-by-name preserving
the other entry and file permissions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
The wizard description was stale — it still described the old
site-defaults-first flow and didn't mention that base now installs Docker,
openssh-server (with SSH key import), and NetBird, or that the wizard ends
by dropping into a fresh login shell so the docker group takes effect.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
capture.sh already replaces motionEye/any NVR itself - it just needs
each camera's RTSP URL. services/sky-cam.sh never actually prompted for
CAM_RTSP_<cam>, so capture/audio never had anything to connect to.
- Prompt per camera for its RTSP URL -> CAM_RTSP_<cam> in .env
- Prompt for sunrise mic / optional ambient audio library
- Fix Mattermost integration: sunrise2mm.py reads mattermost_url/
access_token/channel_id (bot-token REST upload), not the
MM_WEBHOOK_URL/MM_CHANNEL incoming-webhook scheme the installer used
to write - uploads never worked before this
- Add optional ntfy push notifications
- Auto-generate SCHEDULE_SEASONS_<cam> (staggered 30 min apart) for
every configured camera, not just the stock east/north/south, so
install.sh wires up every applicable systemd timer for any camera set
Removes services/sky-cam-frigate.sh entirely - routing sky-cam's frames
through Frigate (via export API or restream) turned out to be solving a
problem that doesn't exist; sky-cam's own capture.sh talking directly to
each camera is simpler and has no quality/resolution tradeoffs. Frigate
continues to run fully independently for NVR/detection.
The real sky-cam repo's capture.sh already replaces MotionEye/any NVR
itself (plain ffmpeg RTSP frame-grab) - it and daily_sunrise_video.sh's
optional audio capture are the only places that touch a camera's RTSP
URL directly. Every other script (4-seasons, montage-mvt, year-end-join,
moon-track, moon-phase-monthly) only reads JPEGs/audio already on disk.
So the entire motionEye->Frigate transition is pointing CAM_RTSP_<cam>
at Frigate's go2rtc restream (rtsp://<frigate-host>:8554/<cam>) instead
of the camera directly - no upstream script changes needed. Replaces
the previous frigate-retime.sh/export-API approach, which solved a
problem (matching an arbitrary recording length to music duration) that
sky-cam's own 4-seasons.sh/montage-mvt.sh already handle via JPEG frame
counts.
Duplicates services/sky-cam.sh into a Frigate-backed variant that pulls
recordings via Frigate's export API instead of a JPEG image folder.
Includes a frigate-retime.sh helper that exports a coarse timelapse,
measures its actual duration with ffprobe, and re-encodes once with a
computed setpts factor to hit an exact target length (e.g. a Four
Seasons movement's runtime).
Group membership added by 'usermod -aG docker' (in require_docker) doesn't
apply to the shell that invoked sudo — only to new logins. Users had to
manually run 'newgrp docker' or reconnect SSH after every install. Since a
child process can't change its parent shell's group list directly, the
practical fix is to exec a fresh 'su - ' login shell at the end
of the guided flow, which re-reads /etc/group and lands the user back in
the same terminal with docker access already active.
Gated on: running via sudo (SUDO_USER set), interactive (not --unattended),
docker group exists and the user is actually a member, and stdin is a real
tty — so this never fires for scripted/explicit-service/piped invocations.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Previously re-running the installer always overwrote config.yml and .env
from scratch, silently discarding any real camera credentials already on
disk. Now install_frigate parses an existing config.yml + .env (best-effort,
matching this installer's own output shape) and presents a numbered list
of detected cameras with a menu:
[1] Keep everything as-is (no changes at all)
[2] Backup existing config and start fresh
[3] Add more cameras (keep these)
[4] Remove cameras (choose numbers, or 'all'), then optionally add more
Implementation switches from building config.yml/.env as concatenated text
blocks inline in the collection loop to parallel CAM_* bash arrays
(name/ip/port/var-names/enabled/notify/substream-suffixes), so cameras can
be parsed, listed, removed, and re-rendered independently:
- _frigate_parse_existing: reads go2rtc streams + cameras: enabled/notifications
from config.yml, and credential values from .env, into the CAM_* arrays
- _frigate_review_existing: numbered menu, mutates arrays per choice
- _frigate_next_suffix_int: kept cameras retain their existing FRIGATE_RTSP_USER[N]
var names unchanged; new cameras get the next unused numeric suffix so
credentials never collide after removals
- _frigate_camera_wizard / _frigate_render_config: same prompts and output
shape as before, now array-driven so kept + new cameras render uniformly
- _frigate_backup_existing: copies config.yml/.env/docker-compose.yml to a
timestamped backup-YYYYMMDD-HHMMSS/ dir before any destructive rewrite
Verified with a 5-scenario test harness (fresh install, keep-as-is producing
byte-identical output, add-camera preserving existing credentials, remove-
by-number without var collisions, backup-and-fresh) before committing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Previously Frigate always wrote a placeholder config.yml the operator had
to hand-edit to add cameras. Now install_frigate prompts to add cameras
one at a time (name, RTSP IP/port/user/password/path, optional sub-stream,
enabled, notifications), matching the go2rtc + cameras structure used in
production frigate configs:
- Each camera gets a go2rtc stream entry (+ optional _sub for detection)
and a cameras: block with ffmpeg inputs/roles, detect, notifications
- RTSP credentials/IPs are written to .env as FRIGATE_* variables (first
camera gets FRIGATE_RTSP_USER/PASSWORD, later cameras get numbered
suffixes _1, _2, ... to avoid collisions) and referenced in config.yml
via Frigate's {FRIGATE_VAR} substitution syntax — secrets never appear
in the YAML directly
- docker-compose.yml now includes env_file: .env so those vars actually
reach the container for substitution to work
- Skipping all camera prompts falls back to the original starter
config.yml for manual editing, preserving existing behavior
- README and DRY-RUN summary updated to reflect the new flow
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Defaults install_asterisk() to FQDN networking mode and prompts for VLAN/VPN
subnets (with host-network auto-detection to filter out noise like Docker
bridges) so phones on other networks get correct NAT/SDP handling from the
first boot.
The container now mounts Caddy's cert store read-only when Caddy is
installed, and the entrypoint syncs a matching Let's Encrypt cert for
DOMAIN_NAME automatically, re-checking every 12h to pick up renewals without
a restart. Falls back to self-signed only when no matching cert is found.
Also fixes a real bug hit in the field: a preserved/migrated pjsip.conf could
be missing the transport-udp/transport-tcp sections entirely, with no bind
error logged, silently blocking any device that registers without TLS. Adds
the same migration-injection already used for transport-tls.
The 'Install Caddy now?' prompt ran regardless of the just-answered
Caddy location question, so choosing 'remote' still asked whether to
install Caddy locally — contradicting the choice made one prompt earlier.
Gate it on CADDY_MODE being local (or unset, for configs predating the
wizard split).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Previously the Caddy-location question lived inside run_site_configure,
gated behind 'Configure site defaults now? (y/n)'. Answering 'n' (e.g.
because Caddy is on a different box and you don't care about domain/tz
autofill) meant CADDY_MODE never got set, which silently disabled Caddy
prompts for every service for the life of the install (configure_caddy_for_service
falls through to mode 'none' and returns immediately).
Split into two steps:
1. ask_caddy_location() — always runs on first setup.sh invocation,
independent of any other prompt, and persists CADDY_MODE immediately.
2. run_site_configure() — now only asks timezone/domain/Caddy-network,
and is only offered when CADDY_MODE=local (those defaults are only
useful for FQDN autofill tied to a locally-managed Caddyfile).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Reorders the site defaults wizard so 'Where does Caddy run?' comes before
timezone/domain, since it's the more fundamental choice and the answer
context matters when explaining the other prompts. Also skips the Caddy
Docker network prompt entirely when Caddy isn't running locally — that
setting is only relevant to services joining a local Caddy container's
bridge network; remote/none mode proxies via localhost:PORT + snippet
files instead.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
The get.docker.com convenience script internally wraps every step in
'sudo -E sh -c ...'. On minimal/cloud Ubuntu images that never installed
the sudo package (common when operating purely as root), those internal
sudo calls silently fail while the outer script still exits 0 — apt never
actually runs, but no error surfaces. require_docker already runs as root,
so there's no need for sudo at all.
Replaced it with Docker's documented apt-repo steps run directly: add the
keyring, add the repo (with architecture/codename detected via dpkg and
os-release), apt-get install docker-ce + compose plugin, enable the
service. Real apt/curl/systemctl failures now propagate and print to the
terminal instead of being masked by the wrapper script.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
require_docker returning non-zero was silently ignored (no set -e).
Add explicit warning so the operator sees the failure; setup.sh already
has an unconditional Docker check after base that will retry.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
The Docker check+install was inside the else branch that only runs when
base has never been installed. On re-runs (base already present) Docker
was silently skipped and only warned about. Move the check outside the
if/else so Docker is always installed if missing, regardless of whether
base was skipped.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
whiptail fix:
- bootstrap.sh: redirect stdout and stderr to /dev/tty alongside stdin so
whiptail has full terminal control for raw mode (arrow keys, highlighting)
- setup.sh: run 'stty sane' on /dev/tty before the menu loop to reset any
stale terminal state from SSH reconnections or prior sessions
Installed-service summary:
- Print a grouped list of all currently-installed services before every
menu session so the operator knows the current state at a glance
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Two fixes:
1. Export TERM (default xterm-256color) early — whiptail needs a valid
TERM to enter raw mode; when bash is started via pipe TERM may be
unset, causing keypresses to leak to the shell instead of the menu
2. Add </dev/tty to both whiptail calls so keyboard input always comes
from the controlling terminal regardless of how stdin was redirected
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Three fixes:
1. configure_caddy_for_service: remove the '!= example.com' filter that
silently dropped any valid domain matching that string; now any non-empty
SITE_DOMAIN is used as the default subdomain suggestion
2. load_site_config: trim leading/trailing whitespace from key and val so
hand-edited .config files with extra spaces still parse correctly
3. setup.sh: call load_site_config after the site wizard saves so the
in-memory values are guaranteed fresh for all subsequent service installs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
- require_docker now runs as part of base so Docker is present on every box
- Install openssh-server, offer GitHub (gh:) and Launchpad (lp:) key import
via ssh-import-id; disable password auth only after keys are confirmed imported
- Handle Ubuntu cloud-init drop-in that re-enables PasswordAuthentication
- Offer NetBird install with optional --allow-server-ssh flag and setup key
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
command -v may miss the binary if sudo stripped PATH; check the canonical
apt install location directly as a fallback before reporting failure, and
use the same fallback when printing the installed version.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Two bugs in require_docker:
1. apt post-install hooks (needrestart etc.) block on stdin which is
at EOF when running via pipe; DEBIAN_FRONTEND=noninteractive skips them
2. bash's command hash table doesn't pick up a newly installed binary;
hash -r flushes it so command -v docker finds /usr/bin/docker
Also moved usermod and success log after the binary check so [OK] only
prints when docker is actually reachable.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
curl | bash consumes stdin from the pipe, so when bootstrap hands off
to setup.sh the script gets EOF immediately and exits with 'Cancelled'.
Redirecting </dev/tty restores keyboard input for the interactive menu.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
BASH_SOURCE[0] is unbound when bash reads from a pipe; set -u turns
this into a fatal error that no amount of :- or set +u reliably fixes
across bash versions. $0 is always set: 'bash' when piped (dirname
gives '.' where no setup.sh exists), and the correct path when run
directly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
${BASH_SOURCE[0]:-} still triggers set -u when BASH_SOURCE is entirely
unset (not just empty) in pipe mode. Temporarily disable -u for that
single assignment, then restore it.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
set -euo pipefail causes ${BASH_SOURCE[0]} to abort with 'unbound variable'
when the script is fed via curl | bash. Use ${BASH_SOURCE[0]:-} so the
variable expands to an empty string in that context, letting SCRIPT_DIR
resolve safely and the pipe path continue to the git-clone branch.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LQJBvqzXeyuhhAcAA3Q5Wq
Three root-cause fixes found by reading the Docker image source:
1. Wine/Proton page fault: Docker's default seccomp profile blocks
system calls that Wine Proton GE requires. Fix: security_opt:
seccomp=unconfined + shm_size: 256m (Xvfb needs /dev/shm for
MIT-SHM extension; 64 MB default is too small).
Added network_mode: host for game traffic (dynamic UDP ports).
2. KYBER_MAP_ROTATION exit 64: the Kyber CLI decodes base64 and
parses newline-separated "MODE;MAP_PATH" lines, not JSON objects.
Our JSON [{map:...,mode:...}] format split on semicolons into one
field → ExitCode.usage (64). Fixed builder to emit MODE;MAP_PATH\n
lines and updated .env comment + README example.
3. GPU passthrough removed: Proton GE includes DXVK which crashes
headlessly when a GPU is passed through (no Vulkan display). The
server needs no GPU; removing passthrough is the correct fix.
Also: install libgamemode0:i386 on the host (Wine/Proton dep),
add alphanumeric-password warning (special chars → INVALID_PASSWORD).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014be1aK9G8CY2msho5LjxR4
dall-e-2/dall-e-3 retired May 12 2026 and gpt-image-1 deprecates Oct 23
2026, so move every OpenAI default (config.py, both provider classes,
both compose files, .env.example, the in-app provider-settings dropdown,
README) to gpt-image-2 for both generation and edits. Also fix response
parsing in ai_provider.py's OpenAIProvider, which never sent a model
param and assumed a url response — gpt-image-1/2 only return b64_json.
Separately, .env.example shipped AI_PROVIDER=replicate by default, but
replicate has no driver in remote_provider.py, so following the
documented "cp .env.example .env" setup silently broke every AI call
and defeated the GPU quick-start (an explicit non-empty .env value
overrides docker-compose.gpu.yml's own local_gpu fallback). Default to
local_gpu instead, mark replicate/stability as not-yet-implemented, and
recommend Lykon/dreamshaper-8-inpainting as a hands/face-tuned
HF_MODEL_INPAINT override for 4-6GB cards (Quadro P2200, GTX 1060/1660).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nb2vJ8W7bHKx1JXVvpCraH
inpaint()/img2img()/outpaint() called /v1/images/edits without a model
field, so OpenAI defaulted every cloud edit to dall-e-2 regardless of
configuration — while txt2img used dall-e-3. Add a separate
OPENAI_EDIT_MODEL (default gpt-image-1, the only current model that
supports masked edits at ChatGPT-comparable quality), thread it through
the provider and both compose files, and handle gpt-image-1's
b64_json-only response shape alongside the url shape dall-e-2/3 return.
setup.sh's run_service() looks up install_<name> using the literal
hyphenated registered name, not an underscore-converted one. borg-backup,
calibre-web, gaming-backup, and stirling-pdf all used underscored function
names and were therefore uninstallable ("has no install_<name>"). Same
fix already applied to ai-gpu/ai-stack; this closes out the rest of the
repo-wide audit.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nb2vJ8W7bHKx1JXVvpCraH
PaintPlus's backend already supported invokeai/comfyui providers (generic
"self-hosted, on another machine" remote APIs) but the installer never
exposed them and the compose file never passed the URLs through. Add a
3rd provider choice — shown only when the ai-stack service is installed —
that sets AI_PROVIDER + INVOKEAI_URL/COMFYUI_URL and joins ai-stack's
Docker network (ai-stack_default) so PaintPlus can reach those containers
by name. No cloud key, no extra GPU download: it rides on ai-stack's
already-running InvokeAI/ComfyUI.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nb2vJ8W7bHKx1JXVvpCraH