The auto-sync condition "[STATUS] < 400" reads as red for any site behind
Authelia's forward_auth: Gatus's probe is never logged in, so it correctly
gets a 401 back every time — the site is completely healthy, Authelia is
just doing its job, but that 401 fails the condition. Confirmed live: every
site the user actually logs into showed permanently red.
That single condition also had the opposite bug in reserve: on a genuine
outage (connection refused, DNS failure, TLS failure), Gatus reports
[STATUS] as 0, and 0 < 400 is true — a fully unreachable site would have
silently read as "up".
Fixed to two conditions together: "[CONNECTED] == true" (catches the actual
outage case) and "[STATUS] < 500" (accepts any real response, including
401/403/redirects from an auth gate, only failing on Caddy's own 502/503/504
when the backend itself is unreachable).
Also changed the sync loop to refresh conditions on already-synced
endpoints, not just add-missing-ones — the old add-if-missing-only logic
meant this fix would only apply to newly discovered domains, leaving every
already-synced site (which is most of them, on a live box) stuck on the
broken condition forever until removed and re-added by hand. Now every
sync run (every 15 minutes via the existing systemd timer, or the one that
happens immediately on a Gatus reinstall) self-heals all of them.
Verified end-to-end against the real mikefarah/yq binary: an existing
caddy-sync entry gets its conditions rewritten in place, an unrelated
manually-added endpoint is left untouched, and a newly-discovered domain
gets the corrected conditions from the start.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
Adds one Gatus endpoint per Caddy site block automatically, tagged
group: caddy-sync — the sync only ever adds/removes entries in that
exact group, so anything added by hand (the default external checks,
a custom endpoint) is never touched regardless of what the Caddyfile
looks like. Offered at install time (syncs once immediately) and, if
systemd is available, scheduled via a timer every 15 minutes so a site
added or removed later gets picked up without re-running the installer
— matches the "schedule that checks the Caddyfile" shape asked for.
Domain extraction tracks actual brace depth (reusing the same approach
as remove_service's Caddy block removal) rather than a naive
line-by-line scan, so it correctly skips the global options block and
parenthesized snippet definitions like (authelia) without needing to
special-case them by name.
Verified end-to-end against a real Caddyfile/config.yaml fixture with
the actual mikefarah/yq binary: initial sync adds the right entries
and leaves the default "external" group alone, a second run with no
Caddyfile changes is a true no-op (0 added, 0 removed), and changing
the Caddyfile (removing one site, adding another) correctly adds the
new endpoint and removes only the stale one.
Also fixes a real gap surfaced while building this: ensure_yq (used by
both gatus.sh now and onlyoffice.sh already) checked `command -v yq`
alone, which a box can satisfy with a completely different, incompatible
yq — confirmed live in this environment, Debian/Ubuntu's own `yq`
apt package is kislyuk/yq (a Python jq-wrapper) which silently errors
on mikefarah/yq's `e '.path' file` syntax every caller here depends on.
Now checks the version string actually identifies as mikefarah's
before trusting it, installing to /usr/local/bin (which precedes
/usr/bin on Ubuntu's default PATH) if not. Promoted ensure_yq itself
from onlyoffice.sh (its only previous user) to lib/common.sh now that
gatus.sh needs the same thing, so both share one implementation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H4k6J1qXXyYxhGEgnJaMvn
With 70+ services sharing a handful of common default ports (emby and
jellyfin both default to 8096, changedetection and frigate both default
to 5000, arm and nextcloud both default to 8080...), nothing previously
checked whether a service's default port was actually free on the host.
Whichever service installed second would silently write a compose file
claiming an already-held port, only failing at `docker compose up` time.
Adds two shared helpers to lib/common.sh:
- port_in_use PORT [PROTO] — true if something's already listening
- find_free_port VARNAME START [PROTO] — scans upward, writes back the
first free port
Every service that publishes a fixed host port now scans before writing
docker-compose.yml, on every install (not just when adding an explicit
additional instance). On a normal single-install host this is a silent
no-op; it only changes behavior when something else already holds the
port.
- The 19 services already given multi-instance support this session had
their port scan moved out of the "add instance" branch to run
unconditionally, since the same collision risk exists on a plain first
install.
- 20 more services with previously-hardcoded ports gained scanning for
the first time: archivebox, arm, calibre-web, changedetection,
drum-rhythm-game, gatus, n8n, nextcloud, onlyoffice, stirling-pdf,
uptimekuma, portainer, iopaint (both GPU/CPU compose branches), koha
(paired), syncthing (paired), wg-easy (paired, plus WG_PORT env so
generated peer configs keep the right Endpoint), homeassistant
(bridge-mode only — host mode can only warn), frigate and
frigate-audio (multi-port stacks, moved together).
- caddy.sh is the deliberate exception: 80/443 stay fixed and only warn
on collision, since silently moving Caddy itself would leave nothing
listening where any client actually looks.
- authelia.sh needs no change — it has no published host port at all.
- Every service's standalone bootstrap fallback (sudo bash services/x.sh
with no sibling files) got the same two helpers duplicated into its
stub block, matching how every other shared helper is already handled
there.
Documents the full pattern in CLAUDE.md's new "Port collision avoidance"
section, including the quoted-heredoc/backtick-escaping gotcha and the
network_mode:host limitation (can only scan ports the app takes as a
configurable env var).
Verified via bash -n on every changed file, plus functional runs seeding
occupied ports for each collision shape used here (single, paired,
multi-port stacks) and confirming the scan/shift and generated
compose/README output are correct — including the emby/jellyfin,
nextcloud/arm, and frigate/changedetection collision scenarios that
originally motivated this.
traccar.sh's caddy_net wiring was fixed to mirror configure_caddy_for_service's
own mode resolution (CADDY_MODE from site config, then a local ~/docker/caddy,
then the legacy CADDY_REMOTE_HOST var) instead of only checking for the local
directory. That same bare directory check was copy-pasted into the caddy_net
wiring of every other Docker service in the repo, so a site with Caddy on a
different box would silently fail to join any of their containers to caddy_net
during setup (or, for homeassistant/koha, only get half the wiring right).
Applied the same fix mechanically across all 37 services using the standard
_CADDY_NET_BLOCK/_CADDY_NET_SECTION pattern (verified identical text via
scripted diff before touching any of them), plus by hand for:
- homeassistant.sh and koha.sh, which use their own differently-shaped
variables (HA_CADDY_NET_LINES / _CADDY_NET_ENTRY) for the same decision
- paintplus.sh and ai-stack.sh, which do a live `docker network connect`
instead of a compose network block
- watchyourlan.sh, whose Caddy note was worded for local-only setups
sms-inbound.sh got more than a mode swap: its Caddy wiring was hand-rolled
(not routed through configure_caddy_for_service) and had no remote-Caddy
path at all — a remote Caddy box would get a misleading "Caddy isn't
installed here" message instead of a snippet. Added
_sms_write_caddy_snippet(), mirroring the snippet-file pattern
configure_caddy_for_service uses everywhere else, and pointed the firewall
gate at the same three-way mode instead of a two-way dir check.
Verified: bash -n across all of services/*.sh, a scripted check that every
touched file has exactly one _CADDY_MODE resolution and no leftover bare
`[ -d "$DOCKER_DIR/caddy" ]` feeding a caddy_net decision, and spot-checked
docker compose config renders (traccar, mattermost) confirming the ${VAR}
interpolation and multi-service usage sites still resolve correctly.
New site config key: CADDY_REMOTE_HOST (set via 'sudo ./setup.sh configure').
When set, configure_caddy_for_service operates in "remote" mode instead of
writing to a local Caddyfile:
- Upstream uses CADDY_REMOTE_HOST:PORT (host IP, not container name)
- Snippet saved to ~/docker/caddy-snippets/<subdomain>.caddy
- User is shown scp/rsync commands to copy it to the Caddy machine
Three modes in configure_caddy_for_service (lib/common.sh and inline stubs):
local: ~/docker/caddy/ exists → write Caddyfile + reload (existing behavior)
remote: CADDY_REMOTE_HOST set → save snippet, print copy instructions
none: neither configured → silent return (unchanged)
All 31 service standalone bootstrap stubs updated with the new logic.
CADDY_REMOTE_HOST global added to all 42 standalone bootstrap sections.
setup.sh configure now prompts for CADDY_REMOTE_HOST with a clear explanation.
wolf.sh: add missing stubs (configure_caddy_for_service, write_readme,
prompt_yn, ensure_docker_dir_ownership) and the Authelia/Caddy/start calls
that were missing from the install function.
https://claude.ai/code/session_014CCYqVwW6d6f5dw1qRokYt
Each service can now be run directly with sudo bash <service>.sh on any
machine with Docker installed, without needing the full post-install repo.
Uses the shared bootstrap pattern from docs/standalone-template.sh.
https://claude.ai/code/session_014CCYqVwW6d6f5dw1qRokYt