Repository navigation
Proposal: Make WireGuard startup resilient to temporary DNS unavailability #2319
Description
Activity
I have already applied the second solution to all my cameras and will monitor them for a week.
Power outages occur at the camera locations with an unknown frequency, but every 1–2 days I find one or two cameras with WireGuard in a non-working state. The cameras sometimes boot faster than the network connection to the DNS server becomes available after a power outage, which previously left them without WireGuard.
If the fix proves reliable, this will support implementing this approach in the official firmware scripts.
Two days — everything has been working perfectly.
Just in case:/etc/init.d/S98wireguard
#!/bin/sh case "$1" in start) wgc=$(fw_printenv -n wg_privkey) if [ -n "$wgc" ]; then wireguard & fi ;; stop) ;; *) echo "Usage: $0 {start}" exit 1 ;; esac
/usr/sbin/wireguard
#!/bin/sh modprobe wireguard || { echo "Error: Failed to load wireguard module." >&2; exit 1; } ip link add dev wg0 type wireguard || { echo "Error: Failed to create wg0 interface." >&2; exit 1; } WG_PRIVKEY="$(fw_printenv -n wg_privkey)" WG_PRESHARED_KEY="$(fw_printenv -n wg_sharkey)" ENDPOINT=$(fw_printenv -n wg_endpoint) case "$ENDPOINT" in [0-9]*) ;; *) HOST="${ENDPOINT%:*}" while ! nslookup "$HOST" >/dev/null 2>&1; do sleep 10 done ;; esac ( echo "#" echo "[Interface]" echo "PrivateKey = $WG_PRIVKEY" echo echo "[Peer]" echo "Endpoint = $ENDPOINT" echo "PersistentKeepalive = $(fw_printenv -n wg_alive)" echo "PublicKey = $(fw_printenv -n wg_pubkey)" [ -n "$WG_PRESHARED_KEY" ] && echo "PresharedKey = $WG_PRESHARED_KEY" echo "AllowedIPs = $(fw_printenv -n wg_allowed)" echo "#" ) >>/tmp/wireguard.conf wg setconf wg0 /tmp/wireguard.conf || { echo "Error: Failed to apply wireguard configuration." >&2; exit 1; } wg_address="$(fw_printenv -n wg_address)" if [ -z "$wg_address" ]; then echo "Error: wg_address environment variable is not set or empty." >&2 exit 1 fi ip address add dev wg0 "$wg_address" ip link set up dev wg0 for i in $(fw_printenv -n wg_allowed | tr ',' ' '); do ip -4 route add "$i" dev wg0 done
As I suggested in the issue description, errors from initialization scripts should be logged to syslog rather than printed only to the console. I will continue testing a script implementing this approach:
#!/bin/sh PROGRAM="${0##*/}" run_cmd() { msg="$1" shift output=$("$@" 2>&1) rc=$? if [ "$rc" -ne 0 ]; then [ -n "$msg" ] && msg="$msg. " logger -p user.err -t "$PROGRAM[$$]" "Error: ${msg}" logger -p user.err -t "$PROGRAM[$$]" "Command \`$*\` finished with code $rc:" printf '%s\n' "$output" | while IFS= read -r line; do logger -p user.err -t "$PROGRAM[$$]" "$line" done kill -s TERM $$ 2>/dev/null exit "$rc" fi printf '%s\n' "$output" } run_cmd "Failed to load wireguard module" modprobe wireguard run_cmd "Failed to create wg0 interface" ip link add dev wg0 type wireguard WG_PRIVKEY=$(run_cmd "" fw_printenv -n wg_privkey) WG_PRESHARED_KEY=$(fw_printenv -n wg_sharkey 2>/dev/null || true) ENDPOINT=$(run_cmd "" fw_printenv -n wg_endpoint) case "$ENDPOINT" in [0-9]*) ;; *) HOST="${ENDPOINT%:*}" while ! nslookup "$HOST" >/dev/null 2>&1; do sleep 10 done ;; esac ( echo "#" echo "[Interface]" echo "PrivateKey = $WG_PRIVKEY" echo echo "[Peer]" echo "Endpoint = $ENDPOINT" echo "PersistentKeepalive = $(run_cmd "" fw_printenv -n wg_alive)" echo "PublicKey = $(run_cmd "" fw_printenv -n wg_pubkey)" [ -n "$WG_PRESHARED_KEY" ] && echo "PresharedKey = $WG_PRESHARED_KEY" echo "AllowedIPs = $(run_cmd "" fw_printenv -n wg_allowed)" echo "#" ) >>/tmp/wireguard.conf run_cmd "Failed to apply wireguard configuration" wg setconf wg0 /tmp/wireguard.conf wg_address=$(run_cmd "" fw_printenv -n wg_address) if [ -z "$wg_address" ]; then echo "Error: wg_address environment variable is not set or empty." >&2 exit 1 fi run_cmd "" ip address add dev wg0 "$wg_address" run_cmd "" ip link set up dev wg0 for i in $(run_cmd "" fw_printenv -n wg_allowed | tr ',' ' '); do run_cmd "" ip -4 route add "$i" dev wg0 done
If this approach were used in the current
/usr/sbin/wireguardscript in the official firmware, the following would appear in syslog if WireGuard starts while the DNS server is unavailable:Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Error: Failed to apply wireguard configuration. Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Command `wg setconf wg0 /tmp/wireguard.conf` finished with code 1: Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.00 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.20 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.44 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.73 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 2.07 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 2.49 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 2.99 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 3.58 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 4.30 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 5.16 seconds... Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 6.19 seconds... Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 7.43 seconds... Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 8.92 seconds... Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 10.70 seconds... Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 12.84 seconds... Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820' Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Configuration parsing errorSince VTun, if it starts while DNS is unavailable, fills all 64 KB of the syslog RAM buffer with these two messages:
Aug 29 09:05:21 gk7205v300-imx335 daemon.err vtund[21576]: syntax error line 5 Aug 29 09:05:21 gk7205v300-imx335 daemon.err vtund[21576]: No hosts definedI suggest applying the same approach to
/etc/init.d/S98vtunand/usr/sbin/tunnel: wait for DNS to become available before starting VTun.I think the testing is sufficient at this point. I haven't run into a single problem during testing.
I suggest that you develop and test the fix for VTun, so that it doesn't flood the syslog buffer, on your side.
Just wondering if this is still relevant.
PR are always welcome. I personally don't use either Wireguard, nor vtun and don't need them and am happy to remove them from the firmware.
Reacted by usa-Created: #2377
- added a commit that references this issue
on Sep 7, 2026 I am planning to raise one more PR based on this issue during next few days.
- added a commit that references this issue
on Sep 13, 2026 @usa- #2401 is merged — 3769227 on master, in every image from the next nightly. That's the follow-up you mentioned on 7 September, so both proposals from the issue body are now in: the DNS wait from #2377, and initialisation errors going to syslog rather than only the console.
Does that close this out for you? Before you answer — the VTun half you raised on 29 August looks like the one thing here still unresolved:
S98vtun//usr/sbin/tunnelfilling the whole 64 KB syslog RAM buffer withsyntax error line 5/No hosts definedwhen it starts with DNS down.That matters a little more now than it did then. The WireGuard failures #2401 just started logging land in that same 64 KB buffer (
syslogd -C64, fromS01syslogd), so a VTun flood evicts exactly the diagnostics we just added — and on a board running both, it's the same power-cut-then-no-DNS boot that sets each of them off.Would you rather close this and open a separate issue for VTun, or keep this one as the VTun tracker? Either is fine, I'd just like it pointing at one thing.
On the VTun half — #2406 is up as a draft.
One thing worth flagging, because it changes what needs testing: the two messages you pasted are not a resolver failure.
config()writes the session name on line 5 of the file it generates, and that name is the MAC of the default-route interface. On a boot that reachestunnelbefore the network is up,ip rmatches nothing,identity_srcis empty,cat /sys/class/net//addressfails, and line 5 comes out as a bare{:4 } 5 { 6 password ;which is exactly
syntax error line 5, and thenNo hosts definedbecause no session parsed. The reason it drowned the ring is the loop underneath it —while true; do identity; interface; config; vtund ...; donewith no delay anywhere. In a stubbed run that launched vtund 17 times in two seconds, two syslog lines each.So the draft waits for the identity as well as for the name to resolve — your DNS point is real too, vtund resolves the server itself and exits when it cannot — and puts a floor under the restart rate for whatever fails after that. The literal-address test is the one
/usr/sbin/wireguardalready uses, so a numeric server never waits on a resolver it does not need.I have no camera running vtund, so the draft carries no hardware evidence and I have not ticked anything I could not check. You have the repro and the fleet — would you give it a run? If it holds up on your boards I will mark it ready.
The WireGuard startup issue is now resolved, so I think we can close this issue. The remaining VTun startup problem is a separate issue and can be continued in #2406. Thanks!
- added a commit that references this issue
on Sep 18, 2026
I have identified the cause of an intermittent WireGuard startup failure. As I suspected, the problem occurs when the server address is specified as a hostname rather than an IP address, and DNS is unavailable during reboot, either because the network is not ready yet or because the DNS server itself is unreachable.
The error occurs after approximately 2.5 minutes in
/usr/sbin/wireguard, at this line:How to reproduce
This can be reproduced in real time on a camera connected to a local network:
wireguardon the camera and observe the output:Possible solutions
The following are alternatives to the workaround proposed in #2242 (which suggested periodically monitoring the WireGuard connection and restarting it when it is detected as inactive).
Change WireGuard so that reading the configuration does not fail when the endpoint hostname cannot currently be resolved. The configuration should be accepted, and WireGuard should retry connecting later, similarly to what happens when an already established connection to the server is lost.
Start
/usr/sbin/wireguardasynchronously from/etc/init.d/S98wireguard, and in/usr/sbin/wireguard, after all U-Boot variables have been read, wait for the DNS server to become available if the WireGuard server address is specified as a hostname:This approach avoids blocking the rest of the system startup while still allowing WireGuard to start automatically as soon as DNS becomes available.
P. S. I also believe that errors from initialization scripts should be logged to syslog, rather than being printed only to the console.