Skip to content

Proposal: Make WireGuard startup resilient to temporary DNS unavailability #2319

Description

@usa-

I have identified the cause of an intermittent WireGuard startup failure. As I suspected, the problem occurs when the server address is specified as a hostname rather than an IP address, and DNS is unavailable during reboot, either because the network is not ready yet or because the DNS server itself is unreachable.

The error occurs after approximately 2.5 minutes in /usr/sbin/wireguard, at this line:

wg setconf wg0 /tmp/wireguard.conf || { echo "Error: Failed to apply wireguard configuration." >&2; exit 1; }

How to reproduce

This can be reproduced in real time on a camera connected to a local network:

  1. Disconnect the local network from the Internet (more precisely, block communication between the camera and the DNS server).
  2. Configure WireGuard to use a hostname as the server address.
  3. On the camera, execute:
chmod -x /etc/init.d/S98wireguard
  1. Reboot the camera.
  2. Run wireguard on the camera and observe the output:
root@gk7205v300-imx335:~# wireguard
Try again: `${Endpoint}'. Trying again in 1.00 seconds...
Try again: `${Endpoint}'. Trying again in 1.20 seconds...
Try again: `${Endpoint}'. Trying again in 1.44 seconds...
Try again: `${Endpoint}'. Trying again in 1.73 seconds...
Try again: `${Endpoint}'. Trying again in 2.07 seconds...
Try again: `${Endpoint}'. Trying again in 2.49 seconds...
Try again: `${Endpoint}'. Trying again in 2.99 seconds...
Try again: `${Endpoint}'. Trying again in 3.58 seconds...
Try again: `${Endpoint}'. Trying again in 4.30 seconds...
Try again: `${Endpoint}'. Trying again in 5.16 seconds...
Try again: `${Endpoint}'. Trying again in 6.19 seconds...
Try again: `${Endpoint}'. Trying again in 7.43 seconds...
Try again: `${Endpoint}'. Trying again in 8.92 seconds...
Try again: `${Endpoint}'. Trying again in 10.70 seconds...
Try again: `${Endpoint}'. Trying again in 12.84 seconds...
Try again: `${Endpoint}'
Configuration parsing error
Error: Failed to apply wireguard configuration.
  1. Reconnect the camera to the Internet and restore the normal startup:
chmod +x /etc/init.d/S98wireguard

Possible solutions

The following are alternatives to the workaround proposed in #2242 (which suggested periodically monitoring the WireGuard connection and restarting it when it is detected as inactive).

  1. Change WireGuard so that reading the configuration does not fail when the endpoint hostname cannot currently be resolved. The configuration should be accepted, and WireGuard should retry connecting later, similarly to what happens when an already established connection to the server is lost.

  2. Start /usr/sbin/wireguard asynchronously from /etc/init.d/S98wireguard, and in /usr/sbin/wireguard, after all U-Boot variables have been read, wait for the DNS server to become available if the WireGuard server address is specified as a hostname:

ENDPOINT=$(fw_printenv -n wg_endpoint)

case "$ENDPOINT" in
    [0-9]*)
        ;;
    *)
        HOST="${ENDPOINT%:*}"

        while ! nslookup "$HOST" >/dev/null 2>&1; do
            sleep 10
        done
        ;;
esac

This approach avoids blocking the rest of the system startup while still allowing WireGuard to start automatically as soon as DNS becomes available.

P. S. I also believe that errors from initialization scripts should be logged to syslog, rather than being printed only to the console.

Activity

  1. usa- commented on Aug 27, 2026

    @usa-
    ContributorAuthor

    I have already applied the second solution to all my cameras and will monitor them for a week.

    Power outages occur at the camera locations with an unknown frequency, but every 1–2 days I find one or two cameras with WireGuard in a non-working state. The cameras sometimes boot faster than the network connection to the DNS server becomes available after a power outage, which previously left them without WireGuard.

    If the fix proves reliable, this will support implementing this approach in the official firmware scripts.

  2. usa- commented on Aug 29, 2026

    @usa-
    ContributorAuthor

    Two days — everything has been working perfectly.
    Just in case:

    /etc/init.d/S98wireguard
    #!/bin/sh
    case "$1" in
    	start)
    		wgc=$(fw_printenv -n wg_privkey)
    		if [ -n "$wgc" ]; then
    			wireguard &
    		fi
    		;;
    	stop)
    		;;
    	*)
    		echo "Usage: $0 {start}"
    		exit 1
    		;;
    esac
    /usr/sbin/wireguard
    #!/bin/sh
    modprobe wireguard || { echo "Error: Failed to load wireguard module." >&2; exit 1; }
    ip link add dev wg0 type wireguard || { echo "Error: Failed to create wg0 interface." >&2; exit 1; }
    WG_PRIVKEY="$(fw_printenv -n wg_privkey)"
    WG_PRESHARED_KEY="$(fw_printenv -n wg_sharkey)"
    ENDPOINT=$(fw_printenv -n wg_endpoint)
    
    case "$ENDPOINT" in
        [0-9]*)
            ;;
        *)
            HOST="${ENDPOINT%:*}"
    
            while ! nslookup "$HOST" >/dev/null 2>&1; do
                sleep 10
            done
            ;;
    esac
    
    ( echo "#"
      echo "[Interface]"
      echo "PrivateKey = $WG_PRIVKEY"
      echo
      echo "[Peer]"
      echo "Endpoint = $ENDPOINT"
      echo "PersistentKeepalive = $(fw_printenv -n wg_alive)"
      echo "PublicKey = $(fw_printenv -n wg_pubkey)"
      [ -n "$WG_PRESHARED_KEY" ] && echo "PresharedKey = $WG_PRESHARED_KEY"
      echo "AllowedIPs = $(fw_printenv -n wg_allowed)"
      echo "#"
    ) >>/tmp/wireguard.conf
    wg setconf wg0 /tmp/wireguard.conf || { echo "Error: Failed to apply wireguard configuration." >&2; exit 1; }
    wg_address="$(fw_printenv -n wg_address)"
    if [ -z "$wg_address" ]; then
        echo "Error: wg_address environment variable is not set or empty." >&2
        exit 1
    fi
    ip address add dev wg0 "$wg_address"
    ip link set up dev wg0
    for i in $(fw_printenv -n wg_allowed | tr ',' ' '); do
        ip -4 route add "$i" dev wg0
    done
  3. usa- commented on Aug 29, 2026

    @usa-
    ContributorAuthor

    As I suggested in the issue description, errors from initialization scripts should be logged to syslog rather than printed only to the console. I will continue testing a script implementing this approach:

    #!/bin/sh
    
    PROGRAM="${0##*/}"
    
    run_cmd() {
        msg="$1"
        shift
    
        output=$("$@" 2>&1)
        rc=$?
    
        if [ "$rc" -ne 0 ]; then
            [ -n "$msg" ] && msg="$msg. "
            logger -p user.err -t "$PROGRAM[$$]" "Error: ${msg}"
            logger -p user.err -t "$PROGRAM[$$]" "Command \`$*\` finished with code $rc:"
    
            printf '%s\n' "$output" |
            while IFS= read -r line; do
                logger -p user.err -t "$PROGRAM[$$]" "$line"
            done
    
            kill -s TERM $$ 2>/dev/null
            exit "$rc"
        fi
    
        printf '%s\n' "$output"
    }
    
    run_cmd "Failed to load wireguard module" modprobe wireguard
    run_cmd "Failed to create wg0 interface" ip link add dev wg0 type wireguard
    
    WG_PRIVKEY=$(run_cmd "" fw_printenv -n wg_privkey)
    WG_PRESHARED_KEY=$(fw_printenv -n wg_sharkey 2>/dev/null || true)
    ENDPOINT=$(run_cmd "" fw_printenv -n wg_endpoint)
    
    case "$ENDPOINT" in
        [0-9]*)
            ;;
        *)
            HOST="${ENDPOINT%:*}"
    
            while ! nslookup "$HOST" >/dev/null 2>&1; do
                sleep 10
            done
            ;;
    esac
    
    (
        echo "#"
        echo "[Interface]"
        echo "PrivateKey = $WG_PRIVKEY"
        echo
        echo "[Peer]"
        echo "Endpoint = $ENDPOINT"
        echo "PersistentKeepalive = $(run_cmd "" fw_printenv -n wg_alive)"
        echo "PublicKey = $(run_cmd "" fw_printenv -n wg_pubkey)"
        [ -n "$WG_PRESHARED_KEY" ] && echo "PresharedKey = $WG_PRESHARED_KEY"
        echo "AllowedIPs = $(run_cmd "" fw_printenv -n wg_allowed)"
        echo "#"
    ) >>/tmp/wireguard.conf
    
    run_cmd "Failed to apply wireguard configuration" wg setconf wg0 /tmp/wireguard.conf
    
    wg_address=$(run_cmd "" fw_printenv -n wg_address)
    if [ -z "$wg_address" ]; then
        echo "Error: wg_address environment variable is not set or empty." >&2
        exit 1
    fi
    
    run_cmd "" ip address add dev wg0 "$wg_address"
    run_cmd "" ip link set up dev wg0
    
    for i in $(run_cmd "" fw_printenv -n wg_allowed | tr ',' ' '); do
        run_cmd "" ip -4 route add "$i" dev wg0
    done

    If this approach were used in the current /usr/sbin/wireguard script in the official firmware, the following would appear in syslog if WireGuard starts while the DNS server is unavailable:

    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Error: Failed to apply wireguard configuration.
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Command `wg setconf wg0 /tmp/wireguard.conf` finished with code 1:
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.00 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.20 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.44 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 1.73 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 2.07 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 2.49 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 2.99 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 3.58 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 4.30 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 5.16 seconds...
    Aug 29 09:05:30 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 6.19 seconds...
    Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 7.43 seconds...
    Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 8.92 seconds...
    Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 10.70 seconds...
    Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'. Trying again in 12.84 seconds...
    Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Try again: `{hidden}:51820'
    Aug 29 09:05:31 gk7205v300-imx335 user.err wireguard[770]: Configuration parsing error
    
  4. usa- commented on Aug 29, 2026

    @usa-
    ContributorAuthor

    Since VTun, if it starts while DNS is unavailable, fills all 64 KB of the syslog RAM buffer with these two messages:

    Aug 29 09:05:21 gk7205v300-imx335 daemon.err vtund[21576]: syntax error line 5
    Aug 29 09:05:21 gk7205v300-imx335 daemon.err vtund[21576]: No hosts defined
    

    I suggest applying the same approach to /etc/init.d/S98vtun and /usr/sbin/tunnel: wait for DNS to become available before starting VTun.

  5. usa- commented on Sep 1, 2026

    @usa-
    ContributorAuthor

    @widgetii, @flyrouter

    I think the testing is sufficient at this point. I haven't run into a single problem during testing.

    I suggest that you develop and test the fix for VTun, so that it doesn't flood the syslog buffer, on your side.

  6. usa- commented on Sep 6, 2026

    @usa-
    ContributorAuthor

    @widgetii

    Just wondering if this is still relevant.

  7. widgetii commented on Sep 6, 2026

    @widgetii
    Member

    PR are always welcome. I personally don't use either Wireguard, nor vtun and don't need them and am happy to remove them from the firmware.

  8. usa- commented on Sep 6, 2026

    @usa-
    ContributorAuthor

    Created: #2377

  9. openipc-ai commented on Sep 7, 2026

    @openipc-ai
    Collaborator

    @usa- #2377 is merged and on master as bf7567a; it reaches every image from the next nightly. Does that cover everything this issue asked for, or is something still open here? If it is all in, close it whenever you are happy, or say so and I will.

  10. usa- commented on Sep 7, 2026

    @usa-
    ContributorAuthor

    I am planning to raise one more PR based on this issue during next few days.

  11. openipc-ai commented on Sep 13, 2026

    @openipc-ai
    Collaborator

    @usa- #2401 is merged — 3769227 on master, in every image from the next nightly. That's the follow-up you mentioned on 7 September, so both proposals from the issue body are now in: the DNS wait from #2377, and initialisation errors going to syslog rather than only the console.

    Does that close this out for you? Before you answer — the VTun half you raised on 29 August looks like the one thing here still unresolved: S98vtun / /usr/sbin/tunnel filling the whole 64 KB syslog RAM buffer with syntax error line 5 / No hosts defined when it starts with DNS down.

    That matters a little more now than it did then. The WireGuard failures #2401 just started logging land in that same 64 KB buffer (syslogd -C64, from S01syslogd), so a VTun flood evicts exactly the diagnostics we just added — and on a board running both, it's the same power-cut-then-no-DNS boot that sets each of them off.

    Would you rather close this and open a separate issue for VTun, or keep this one as the VTun tracker? Either is fine, I'd just like it pointing at one thing.

  12. openipc-ai commented on Sep 13, 2026

    @openipc-ai
    Collaborator

    On the VTun half — #2406 is up as a draft.

    One thing worth flagging, because it changes what needs testing: the two messages you pasted are not a resolver failure. config() writes the session name on line 5 of the file it generates, and that name is the MAC of the default-route interface. On a boot that reaches tunnel before the network is up, ip r matches nothing, identity_src is empty, cat /sys/class/net//address fails, and line 5 comes out as a bare {:

         4	}
         5	 {
         6		password ;
    

    which is exactly syntax error line 5, and then No hosts defined because no session parsed. The reason it drowned the ring is the loop underneath it — while true; do identity; interface; config; vtund ...; done with no delay anywhere. In a stubbed run that launched vtund 17 times in two seconds, two syslog lines each.

    So the draft waits for the identity as well as for the name to resolve — your DNS point is real too, vtund resolves the server itself and exits when it cannot — and puts a floor under the restart rate for whatever fails after that. The literal-address test is the one /usr/sbin/wireguard already uses, so a numeric server never waits on a resolver it does not need.

    I have no camera running vtund, so the draft carries no hardware evidence and I have not ticked anything I could not check. You have the repro and the fleet — would you give it a run? If it holds up on your boards I will mark it ready.

  13. usa- commented on Sep 13, 2026

    @usa-
    ContributorAuthor

    The WireGuard startup issue is now resolved, so I think we can close this issue. The remaining VTun startup problem is a separate issue and can be continued in #2406. Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions