[PR #6518] [client] Skip re-resolving cached management cache domains #26009

Closed
opened 2026-08-05 07:06:45 -04:00 by saavagebueno · 0 comments
Owner

Original Pull Request: https://github.com/netbirdio/netbird/pull/6518

State: closed
Merged: Yes


Describe your changes

The management DNS cache re-resolved every infrastructure domain (signal, relay, STUN, TURN) on each config-carrying sync, regardless of whether it was already cached. Those resolves run under the engine sync lock, so when DNS is degraded each one waits out its timeout and starves signal/peer processing. This skips domains already in the cache and resolves the rest off the serial path.

  • Skip re-resolving infrastructure domains that are already cached; staleness is handled by the existing stale-while-revalidate refresh path
  • Resolve newly added domains concurrently instead of one after another
  • Back off a domain that fails its initial resolve so it isn't retried on every sync until the backoff elapses
  • Dedupe domains that appear under more than one type (e.g. STUN and TURN on the same host)

Stack

Checklist

  • Is it a bug fix
  • Is a typo/documentation fix
  • Is a feature enhancement
  • It is a refactor
  • Created tests that fail without the change (if possible)
  • This change does not modify the public API, gRPC protocols, functionality behavior, CLI / service flags, or introduce a new feature — OR I have discussed it with the NetBird team beforehand (link the issue / Slack thread in the description). See CONTRIBUTING.md.

By submitting this pull request, you confirm that you have read and agree to the terms of the Contributor License Agreement.

Documentation

Select exactly one:

  • I added/updated documentation for this change
  • Documentation is not needed for this change (explain why)

Internal client behavior change to the management DNS cache; no user-facing surface.

Docs PR URL (required if "docs added" is checked)

Paste the PR link from https://github.com/netbirdio/docs here:

https://github.com/netbirdio/docs/pull/__

Summary by CodeRabbit

  • Bug Fixes

    • Improved DNS resolution by tracking failed management-domain lookups and applying backoff to avoid immediate retries.
    • Handle partial DNS family failures by keeping successful records cached while retrying only the missing family.
    • Treat NODATA as non-failure to prevent unnecessary re-resolution.
    • Strengthened deduplication to resolve each needed domain only once per update.
  • Tests

    • Added/expanded resolver test coverage for caching, concurrency, backoff, pruning failure markers, duplicate suppression, partial failures, and NODATA handling.
**Original Pull Request:** https://github.com/netbirdio/netbird/pull/6518 **State:** closed **Merged:** Yes --- ## Describe your changes The management DNS cache re-resolved every infrastructure domain (signal, relay, STUN, TURN) on each config-carrying sync, regardless of whether it was already cached. Those resolves run under the engine sync lock, so when DNS is degraded each one waits out its timeout and starves signal/peer processing. This skips domains already in the cache and resolves the rest off the serial path. - Skip re-resolving infrastructure domains that are already cached; staleness is handled by the existing stale-while-revalidate refresh path - Resolve newly added domains concurrently instead of one after another - Back off a domain that fails its initial resolve so it isn't retried on every sync until the backoff elapses - Dedupe domains that appear under more than one type (e.g. STUN and TURN on the same host) ## Issue ticket number and link ## Stack <!-- branch-stack --> ### Checklist - [x] Is it a bug fix - [ ] Is a typo/documentation fix - [ ] Is a feature enhancement - [ ] It is a refactor - [x] Created tests that fail without the change (if possible) - [ ] This change does **not** modify the public API, gRPC protocols, functionality behavior, CLI / service flags, or introduce a new feature — **OR** I have discussed it with the NetBird team beforehand (link the issue / Slack thread in the description). See [CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#discuss-changes-with-the-netbird-team-first). > By submitting this pull request, you confirm that you have read and agree to the terms of the [Contributor License Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md). ## Documentation Select exactly one: - [ ] I added/updated documentation for this change - [x] Documentation is **not needed** for this change (explain why) Internal client behavior change to the management DNS cache; no user-facing surface. ### Docs PR URL (required if "docs added" is checked) Paste the PR link from https://github.com/netbirdio/docs here: https://github.com/netbirdio/docs/pull/__ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved DNS resolution by tracking failed management-domain lookups and applying backoff to avoid immediate retries. * Handle partial DNS family failures by keeping successful records cached while retrying only the missing family. * Treat NODATA as non-failure to prevent unnecessary re-resolution. * Strengthened deduplication to resolve each needed domain only once per update. * **Tests** * Added/expanded resolver test coverage for caching, concurrency, backoff, pruning failure markers, duplicate suppression, partial failures, and NODATA handling. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
saavagebueno added the pull-request label 2026-08-05 07:06:45 -04:00
Sign in to join this conversation.
No Label pull-request
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: DYNR/netbird#26009