[PR #6518] [MERGED] [client] Skip re-resolving cached management cache domains #29614

Closed
opened 2026-08-05 08:08:29 -04:00 by saavagebueno · 0 comments
Owner

📋 Pull Request Information

Original PR: https://github.com/netbirdio/netbird/pull/6518
Author: @lixmal
Created: 6/23/2026
Status: Merged
Merged: 6/23/2026
Merged by: @lixmal

Base: mainHead: fix/mgmt-cache-skip-resolved


📝 Commits (2)

  • b8d0a04 Skip re-resolving cached management cache domains and resolve new ones in parallel
  • b8cce80 Retry mgmt cache domains whose initial resolve only partially succeeded

📊 Changes

3 files changed (+292 additions, -10 deletions)

View changed files

📝 client/internal/dns/mgmt/mgmt.go (+98 -10)
📝 client/internal/dns/mgmt/mgmt_refresh_test.go (+11 -0)
client/internal/dns/mgmt/mgmt_resolve_test.go (+183 -0)

📄 Description

Describe your changes

The management DNS cache re-resolved every infrastructure domain (signal, relay, STUN, TURN) on each config-carrying sync, regardless of whether it was already cached. Those resolves run under the engine sync lock, so when DNS is degraded each one waits out its timeout and starves signal/peer processing. This skips domains already in the cache and resolves the rest off the serial path.

  • Skip re-resolving infrastructure domains that are already cached; staleness is handled by the existing stale-while-revalidate refresh path
  • Resolve newly added domains concurrently instead of one after another
  • Back off a domain that fails its initial resolve so it isn't retried on every sync until the backoff elapses
  • Dedupe domains that appear under more than one type (e.g. STUN and TURN on the same host)

Stack

Checklist

  • Is it a bug fix
  • Is a typo/documentation fix
  • Is a feature enhancement
  • It is a refactor
  • Created tests that fail without the change (if possible)
  • This change does not modify the public API, gRPC protocols, functionality behavior, CLI / service flags, or introduce a new feature — OR I have discussed it with the NetBird team beforehand (link the issue / Slack thread in the description). See CONTRIBUTING.md.

By submitting this pull request, you confirm that you have read and agree to the terms of the Contributor License Agreement.

Documentation

Select exactly one:

  • I added/updated documentation for this change
  • Documentation is not needed for this change (explain why)

Internal client behavior change to the management DNS cache; no user-facing surface.

Docs PR URL (required if "docs added" is checked)

Paste the PR link from https://github.com/netbirdio/docs here:

https://github.com/netbirdio/docs/pull/__

Summary by CodeRabbit

  • Bug Fixes

    • Improved DNS resolution by tracking failed management-domain lookups and applying backoff to avoid immediate retries.
    • Handle partial DNS family failures by keeping successful records cached while retrying only the missing family.
    • Treat NODATA as non-failure to prevent unnecessary re-resolution.
    • Strengthened deduplication to resolve each needed domain only once per update.
  • Tests

    • Added/expanded resolver test coverage for caching, concurrency, backoff, pruning failure markers, duplicate suppression, partial failures, and NODATA handling.

🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.

## 📋 Pull Request Information **Original PR:** https://github.com/netbirdio/netbird/pull/6518 **Author:** [@lixmal](https://github.com/lixmal) **Created:** 6/23/2026 **Status:** ✅ Merged **Merged:** 6/23/2026 **Merged by:** [@lixmal](https://github.com/lixmal) **Base:** `main` ← **Head:** `fix/mgmt-cache-skip-resolved` --- ### 📝 Commits (2) - [`b8d0a04`](https://github.com/netbirdio/netbird/commit/b8d0a04fa169386a54e97abed2962d6e592f3e05) Skip re-resolving cached management cache domains and resolve new ones in parallel - [`b8cce80`](https://github.com/netbirdio/netbird/commit/b8cce80e242b8f4164dad3fb8bfe75c09c868966) Retry mgmt cache domains whose initial resolve only partially succeeded ### 📊 Changes **3 files changed** (+292 additions, -10 deletions) <details> <summary>View changed files</summary> 📝 `client/internal/dns/mgmt/mgmt.go` (+98 -10) 📝 `client/internal/dns/mgmt/mgmt_refresh_test.go` (+11 -0) ➕ `client/internal/dns/mgmt/mgmt_resolve_test.go` (+183 -0) </details> ### 📄 Description ## Describe your changes The management DNS cache re-resolved every infrastructure domain (signal, relay, STUN, TURN) on each config-carrying sync, regardless of whether it was already cached. Those resolves run under the engine sync lock, so when DNS is degraded each one waits out its timeout and starves signal/peer processing. This skips domains already in the cache and resolves the rest off the serial path. - Skip re-resolving infrastructure domains that are already cached; staleness is handled by the existing stale-while-revalidate refresh path - Resolve newly added domains concurrently instead of one after another - Back off a domain that fails its initial resolve so it isn't retried on every sync until the backoff elapses - Dedupe domains that appear under more than one type (e.g. STUN and TURN on the same host) ## Issue ticket number and link ## Stack <!-- branch-stack --> ### Checklist - [x] Is it a bug fix - [ ] Is a typo/documentation fix - [ ] Is a feature enhancement - [ ] It is a refactor - [x] Created tests that fail without the change (if possible) - [ ] This change does **not** modify the public API, gRPC protocols, functionality behavior, CLI / service flags, or introduce a new feature — **OR** I have discussed it with the NetBird team beforehand (link the issue / Slack thread in the description). See [CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#discuss-changes-with-the-netbird-team-first). > By submitting this pull request, you confirm that you have read and agree to the terms of the [Contributor License Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md). ## Documentation Select exactly one: - [ ] I added/updated documentation for this change - [x] Documentation is **not needed** for this change (explain why) Internal client behavior change to the management DNS cache; no user-facing surface. ### Docs PR URL (required if "docs added" is checked) Paste the PR link from https://github.com/netbirdio/docs here: https://github.com/netbirdio/docs/pull/__ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved DNS resolution by tracking failed management-domain lookups and applying backoff to avoid immediate retries. * Handle partial DNS family failures by keeping successful records cached while retrying only the missing family. * Treat NODATA as non-failure to prevent unnecessary re-resolution. * Strengthened deduplication to resolve each needed domain only once per update. * **Tests** * Added/expanded resolver test coverage for caching, concurrency, backoff, pruning failure markers, duplicate suppression, partial failures, and NODATA handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --- <sub>🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.</sub>
saavagebueno added the pull-request label 2026-08-05 08:08:29 -04:00
Sign in to join this conversation.
No Label pull-request
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: DYNR/netbird#29614