Fluxtail
Log Management Guides

How to Fix DNS Issues: A Step-by-Step Diagnosis

Find the cause of DNS failures with dig checks. Distinguish NXDOMAIN, no data, SERVFAIL, timeouts, caches, delegation, and application resolver differences.

By Fluxtail Engineering Updated

To fix a DNS issue, first identify the exact hostname, record type, client, and resolver that produced the failure. Query that same name and type from the affected network, read the DNS response code, and compare the recursive resolver's answer with the authoritative source only if needed. Do not start by flushing every cache or changing to a public DNS provider: those actions can hide the evidence and break private or split-horizon names.

For a person seeing “server not found,” check the spelling and full hostname, then try the same site from another device or network once. If only one browser fails, its DNS settings may differ from the operating system's. If the problem persists, give the site or network operator the hostname, time, network, and exact error text—without sharing cookies or private URLs. The rest of this guide shows how to narrow the fault when you operate the affected client or DNS service.

Ask the exact DNS question that failed

DNS answers are specific to a query name and record type. A asks for an IPv4 address, AAAA for an IPv6 address, MX for mail routing, and so on. An A answer does not prove the AAAA path works. A trailing dot makes the name explicitly absolute, avoiding accidental search-suffix expansion in a diagnostic command.

These read-only examples use the reserved example.com domain. Replace it with a hostname you are authorized to test:

dig example.com. A
dig example.com. AAAA

Keep the normal, full output. It shows the question, status, answer, TTL, and the server that responded; +short hides details needed to distinguish an empty answer from a failure. BIND's current dig manual documents the name/type arguments and says that, unless an explicit server is supplied, dig consults /etc/resolv.conf on the system where it runs. A local stub may appear as the responding server even though it forwards to another resolver. If your dig output differs from the expected defaults, also check for a user .digrc file or specify the relevant query options explicitly.

Record the source host or container, time with timezone, exact FQDN, type, resolver address shown by dig, response code, answer and TTL. Repeat from one known-good environment using the same FQDN and type. A comparison with a different name, record type, VPN state, or resolver view is not evidence of inconsistency.

Read the response code before editing records

The response categories point to different next checks:

  • NOERROR with an answer: The queried resolver supplied a record. Check whether it is the expected value for this client, whether a CNAME chain ends at the intended destination, and whether the application actually used this resolver. A valid DNS answer does not prove the destination accepts TCP connections or serves the right application.
  • NXDOMAIN: The name was reported not to exist. Check for a typo, an unintended search suffix, a missing label, or a wrong delegation. Follow the CNAME chain if present: a negative result can refer to the target of an alias, not the alias text you first queried. RFC 2308 defines this negative response and its caching behavior.
  • NOERROR with no record of the requested type: This can be a valid NODATA response: the name exists, but this type is unavailable. For example, AAAA may be absent while A exists. An empty answer alone is not enough to classify it; inspect the authority section and alias chain. RFC 2308 explains the distinction from a referral.
  • SERVFAIL: The server could not complete the query. It does not mean the name is absent. Investigate the resolver's upstream path, authoritative availability, delegation, or DNSSEC validation as appropriate. RFC 1035 defines the response code; RFC 9520 describes several failure contexts.
  • Timeout or no response: The query did not get a usable reply from that path. Check reachability to the selected resolver, UDP/TCP 53 policy as applicable, resolver health, and whether an upstream authoritative server responds. A firewall, routing fault, or overloaded server can resemble a DNS record problem. A timeout is not NXDOMAIN.
  • REFUSED: The queried server declined to answer, often because of policy or role. Confirm that it is intended to serve this client or zone before changing an access list. It is not equivalent to a missing record.

The DNS response-code definitions are a better guide than a generic “DNS failed” message. Avoid treating NOERROR as proof that the desired application can connect, or SERVFAIL as permission to create a new record.

Compare the application's lookup path with dig

dig asks a DNS server; many applications use more than that path. On Linux, the Name Service Switch can consult /etc/hosts and other sources before DNS. Resolver search domains can turn a short name into another FQDN. A container or Kubernetes pod can have different DNS settings from its host. A browser may use a secure-DNS resolver, and some runtimes cache answers or use their own resolver behavior. A dig result from your laptop therefore does not prove what the failing process asked or received.

Reproduce from the affected network namespace or client environment using the name as the application actually supplied it. Inspect its configured resolver, search domains, hosts file, VPN or private-DNS policy, and any proxy setting that could resolve the destination elsewhere. Compare a fully qualified name ending in . with the short name only when the application uses a short name. Do not publish private hostnames by sending them to an unrelated public resolver as a convenience check.

If the application receives an address but its connection still fails, shift to the next layer: port reachability, TLS, proxy routing, or application response. A DNS lookup cannot establish that HTTPS will succeed. Likewise, an ICMP ping failure alone does not prove DNS is broken, because ICMP may be blocked while DNS and HTTPS work.

Compare recursive and authoritative answers

A recursive resolver follows the DNS hierarchy on the client's behalf and may return a cached answer. Authoritative name servers hold the zone's data for their delegated portion of the namespace. RFC 1034 explains those roles and the distributed, cached design. When the recursive result looks wrong, query the relevant authoritative server directly—but first verify that it is actually authoritative for the affected zone, including the parent delegation.

The following command format is from the dig manual. 192.0.2.53 is a documentation-only address: replace it with the approved resolver or authoritative server IP you identified. +norecurse asks that server without requesting recursion; it does not turn an arbitrary server into an authoritative one.

dig @192.0.2.53 example.com. A +norecurse

Check the answer and authority flags, not just the IP. A referral points to another zone's name servers. If the parent delegates to one set of servers but those servers lack the zone or disagree on its records, repair the exact delegation or zone publication after verifying ownership. A direct authoritative response can be current while a recursive resolver still has a cached older answer. Conversely, a recursive response can be intentionally different because the client is inside a private DNS view. Do not overwrite private records to match the public internet.

For an external name, look at the parent-to-child delegation, CNAME chain, and the authoritative responses from each advertised name server. For an internal name, identify the intended private authority and forwarding rules; a public resolver is usually the wrong control. Avoid a broad zone transfer or full resolver configuration dump just to check one record.

Account for caching before calling a change ineffective

DNS TTLs allow answers to remain in recursive and local caches. A TTL in a dig response from a recursive resolver is the remaining lifetime of that cached record, not the original TTL necessarily. Positive answers are cached according to their record TTL. Negative NXDOMAIN and NODATA answers have their own TTL derived from the zone's SOA information under RFC 2308. Newer RFC 9520 also addresses caching of resolution failures such as SERVFAIL and timeouts. Some resolvers may serve stale data during upstream trouble. There is no universal “propagation takes exactly N minutes” rule.

If a record was just changed, compare the authority's answer with the affected recursive resolver and note the TTL or negative-cache evidence. Wait for the relevant cache lifetime when safe; if urgent, use a narrowly scoped cache action only after identifying the cache and its role. Flushing every workstation, restarting resolvers, or lowering a new TTL after old answers are already cached will not reliably erase earlier cached state. Lower TTL ahead of a planned change only when the operational tradeoff has been assessed.

Investigate DNSSEC and split views without bypassing them

A validating resolver can return SERVFAIL when signed data fails validation, even if a direct query to an authoritative server returns records. The DNSSEC introduction distinguishes secure, insecure, bogus, and indeterminate validation states. Check the resolver's validation logs and the domain's delegation, DS, DNSKEY, and signature state. Do not turn off DNSSEC validation or bypass it in production to make an answer appear; that hides a trust failure. An unsigned zone is not automatically broken—interpret it against the delegation chain.

Split-horizon DNS can intentionally serve different answers inside a VPN, VPC, or enterprise network than outside it. Compare sources within the same view before declaring a mismatch. When a private name works only through the VPN, first inspect that VPN's resolver routing and search policy. Replacing it with a public resolver can leak the queried name and remove the private answer entirely.

Verify the exact fix and keep useful evidence

Make the smallest approved change at the responsible layer: correct the application's hostname or resolver configuration, repair a wrong zone record, restore an authoritative server, fix a delegation, or resolve a DNSSEC signing mismatch. DNS record, forwarding, resolver, and cache changes affect other callers and require change control. Afterward, retest the original FQDN and type from the affected client, compare with the authoritative answer, and verify the intended application connection—not merely a successful DNS lookup. Check one neighboring name that should remain unchanged.

Useful, privacy-safe DNS evidence includes time, queried name or a suitable redacted class, type, resolver identity, response code, answer/TTL where permitted, and the client service or environment. Query names can reveal internal systems or user activity, so restrict access and retention. Application logs may show a lookup error but not the authoritative cause; resolver logs may show the DNS result but not whether the application connected afterward. The log-reading guide helps keep those boundaries clear.

Fluxtail is a paid Starter/Pro, logs-focused destination for resolver or application events already delivered through a supported collector or receiver. Search, filters, and Live Tail can help inspect those records if the source emits and maps the relevant fields. It does not query authoritative DNS, validate DNSSEC, clear caches, or infer a failure from a missing log row. For planned DNS traffic steering rather than troubleshooting a lookup, see DNS load balancing.