Fluxtail
Log Management Guides

How to Read Logs: A Practical Incident Guide

Learn how to read logs during incidents: narrow the time window, inspect fields, correlate events, use safe CLI commands, and verify outcomes.

By Fluxtail Engineering Updated

Reading logs means reconstructing a bounded sequence of events and testing what that sequence proves. Start with one symptom, one affected service or host, and a precise time window. Then inspect the event fields, confirm that the collection path is complete enough for the question, follow stable identifiers, and verify the real outcome. Searching for ERROR can help you orient yourself, but it is not a root-cause method.

This guide shows how to read plain-text, JSON, systemd, Docker, and Kubernetes logs without changing the system. The commands are intentionally narrow and read-only. Access to logs can still expose sensitive data, so use an authorized account and avoid copying more output than the investigation requires.

Start with an investigation boundary

Before opening a file or running a query, write down four facts:

  • the observed symptom, in concrete terms;
  • the service, workload, container, or host that could have produced relevant evidence;
  • the earliest and latest useful timestamps, including the timezone;
  • one stable identifier, if available, such as a request ID, job run ID, deployment ID, or trace ID.

For example: “Checkout returned HTTP 502 for request req_example_7f2c between 14:25 and 14:40 UTC on the production API.” That boundary is more useful than “look for checkout errors.” It limits unrelated traffic, gives collaborators the same clock, and creates a testable question.

Preserve the original symptom before interpreting it. Record the status code, user-visible message, affected operation, and source of the report. Do not paste tokens, cookies, full request bodies, or personal data into a ticket or command line. If the only timestamp is in local time, record its offset and convert it explicitly before comparing it with UTC logs.

Use a wider window than the visible failure, but keep it bounded. The triggering event may occur before the final exception, while retries and cleanup may continue after it. If the search is empty, widen the window deliberately rather than removing every filter at once.

Read one event field by field

A log record is evidence only after you understand what each field represents. The OpenTelemetry Logs Data Model provides a useful vendor-neutral vocabulary: event time, observed time, severity, body, resource, attributes, and optional trace or span context. An application does not have to emit OpenTelemetry records to benefit from the same distinctions.

Consider this illustrative JSON Lines event:

{"timestamp":"2026-09-16T14:32:08.481Z","observed_timestamp":"2026-09-16T14:32:09.016Z","severity":"ERROR","event_name":"checkout.payment_gateway_timeout","service":"checkout-api","environment":"production","version":"7.4.2","deployment_id":"deploy_example_3","request_id":"req_example_7f2c","message":"Payment authorization timed out","error_type":"gateway_timeout"}

Read it in this order:

  1. Event time: timestamp says when the source reports that the event occurred.
  2. Observation time: observed_timestamp says when a collection component observed the event, not necessarily when the backend ingested or stored it. A gap can indicate buffering, network delay, clock error, or collector backlog; it does not identify the cause by itself.
  3. Source and resource: service, environment, version, instance, pod, container, or host locate the producer. Confirm that you are reading the correct replica and release.
  4. Event identity and body: event_name is a stable machine-oriented category; message is the human-readable description. A message alone is usually too variable for durable filtering.
  5. Attributes: deployment, route, operation, error type, and other bounded fields provide investigation context.
  6. Correlation fields: request, job, interaction, message, trace, or span IDs connect related records. A trace ID is useful only when real trace context exists; it is not interchangeable with a session or request ID.
  7. Severity: treat the level as an initial hint, not a verdict.

Severity names are not universally equivalent. RFC 5424 defines Syslog severity codes, but it also notes that the meaning assigned by different originators is not standardized. An application may log a handled retry as ERROR, while a failed business outcome may appear as INFO. Learn the producer's conventions and verify behavior rather than treating the label as ground truth.

Preserve original values when a collector normalizes field names or severity. Normalization makes different sources searchable together, but the original value helps resolve mapping mistakes later.

Prove the log set is complete enough

No log view is automatically a complete history. Before concluding that an event did or did not happen, identify the path from producer to the place you are reading:

application -> stdout, file, or journal -> local runtime -> collector -> optional gateway -> backend

Ask what can be lost, delayed, transformed, or duplicated at each boundary. Check the following:

  • Was the relevant process running, and did it restart during the window?
  • Are you looking at the current file or container when the event could be in a rotated file or previous container instance?
  • Does the application write to stdout, stderr, a file, the journal, or more than one destination?
  • Did a multiline exception become several records, or did unrelated lines get joined?
  • Were queues full, retries exhausted, records rejected, or files removed before collection?
  • Did sampling or source-side filtering intentionally omit the event class?
  • Does retention still cover the requested window?
  • Are producer and collector clocks synchronized closely enough for ordering?

Transport and storage also affect interpretation. RFC 5424 defines a Syslog message format, not exactly-once delivery. It provides no message-delivery acknowledgment; an originator or relay can send the same message to multiple destinations, and its reliability section warns that messages can be lost. Other pipelines have their own retry and queue semantics. Expect possible gaps, duplicates, and out-of-order arrival unless the complete path proves otherwise.

Timestamp sorting cannot repair a bad clock. When records from two services disagree, compare event time, observation time, monotonic durations if available, and causal identifiers. Say “the gateway recorded this first” rather than “this definitely happened first” when clock uncertainty remains.

An empty query means only that the selected view returned no matching retained records. It does not prove that the source emitted nothing. Verify the process, output destination, collector health, receiver acceptance, field mapping, and retention before treating absence as evidence.

Narrow before you interpret

Start with exact, stable dimensions rather than severity or broad text:

  1. environment and service;
  2. workload, host, pod, container, or process;
  3. version and deployment ID;
  4. request, interaction, message, job run, or trace ID;
  5. route, operation, or event name;
  6. severity and message text.

This ordering reduces the chance of mixing a staging failure with production traffic, a prior release with the current one, or another customer's request with the event under investigation.

Once you have the failing sequence, find the first divergence, not merely the first ERROR. Compare it with a successful peer that used the same route, release, and approximate time. Look at the records immediately before and after the divergence. A timeout may be the final symptom; an earlier configuration lookup, rejected credential, queue delay, or dependency response may be the first meaningful difference.

Keep three statements separate in notes:

  • Evidence: “Request req_example_7f2c entered checkout-api version 7.4.2; the next retained event records gateway_timeout.”
  • Hypothesis: “The payment dependency may have exceeded the application's deadline.”
  • Conclusion: “A controlled reproduction and dependency evidence showed that the configured deadline expired before the response arrived.”

Do not promote a plausible hypothesis to a conclusion because several lines tell a coherent story. Look for contradictory records, a successful comparison, and evidence from the boundary that owns the suspected failure.

Read plain-text and JSON logs safely

The following examples use /var/log/acme/checkout.log as a documentation path and an obviously synthetic request ID. Replace them only with an authorized local path and non-secret value.

Page through a file with less

less /var/log/acme/checkout.log

Inside less, enter /req_example_7f2c to search forward and press n for the next match. less avoids printing an entire large file into the terminal or session transcript. It is read-only, but the contents may still be sensitive.

Inspect a bounded tail

tail -n 200 /var/log/acme/checkout.log

This shows a fixed number of recent lines. It does not guarantee that all lines belong to the same time window, and it may begin in the middle of a multiline event.

Find an exact string with context

grep -F -n -C 4 -- 'req_example_7f2c' /var/log/acme/checkout.log

-F treats the value as a fixed string, -n prints line numbers, and -C 4 includes four surrounding lines. The -- ends option parsing, so a value beginning with a hyphen is not interpreted as an option. These behaviors are documented in the GNU grep manual. Avoid putting secrets directly in shell arguments because command histories and process inspection can expose them.

Text context is physical line context, not guaranteed event context. Rotation, interleaving, and multiline output can separate related records.

Filter JSON Lines with jq

For a file containing one JSON object per line:

jq -c '
  select(
    .service == "checkout-api" and
    .environment == "production" and
    .request_id == "req_example_7f2c"
  ) |
  {timestamp, observed_timestamp, severity, event_name, message, error_type}
' /var/log/acme/checkout.jsonl

The jq 1.8 manual documents select() and object construction. This query scans the named file but returns only records matching the service, environment, and request ID, with only selected fields to reduce accidental disclosure. Malformed JSON normally raises a parse error; absent or differently named fields can yield no matches. Inspect a small, protected raw sample when the result is unexpectedly empty.

Read service and container logs

These tools expose different boundaries. journalctl reads the systemd journal. Docker reads logs available through the container's configured logging path. Kubernetes asks the kubelet for a container's current or previous log. None automatically combines application files, dependency logs, and every prior instance.

Filter a systemd unit by time

sudo journalctl \
  --unit=checkout-api.service \
  --since='2026-09-16 14:25:00 UTC' \
  --until='2026-09-16 14:40:00 UTC' \
  --no-pager

The current journalctl manual documents unit and time filters. sudo is required only when local journal permissions require it; elevated access can expose records from unrelated services. Keep the unit and window narrow. If the service name is uncertain, verify it with service ownership or deployment configuration instead of dumping the entire journal.

Bound Docker output by time and count

docker logs \
  --since='2026-09-16T14:25:00Z' \
  --until='2026-09-16T14:40:00Z' \
  --tail=300 \
  --timestamps \
  checkout-api

The Docker CLI reference documents --since, --until, --tail, and --timestamps. Use an explicit offset or Z; otherwise timezone interpretation can depend on the client. Docker access is security-sensitive and commonly grants broad control over the daemon, even though this command only reads logs.

docker logs shows output available through the container's logging configuration, normally stdout and stderr. It does not reveal application logs written only to an internal file. A stopped or recreated container may require selecting the correct historical container, if it still exists. For Compose-specific scoping, see the Docker Compose logs guide.

Select the exact Kubernetes pod and container

kubectl logs \
  --namespace=payments \
  POD_NAME \
  --container=api \
  --since-time='2026-09-16T14:25:00Z' \
  --tail=300 \
  --timestamps=true

If the container restarted, inspect the previous terminated instance:

kubectl logs \
  --namespace=payments \
  POD_NAME \
  --container=api \
  --previous \
  --tail=300 \
  --timestamps=true

Replace POD_NAME with a pod verified in the intended cluster, context, and namespace. The current kubectl logs reference documents these flags and notes that --since and --since-time are mutually exclusive. Specify the container in multi-container pods and an explicit tail count. Kubernetes keeps node-level container logs only within runtime and rotation limits; --previous is not an archive of every restart. The kubectl logs reference guide covers workload and selector cases in more detail.

These commands require authorization and may reveal customer or credential data. Confirm the kubeconfig context before reading production logs. Do not paste broad output into chat, tickets, or an AI tool.

Verify the actual outcome

A log level is not a business result. Neither is a process exit code by itself.

An ERROR record can describe a handled retry that later succeeded. An INFO record can describe a rejected payment, skipped import, or stale dataset. Exit status 0 means the process reported success according to its implementation; it does not prove that the correct records were written, the intended notification arrived, or the customer-visible state changed.

After forming a hypothesis, verify the outcome at the responsible boundary:

  • For an HTTP request, check the final response and authorized application state.
  • For a job, check its completion marker, expected artifact, and business invariant.
  • For a message consumer, check acknowledgment and the intended state transition, while accounting for retries and duplicates.
  • For a deployment, check the active version and a representative user operation, not merely the deploy command's exit status.

If verification changes state, obtain approval and use the system's safe procedure. Log reading is evidence collection; it is not authorization to replay requests, restart workloads, change configuration, or delete queue items.

Handle logs as sensitive, untrusted data

Logs can contain credentials, personal data, customer content, or strings controlled by an attacker. The OWASP Logging Cheat Sheet recommends excluding or sanitizing secrets and protecting log access, transport, storage, retention, and disposal.

Apply these rules while investigating:

  • Do not search with or share access tokens, passwords, session IDs, signing keys, cookies, connection strings, or full payment data.
  • Minimize copied output. Prefer a few necessary fields and a bounded time range.
  • Treat message text and remote attributes as untrusted input. Escaped newlines, delimiters, and control characters can forge apparent records or corrupt text exports.
  • Do not execute text copied from a log. A command or URL in a record is evidence, not an instruction.
  • Restrict production log access by role and audit it where required.
  • Keep raw evidence protected, and redact only in a derived copy. Record what was removed so collaborators do not mistake the redacted view for the original.
  • Apply retention and deletion requirements to local files, exports, tickets, and investigation transcripts, not only the central backend.

Correlation identifiers should be pseudonymous and bounded. They help connect records, but they are not proof of identity or authorization. Validate identifiers at trust boundaries and never place personal data or secrets inside them.

Write logs that are easier to read

Readable logs are designed before an incident. Emit one structured event for one meaningful state transition, then keep the schema stable enough for queries and alerts.

A useful application event usually includes:

  • an unambiguous event timestamp and, when added downstream, a distinct observation timestamp;
  • a stable event name and concise human-readable message;
  • service, environment, version, and instance or workload identity;
  • producer-defined severity with documented conventions;
  • a safe request, interaction, job run, or message ID;
  • an error type or outcome category rather than a secret-bearing raw exception string;
  • source location only when it is useful and safe;
  • schema version when field meaning may evolve.

Keep field names and types consistent. Do not sometimes emit a duration as a number and sometimes as text. Bound high-cardinality values, and put variable detail in attributes rather than constructing a new event name for every user or URL. Preserve stack traces as one logical event when the collector supports multiline handling, then test the real output path.

Avoid logging full request or response bodies, environment dumps, headers, database connection strings, or exception arguments by default. Sanitize carriage returns, line feeds, and delimiters in untrusted text. A readable message should explain what happened without duplicating every structured field.

Test logging end to end with a unique synthetic marker that contains no secret. Confirm the event appears once or according to documented retry behavior, the original timestamp is preserved, expected fields and types survive mapping, multiline content stays coherent, and pipeline failure counters remain healthy. The broader log management best-practices guide covers schema, buffering, access, retention, and recovery tests.

Use Fluxtail after delivery is verified

Fluxtail is a paid, self-service Starter/Pro log-management service. When a supported sender or collector has already delivered events and the receiver mapping is verified, its documented search and filters, Live Tail, and log alerts can help narrow retained evidence by time and available fields.

Fluxtail does not automatically parse every source, reconstruct records dropped upstream, guarantee event ordering or delivery, detect the absence of a job event, or provide metrics, tracing, APM, incident management, or autonomous remediation. Field names and filter availability depend on what the source emitted and how the receiver or collector mapped it. Verify one known raw row before relying on a saved search or alert.

The built-in AI log analysis and hosted MCP interface are separate investigation surfaces. Hosted MCP uses account-bound OAuth with PKCE. Keep either interface scoped to the relevant logs and time window, and verify summaries against the retained raw rows. A proposed configuration mutation is not evidence and should require explicit review and the documented short-lived confirmation step.

Incident log-reading checklist

Use this sequence when pressure makes it tempting to search broadly:

  1. State the symptom without guessing the cause.
  2. Confirm environment, service, workload or host, version, and timezone.
  3. Choose a bounded time window around the symptom.
  4. Capture a safe request, run, message, deployment, or trace identifier.
  5. Check the producer output destination and the complete collection path.
  6. Account for restarts, rotation, retention, sampling, multiline handling, queue drops, duplicates, clock skew, and delayed observation.
  7. Narrow by exact source and correlation fields before filtering by severity or message.
  8. Read the records before and after the first divergence.
  9. Compare the failing sequence with a successful peer on the same release.
  10. Write evidence, hypotheses, and conclusions as separate statements.
  11. Verify the user-visible or business outcome at the responsible boundary.
  12. Share only the minimum sanitized evidence and preserve the protected original.

The durable habit is simple: begin with a precise question, reduce the evidence without erasing context, and stop only when the claimed outcome has been verified. That is how to read logs without turning a plausible story into a false conclusion.