Server monitoring checks whether the services running on a host work for their users, whether the host has enough capacity to keep them working, and whether the monitoring path itself is trustworthy. Start with a user-visible operation or an external check. Then use CPU, memory, storage, network, process, and log evidence to explain a symptom or catch an imminent limit. A green host dashboard cannot prove that an application works; a red CPU graph does not by itself prove that users are affected. Google SRE's monitoring guidance distinguishes externally observed symptoms from internal signals for exactly this reason.
The practical design has three views: outside-in service health, inside-out host and process health, and telemetry-path health. Alert each view for a different decision. A failed user operation may justify an urgent page; a slowly filling disk may justify planned capacity work; a stopped metrics scrape means the system can no longer make a reliable claim about that host.
Start with the service, not the machine
Write down one important user journey before choosing a dashboard. For an API, that might be an eligible request reaching a response. For a background service, it might be accepted work reaching a valid completed state. Define success, failure, and unacceptable delay at one observation boundary. Record the service, environment, owner, and how the check works when traffic is sparse. Google SRE describes latency, traffic, errors, and saturation as its four golden signals, but their exact measurement depends on the service.
An external or black-box check tests what a client can see. It can catch a DNS, network, TLS, routing, or authentication problem while the server's internal counters remain normal. Keep the check safe: use a read-only or dedicated synthetic operation where possible, avoid creating customer-visible data, and verify that its credentials and network path represent the behavior you care about. A check from one region does not establish worldwide availability.
Internal or white-box signals show work being done and resources being consumed. They help explain why a service is slow or close to failure. Keep the two views connected but independent enough that a broken server or collector cannot hide its own failure. A service might return errors while CPU is idle because its database is unavailable; it might use most of a CPU core while serving every request within its target. That is why a utilization number belongs beside the service outcome, not in place of it.
Define an SLI or alert ratio with numerator and denominator measured at the same boundary. A count of application errors divided by all load-balancer requests can mix different request populations. At low traffic, one failure can swing a percentage sharply, while no traffic can make a ratio undefined. The alert must say how it handles a small sample and how it distinguishes no requests from no measurements. For SLO design, see the service level objective guide.
Collect host and process signals that answer a question
Host metrics are useful when they indicate saturation, imminent exhaustion, or a likely explanation for a user symptom. Do not collect every available counter simply because an exporter exposes it. Give each dashboard panel a question: “Is the process stalled?”, “Is this filesystem approaching an operational limit?”, or “Did receive traffic stop when clients reported failures?”
CPU and runnable work
CPU utilization describes time spent in modes such as user, system, idle, or wait as exposed by the operating system and collector. It does not directly measure useful throughput or user latency. A sustained high value may be healthy for a compute-bound batch job; low utilization may coexist with a blocked thread pool or dependency. Compare CPU time with request rate, latency, process-level consumption, and runnable work before concluding that CPU is the cause.
On Linux, Prometheus Node Exporter exposes kernel and hardware metrics, including node_cpu_seconds_total. It is a cumulative counter by CPU and mode, so a graph of the raw total is not a CPU percentage. Use an interval rate or another correctly defined derived view, and document the unit and averaging window. Name and label metrics consistently; the Prometheus naming guide notes that each metric should describe one quantity and unit.
Memory and pressure
Memory questions are about reclaimability and work being stalled, not one generic “RAM used” percentage. Check available memory, the working set or resident memory of the relevant process, allocation or heap behavior where the runtime exposes it, reclaim activity, swap, and out-of-memory events. A cache can use otherwise idle memory without harming the service. Conversely, repeated reclaim and stalled work can hurt latency before a simple capacity percentage looks dramatic.
Linux Pressure Stall Information (PSI) reports how much time tasks are stalled on CPU, memory, or I/O contention. Its some line means at least some tasks were stalled; full means all non-idle tasks were stalled. The avg10, avg60, and avg300 values are recent-window percentages, not a count of stalled tasks. System-level CPU full is undefined and reported as zero for compatibility, so do not use it as a CPU-saturation signal. A read-only check on a Linux host with PSI enabled is:
cat /proc/pressure/memory
cat /proc/pressure/io
These files describe the host where the command runs; a container or VM boundary may require a different scope. Read a trend and compare it with service latency and workload behavior rather than paging on an invented universal PSI cutoff.
Filesystems and storage I/O
Monitor each filesystem that matters to the service: available bytes, inode availability where relevant, usage growth, and whether writes succeed. The Node Exporter guide identifies node_filesystem_avail_bytes as bytes available to non-root users. That is not the same as total free blocks for every privilege level. A small application data mount can matter more than a mostly empty root mount; exclude or label ephemeral, read-only, and irrelevant mounts deliberately.
Disk utilization alone does not explain storage latency. Look at I/O throughput and device or application wait, queueing and error events if the platform exposes them, and the affected operation's latency. A host can have free space but suffer slow I/O, or be close to full while the application is still serving normally. The alert depends on growth, time to intervention, and actual failure mode. A capacity ticket may be more useful than a page until the margin becomes urgent or writes fail.
Network and dependencies
Track received and transmitted bytes, errors and drops where available, connection behavior, and the application result. Node Exporter documents node_network_receive_bytes_total as a cumulative receive counter; use an interval rate for traffic rather than graphing the raw counter as bandwidth. A flat line can mean quiet traffic, a broken scrape, or a disconnected path. Compare it with load-balancer or client-side observations before deciding.
Do not assume a healthy host network proves that an upstream API, DNS resolver, or database works. A bounded external request or dependency-specific service check may identify a path problem that interface counters cannot. Keep probes safe and authorized; do not repeatedly submit state-changing customer transactions just to create a health signal.
Processes and service dependencies
Watch whether the service process exists, whether it is ready to handle its intended work, and whether restarts, crashes, thread or connection exhaustion, or a stuck queue affect outcomes. A running PID is weaker evidence than a successful operation. For a server with multiple services, label ownership explicitly so one healthy process does not mask another failed service on the same host.
Use application-level readiness and success signals only where their contract is clear. A liveness endpoint may report that a process responds while a required dependency is unavailable; a deep readiness check may itself overload a struggling dependency. Understand what each check tests, what it omits, and how frequently it runs.
Assemble a safe collection path
An agent or exporter on the host can expose hardware and kernel metrics to a metrics system. Prometheus's Node Exporter guide demonstrates a local /metrics endpoint and scraping it from Prometheus. That endpoint is a telemetry interface, not a public web page. Bind or firewall it according to the monitoring network and authentication model; do not expose a host-metrics port to the internet for convenience.
On a host where Node Exporter is already running on loopback, this read-only check returns a bounded status and body-size summary:
curl --fail --silent --show-error --max-time 5 --output /dev/null --write-out 'HTTP %{http_code}; bytes %{size_download}\n' http://127.0.0.1:9100/metrics
Check for HTTP 200 and a nonzero body size. The command downloads the endpoint but does not print the metric body; it proves only that the local exporter responded, not that Prometheus scraped it, a recording rule evaluated, or an alert reached the responder. Test those stages separately, and use the metrics system's expression browser for normal inspection.
Keep a stable host or instance identity, service and environment labels, units, and metric definitions. Avoid request IDs, customer IDs, or other unbounded values as metric labels; each unique label set creates more time series. The Prometheus naming guidance recommends consistent units and one logical quantity per metric. If hosts are ephemeral, keep the mapping from instance to service and deployment history so a new instance does not look like an unexplained disappearance of the old one.
Collection needs its own failure policy. Scrapes can time out, credentials can expire, agents can restart, and remote transport can lag. Decide whether data is sampled, buffered, retried, or dropped; monitor gaps and staleness. A graph with no new points is a monitoring incident, not a reassuring zero. An independent external check can catch failures shared by the target and the telemetry system.
Alert for decisions, not every abnormal number
Separate three actions:
- Page for an urgent service symptom. A sustained user-visible failure or severe latency regression has an owner and an immediate safe response. If a ratio is used, define the same-boundary numerator and denominator, minimum traffic or absolute fallback, and missing-data behavior.
- Create a capacity ticket or planned action. Disk growth, memory trend, or CPU headroom may need work before users suffer, but not a nighttime interruption. Page only when the limit is imminent and a timely action is required.
- Alert on a broken observation path. Stale scrapes, exporter failure, missing expected targets, or failed notification delivery make other alerts untrustworthy. Test the full route from a synthetic signal through evaluation and notification.
Prometheus alerting practices recommend symptom-oriented, simple alerts and external black-box checks alongside internal monitoring. Google SRE similarly emphasizes actionable pages and keeps symptom versus cause distinct. Neither source establishes a universal CPU, RAM, disk, or latency threshold for your application. Choose conditions from the service's target, baseline, traffic, recovery time, and tested capacity. Add a persistence window only when it suppresses harmless blips without hiding a total outage.
Every alert should identify its service and environment, affected operation or mount, observed condition and time window, owner, evidence link, runbook, escalation route, and what missing data means. Do not put credentials, customer data, or unbounded log payloads into a notification. Review alert quality after real firings, incidents, and service changes, rather than assuming a fixed calendar alone makes rules safe. The alerting best practices guide expands the response contract.
A bounded investigation example
Consider an illustrative case: an external check reports slow responses from a service, while its host dashboard shows rising I/O pressure. This is a diagnostic exercise, not a claim that I/O caused the slowdown.
First, verify the symptom at the same edge and time window. Check whether request traffic and error/latency measurements are present, whether the external probe is healthy, and whether the affected release or region changed. If only one probe fails, compare from a second vantage point before declaring a server-wide problem. If the metrics scrape is stale, treat the host charts as unavailable evidence, not as normal values.
Next, compare host and process signals. Check the relevant filesystem's available bytes and inode state, I/O throughput or wait, memory pressure, and the service process's restarts. Read /proc/pressure/io on the correct Linux host if PSI is enabled, and compare its recent windows with the time of the symptom. Inspect only the affected device or mount; a quiet root disk says little about a saturated data volume. A high PSI value indicates stalled tasks, but it does not by itself identify which application or device is responsible.
Then open a bounded set of service and system log records around the first slow operation. Look for write failures, storage errors, timeouts, retries, or a deploy marker, and compare with a successful peer. Preserve the original timestamps and distinguish event time from collection time when logs arrive late. Do not dump secrets or entire request bodies into the investigation record. A trace, if the application already emits one, may show where an instrumented request spent time; it is not required to explain every host problem. OpenTelemetry's signal model keeps metrics, logs, and traces separate.
Finally, test the leading hypothesis using the smallest safe, authorized change or a read-only comparison. If the cause is still uncertain, retain the alternative explanations. Verify recovery at the original user-impact boundary and confirm that the host signal, service outcome, and telemetry pipeline agree. A quiet error log cannot prove recovery if the collector stopped sending records.
Logs complement server metrics
Metrics are efficient for trends and conditions such as capacity, throughput, and latency distributions. Logs record specific events and context: a failed write, an exception, a service restart, or a dependency timeout. Logs can alert on a configured event, and metrics can help diagnose a fault; there is no rigid “metrics detect, logs explain” rule. Metrics vs logs covers their different data models in more detail.
Collect logs through an appropriate source and collector, normalize only fields whose mapping is verified, and protect sensitive data. Search by service, environment, host, version, and a bounded time window before broad text searches. A correlation ID is useful for finding one operation, but it is not identity or authorization. Logs may be delayed, sampled, dropped, duplicated, or retained for a different period than metrics. Those limits matter before using a log count as a service-level denominator or absence as a health check.
Fluxtail is a paid Starter/Pro, logs-focused complement to server monitoring. Once supported inputs deliver records to named streams, its search and filters and Live Tail can help investigate received rows; log alerts can surface configured conditions in delivered data. Field-level filters depend on what the source and receiver actually map. Fluxtail does not provide native host metrics, traces, APM, or SLO computation, and it cannot show an event that never reached it.
Its built-in AI chat is separate from the hosted MCP server. Hosted MCP uses account-bound OAuth with PKCE; authorized read tools can query logs, while operator changes are proposed first and require short-lived confirmation. Keep any investigation bounded and verify summaries against raw rows and the service outcome. Neither interface replaces the metrics system or makes a server change on its own.
Server-monitoring checklist
For each important service and host, confirm that:
- A user-visible operation and an external check establish whether the service works, with documented scope and low-traffic behavior.
- Host and process signals cover CPU, memory/pressure, filesystem and inode capacity, I/O, network, and readiness where relevant; units and label identity are stable.
- Alerts distinguish urgent symptoms, planned capacity work, and failures of collection or notification, with an owner and safe first action.
- The metrics path, log path, and external check have each been tested during a target outage, collector gap, restart, and credential change.
- Investigation can narrow to a service, environment, version, host, mount or operation, and time window without exposing secrets or assuming perfect delivery.
- A mitigation is verified at the user-impact boundary, not merely by a green host panel or an empty log search.
Server monitoring is useful when it supports a decision. Measure the service's outcome first, use host telemetry to understand constraints, and keep an independent way to notice when the measurement system itself goes silent.