Technology

Diagnosing Linux Server Issues with eBPF

Quick summary

CPU spikes, network latency, disk waits, and unexpected process behavior may not be explained by logs alone. eBPF lets you observe selected Linux events with small, verifier-checked programs that run in the kernel. Many diagnostic probes can attach without restarting the system, but support depends on the kernel, tool, permissions, and event type.

  • Identify the affected layer first; do not investigate CPU, network, disk, and process problems with the same command.
  • Choose bpftrace, BCC, or bpftool for the question you need to answer, and start with narrow probes on production systems.
  • Check kernel, tool, permission, and symbol support before running an eBPF diagnostic.
  • Limit measurements to a time window, process, port, or disk so the observation itself does not create unnecessary noise or overhead.

When CPU usage rises on a server, application logs are often the first place you look. Logs may show when a request was rejected, but they rarely show which processes the kernel delayed, how long a disk request stayed in a queue, or which stage of a TCP connection became slow.

eBPF is a Linux technology that lets you attach observation code to selected event points in the kernel and related user-space execution paths. Without changing kernel source code, compatible tools can observe system calls, scheduler behavior, network events, and file-access activity. This does not replace application logs; it adds timing and process context that logs may not contain.

The goal is not to collect every metric at once. It is to narrow the gap between a symptom and its cause. The commands in this guide assume sudo access, an eBPF-compatible Linux kernel, and the relevant tools installed. Kernel versions, package versions, event support, and permission policies can change the exact command or its output on your distribution. Check each tool’s help output before using a command on a production server.

What questions can eBPF answer?

Start by turning the behavior you want to observe into a specific question. “The server is slow” is too broad for a useful diagnosis. “Is the request-handling process waiting for CPU, or is a disk read delayed?” gives you a measurable direction.

High CPU that process usage does not explain

Standard tools show CPU percentages, but they do not always distinguish between time spent in user space, kernel space, or time stolen by a virtual machine. eBPF-based profile sampling collects call stacks at selected intervals. This gives you a direction for identifying which code path inside a process is being sampled most often.

This is not an exact calculation of execution time. Sampling frequency, available symbols, compiler optimizations, and the selected stack type affect the result. Even so, a profile can reveal a busy loop, lock-waiting pattern, or kernel-call trend that does not appear in application logs.

Next, record the affected PID and the time window with a standard tool such as top or pidstat. Then use a PID-filtered profile rather than sampling every process.

Is network latency caused by the application?

A slow HTTP request may involve DNS, TCP connection setup, TLS negotiation, a remote service, a local queue, or application processing time. eBPF tools can help you observe TCP connection attempts, retransmissions, and socket behavior independently of application logs.

Latency observed at the network layer is not automatically the total response time seen by the user. Compare network events with application request duration and external monitoring data.

Next, identify the connection and process involved. A PHP-FPM worker making an outbound API call is a different investigation from a CLI queue consumer or a database connection. Filter the observation to the relevant PID or destination where the tool supports it.

Is the disk slow, or is the process behaving poorly?

An application can wait even when overall disk utilization is not high. Small random I/O operations, queues in a virtual disk layer, or contention around one file may explain the behavior. eBPF can help you examine block I/O latency together with the process that initiated the operation and its read or write pattern.

Do not confuse this information with file contents. eBPF generally shows timing, process, and system-call context. It does not mean that you need to inspect sensitive application data.

Next, use iostat or an equivalent system metric for the broad storage view, then measure block latency or process activity during the same time window.

Example scenario

Assume a WooCommerce store becomes slow on its payment page during busy periods, while the PHP error log shows no obvious error and CPU usage remains moderate. Your first hypothesis does not have to be “PHP needs more CPU.” You could use eBPF to observe disk waiting and outbound connection behavior for PHP-FPM processes separately, narrowing the investigation. This is a constructed scenario based on an unverified hypothesis; the actual cause must be confirmed with data from the server.

Plan safe observation before installation

Installing eBPF tools and attaching an eBPF program to the kernel are different stages. Installation depends on your package manager and distribution repositories. During observation, permissions, kernel support, symbols, and tool versions become important.

Check the prerequisites

Begin by recording the state of the kernel and available tools without changing the system. The following commands only collect information:

uname -r
command -v bpftrace
command -v bpftool
sudo bpftool feature probe kernel

The final command may require root privileges and can produce a long feature report. Save the kernel version, tool paths, and any reported unsupported features with your diagnostic notes. If bpftool is not found, the command cannot run; that alone does not prove that the kernel lacks eBPF support. You may need to install the appropriate distribution package.

Some tools use BCC libraries, while others use LLVM and Clang components. Select versions from packages supported by your distribution and check the tool’s documented requirements. A probe that uses a tracepoint or helper unavailable on the running kernel will need a different event or tool.

If you install packages on a production server, review the package manager’s change list first and follow your maintenance and rollback procedures. Updating the kernel only for observation is a broader change than diagnosing an unexplained delay. Record the package versions so you can remove or revert the diagnostic tooling if required.

Set permission and security boundaries

Many eBPF observation commands require root or specific capabilities. Running a command with sudo should not mean granting a user unlimited and permanent access. Define a separate diagnostic procedure for the administrators who need it, and account for sensitive information such as command lines, file names, or network addresses in the output.

Start with a short, narrow test in production. Instead of collecting every system call from every process, use a specific PID, command name, port, or time window. Check output permissions before sending results to shared channels.

Verify that the command attached and produced the expected event type. A successful tool launch is not proof that the selected process generated relevant events, so compare the output with the process list and the same-period application metrics.

Separate FPM, CLI, and queue processes

A PHP-FPM worker handling a web request does not have the same operating model as a CLI process running a cron job or queue consumer. A CLI job that does not make an HTTP request does not consume a PHP-FPM web worker. This distinction prevents you from measuring the wrong process and drawing the wrong conclusion.

List the processes first:

ps -eo pid,ppid,comm,args,%cpu,%mem --sort=-%cpu | head -n 20

Separate the PHP-FPM master PID from the worker PIDs, and distinguish both from cron commands and queue consumers. Where possible, apply the eBPF filter to one worker PID. Low activity in the master process may not represent what the workers are doing. A CLI worker can still create network or disk activity, but it does not occupy an FPM worker unless it makes an HTTP request to one.

Investigate CPU and process behavior

When investigating a CPU symptom, use standard measurements first to establish the time window. top, pidstat, or your existing monitoring system should show which PID increased and how long the increase lasted. Use eBPF next to answer a narrower question: which call stack is being sampled for that process?

A narrow profiling example

The following bpftrace example profiles the process with PID 1234 at 49 samples per second. Replace the PID with the worker identified during your investigation:

sudo bpftrace -e 'profile:hz:49 /pid == 1234/ { @[ustack] = count(); }'

The command continues collecting samples until you stop it. Press Ctrl+C to end a short observation window. Meaningful user-space call stacks may require symbol information in the binary or suitable debug symbols. Without symbols, you may see addresses or incomplete names.

After stopping the command, compare the reported stacks with the PID’s CPU usage and application activity from the same period. Do not treat the highest line in the output as the root cause by itself. A profile points to a frequently sampled code path; it does not prove that a business rule is incorrect.

Measure system-call activity

A process may open many files, make frequent timer calls, or generate more system calls than expected. The following example counts openat calls by command name:

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat { @[comm] = count(); }'

This example has no PID filter, so use it briefly and only where the event rate is manageable. On a production system, narrowing it to one process is safer:

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat /pid == 1234/ { @[comm] = count(); }'

Stop the probe after the planned interval and record the count with the interval length. A high call count is not automatically a performance problem. An application may be expected to perform many small file accesses. Validate a suspicious result with system-call duration, file descriptors, and application behavior before changing limits or code.

Tip

Begin eBPF Linux server observability as a diagnostic tool before designing it as a permanent production metric. Sampling cost, data retention, permission models, and alert thresholds require separate decisions for a lasting alerting system. A centralized monitoring and alerting design is a separate operational layer.

Separate network latency into layers

When a network problem is reported, first identify which connection is slow. A single server may handle client requests, database connections, DNS queries, and external API calls at the same time. Instead of collecting every TCP event, narrow the observation to the process creating the connection and the destination port.

Observe TCP connection attempts

If the BCC package is installed, the tcpconnect tool can summarize new TCP connections. For example, to inspect PID 1234:

sudo tcpconnect -p 1234

The syntax may differ slightly between BCC versions, so check tcpconnect --help before running it. The tool may show a connection attempt, but it does not prove that the remote service responded at the application level.

If BCC is unavailable, you can inspect the same question with less detail using bpftrace:

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_connect /pid == 1234/ { @[comm] = count(); }'

This command counts connection system calls; it does not by itself explain the destination address or the elapsed time. Its purpose is to identify unusually frequent connection attempts during a short diagnostic window. Verify the result against ss, application request timing, and the expected traffic pattern.

Interpret retransmissions and possible loss

TCP retransmissions may strengthen the possibility of packet loss or congestion along the network path. A retransmission alone does not prove that the remote server is slow. The local network interface, virtual network, firewall, and destination system must be considered together.

Compare eBPF output with ss -s, connection states, application timeouts, and external monitoring data. If many connections are waiting on a port, do not immediately change kernel or network settings. First confirm which process created them and whether they match expected traffic. If the event rate is high, stop the broad probe and replace it with a narrower process or port filter where supported.

For web performance, combine network measurements with browser and application timings. Finding Core Web Vitals Opportunities in Your Server Logs can help connect server-side observations with perceived performance, but eBPF does not replace browser metrics.

Separate disk access from I/O waiting

For a disk-latency problem, separate two questions: Is the storage layer responding slowly, or is the application generating too many I/O requests? Use iostat or a similar tool for the broad view, then use eBPF to narrow the process and latency distribution.

Measure block I/O latency

If BCC tools are installed, biolatency can show the latency distribution of block-device operations:

sudo biolatency 1 10

Depending on the tool version, this example is expected to produce ten outputs at one-second intervals. Confirm supported options with biolatency --help on your system. Compare the result with the disk device, virtual storage layer, and application request volume during the same period. If the tool reports no events, verify that the server is using a supported block-device path and that I/O occurred during the interval.

To observe file access from a particular process, you can use a broad system-call trace:

sudo bpftrace -e 'tracepoint:syscalls:sys_enter_read /pid == 1234/ { @[comm] = count(); }'

This example counts only read calls; it does not show how long each read waited. Do not conclude that the disk is failing simply because the count is high. The application cache, database query plan, or file-access pattern may produce the same symptom.

Keep the boundary before changing data

Deleting files, clearing a cache, changing a database setting, or changing the I/O scheduler because of a suspected disk problem is not a diagnostic step; it changes system behavior. Record the current configuration, disk-health information, and rollback plan first.

If a data-changing step is necessary, verify that the backup is current and that the restore path has been tested. Save the configuration before editing it, make one targeted change, and define the exact command or file restoration needed to roll back. After the change, repeat the same I/O measurement and confirm application behavior.

Caution

Passing the kernel verifier gives eBPF programs a constrained execution model, but observation commands are not free. Broad filters, high event rates, and long-running stack collection can consume CPU, memory, or output storage. Start with a short duration, one PID, or one event type, and monitor the cost of measurement. Stop temporary probes after verification.

Validate process, file, and kernel events together

A single tool rarely proves a root cause. If PHP-FPM is waiting, the cause could be disk access, a database socket, a file lock, or an external API. eBPF supplies event timing; process lists, application logs, network connections, and storage data give that timeline meaning.

Build a diagnostic flow

  1. Time the symptom: Record when the problem started and ended, which services were affected, and whether traffic changed.
  2. Identify the process: Keep FPM workers, CLI cron jobs, queue consumers, and database processes separate.
  3. Choose one hypothesis: Write a bounded question such as “PID 1234 is waiting on disk reads.”
  4. Run the narrowest eBPF measurement: Use a PID, event type, and short time window.
  5. Compare independent data: Review application duration, system metrics, network state, and relevant log lines from the same time window.
  6. Apply a targeted change: If a setting must change, touch only the confirmed cause and back up the current configuration.
  7. Measure again: Check whether the symptom disappeared or merely moved to another layer.

This flow keeps observation separate from remediation. For example, seeing that a process opens many files does not immediately justify raising the file-descriptor limit. First confirm the application’s expected behavior and the condition that triggered the problem.

Take care with kernel-level symptoms

For a kernel error, lockup, or unexpected restart, eBPF is not a recovery tool by itself. Persistent logs, console output, watchdog information, and kernel dumps must also be examined. If the system has experienced a kernel panic, Server Kernel Panic: Root-Cause Analysis and Recovery can complement an investigation of boot logs and hardware or virtual-infrastructure records.

Capturing an event with eBPF does not recreate a failure that happened in the past and was never recorded. Plan the measurement while the symptom is occurring, while staying within security, performance, and data-privacy boundaries. Verify that the probe itself stopped after the planned observation window.

Choose between bpftrace, BCC, and bpftool

bpftrace is a practical starting point for short, single-purpose observations. You can test a hypothesis quickly with system calls, tracepoints, and profile samples. You still need to check syntax, event names, and compatibility with the running kernel and installed bpftrace version.

BCC provides prepared tools for network, disk, and process behavior. Tools such as tcpconnect and biolatency let you measure common questions without writing every probe from scratch. The BCC package version may not match your kernel or Python dependencies, so verify every command with its help output and confirm that the output represents the event you intended to measure.

bpftool is a lower-level management tool for examining maps, programs, and kernel features. It is useful when you need to verify that an eBPF program is actually loaded. If you build a permanent observability system, design its data model, sampling, permissions, retention, and upgrade procedure separately.

Frequently Asked Questions

Can eBPF show why a process is waiting on a lock?

It can help identify scheduler, wait, or lock-related patterns, but the exact interpretation depends on the event, lock implementation, and available symbols. Confirm the result with application behavior and other system metrics. If the suspected lock is in application code, an application profiler or lock-specific instrumentation may provide more direct evidence.

Can eBPF inspect encrypted application traffic?

eBPF can observe connection and socket behavior without reading encrypted payloads. Payload-level details remain subject to where encryption occurs, the application’s instrumentation, and your security and privacy rules. Start with connection timing and process context unless payload inspection is explicitly required and authorized.

Should eBPF probes run continuously on a web server?

Not automatically. Continuous collection needs an explicit cost, retention, access, and alerting design. A short, filtered diagnostic is often more appropriate when you are investigating one incident. If a probe becomes permanent, document its ownership, upgrade compatibility, stop procedure, and data exposure.

What should I do if a trace produces no useful symbols?

Check whether the relevant binary has symbols, whether debug symbols are available, and whether the selected probe supports your kernel and tool version. An address-only stack can still show repeated activity, but it limits interpretation. Preserve the raw output and correlate it with the PID, binary version, and deployment time before drawing a conclusion.

Actionable checklist

  • Limit the symptom to CPU, network, disk, or process behavior.
  • Check the kernel version, installed tools, required permissions, and supported event types.
  • Separate PHP-FPM, CLI cron, and queue-process PIDs.
  • Start with a short measurement covering one PID or one event type.
  • Compare eBPF output with logs, standard system metrics, and application timing.
  • Record the output and confirm that the probe measured the intended event.
  • Check the backup and rollback plan before changing configuration or data.
  • Repeat the same measurement after remediation and stop temporary observation programs.

Your next step is to write one diagnostic hypothesis for the recurring symptom and select the narrowest suitable tool. For a CPU increase, start with a PID-filtered profile; for network latency, observe connection behavior; for disk waiting, measure I/O latency. As the scope narrows, eBPF Linux server observability can turn an uncertain “the server is slow” report into a timeline that you can test.

Emre

↑