Server

Why Does a systemd Service Keep Restarting? A Debugging Guide

Quick summary

If a systemd service keeps restarting, inspect its status and journal first; adding a longer RestartSec value immediately can hide the actual failure.

  • Use systemctl status to see the latest exit code, signal and systemd warning.
  • Search journalctl -u for the application error, missing environment variable or port conflict.
  • After fixing the application problem, reload the unit only if its definition changed, then start the service in a controlled way.
  • Use reset-failed when systemd has rate-limited attempts; this does not repair the application.

When an application closes and reopens every few seconds, the problem may not be limited to systemd. systemd may be starting the process again because Restart= is enabled in the service definition. The underlying cause could be an incorrect command, missing file, wrong permission, unavailable port, missing environment variable or an application startup failure.

This guide narrows the diagnosis from the systemd layer toward the application layer. First confirm the service name and its actual runtime settings, then interpret the exit code and journal entries. Before changing anything, back up the current unit file and application configuration. After each fix, verify that the service remains healthy instead of only checking whether it starts once.

Confirm the restart loop

The first sign usually appears in systemctl status. The service may briefly show active (running), then become failed or start again. Run the command with your service name:

sudo systemctl status example-service.service --no-pager -l

Replace example-service.service with the real unit name. You need sudo access and a systemd-based operating system where the service is defined. In the output, focus on these lines:

  • Active: Shows the current state and the time of the latest state change.
  • Main PID: Identifies the main process. If it changes on every attempt, systemd may be creating a new process each time.
  • code=exited and status=: Show whether the application ended normally or returned an error code.
  • code=killed and the signal information: Suggest that the process was terminated by a signal.
  • Start request repeated too quickly: Means systemd has seen many failed starts in a short period and is temporarily limiting new attempts.

The last warning is not the root cause. It describes the result of repeated failures and systemd’s protective behavior. The real error is usually in earlier application log lines.

Identify the source of the unit file

A service file may have been supplied by the distribution, installed by a package manager or extended with an override you created. Inspect the effective configuration and drop-ins together:

sudo systemctl cat example-service.service

To list the working directory, user, start command, environment files and restart settings separately, run:

sudo systemctl show example-service.service -p FragmentPath -p DropInPaths -p ExecStart -p User -p Group -p WorkingDirectory -p EnvironmentFiles -p Restart -p RestartSec

The systemctl cat output shows the main unit and additional settings under /etc/systemd/system/example-service.service.d/. If the same setting is defined in several places, diagnosis becomes harder, so assess the main file and every override together. The available properties can vary with the systemd version; the unit text and drop-in files remain the primary reference.

Caution

Direct changes to a package-provided unit can disappear during an update. If a persistent change is needed, back up the existing definition and use systemctl edit to create a narrowly scoped override where possible. Record how to undo any change that affects the way the application starts.

Find the actual error in the journal

The last few lines shown by the status command are often not enough. Review all recent start attempts with timestamps:

sudo journalctl -u example-service.service -b --no-pager -n 200

-b limits the output to the current boot, while -n 200 returns the last 200 lines. If the problem began during an earlier boot, use -b -1 to inspect the previous one. To filter a specific period:

sudo journalctl -u example-service.service --since "30 minutes ago" --no-pager

Look at application lines immediately before or after systemd’s service-start messages. For example, address already in use indicates that another process has the port, permission denied indicates rejected access to a file or directory, and no such file or directory may refer to the command, working directory or an expected file.

Interpret the exit code and signal

A value such as status=1/FAILURE means the application exited with an error, but it does not identify the cause by itself. Match the code with the application’s documentation and the surrounding journal lines. If you see status=0/SUCCESS but the service still restarts, the application may be exiting normally while the unit contains Restart=always.

SIGTERM or SIGINT can be consistent with a controlled stop request. SIGKILL may be related to a timeout, an external termination or memory pressure. Check kernel messages as well:

sudo journalctl -k --since "30 minutes ago" --no-pager | grep -Ei "oom|out of memory|killed process"

An empty result does not definitively rule out a memory problem. It only means those expressions were not found in the kernel journal. If the service writes its application logs to a separate file, check that file’s permissions and the same time range.

Test the start command outside systemd

If the journal shows an application error, the narrowest next test is to run the command used by systemd with the same user and working directory. Take the command from systemctl cat. Testing as root can be misleading; use the same account as the service.

sudo -u application-user sh -c 'cd /srv/example-app && exec /usr/bin/example-app --config /etc/example-app/config.yml'

Replace application-user, the directory and the command according to your unit file. Do not use a diagnostic command that writes live data or runs a migration. Test the normal start command only within an appropriate maintenance plan. If the command needs interactive environment variables, check separately whether those variables are defined for the systemd service.

A command that fails in the terminal points to an application problem. A command that works in the terminal does not prove that it will work under systemd, because the two environments can differ:

  • The service uses another user and cannot access a file or socket.
  • WorkingDirectory is different, so relative paths resolve elsewhere.
  • PATH, HOME or application variables from a shell profile are not passed to the service.
  • Secrets or configuration files are not readable by the service account.
  • A terminal-launched process already uses the port, so the new systemd attempt cannot bind to it.

For diagnosis, do not copy an entire interactive environment into the unit. Add only the variables the application needs, either in the unit or an EnvironmentFile. Do not make secret files broadly readable; verify that the service account can read the required file instead.

Narrow down common root causes

Incorrect ExecStart or a missing file

Check the binary path, working directory and arguments in ExecStart exactly. Python applications installed in a virtual environment may use a different interpreter from the system one. Node.js applications may use a different Node path. Check the absolute path and file access:

test -x /usr/bin/example-app
sudo -u application-user test -x /usr/bin/example-app
sudo -u application-user test -r /etc/example-app/config.yml

The first command confirms that the path exists and is executable; the second checks the same condition as the service account. Before correcting a path, determine where the unit file comes from. Change only the missing or incorrect argument; do not alter unrelated resource or restart settings at the same time.

Port conflicts and dependency order

If the application reports address already in use while binding to a TCP port, find the process using that port:

sudo ss -ltnp | grep ':8080'

Replace 8080 with the application’s real port. If the process is an older copy of the same application, do not terminate it at random. First determine how it was started and which unit manages it. If another service owns the port, assess the conflicting configuration before changing the application port.

An application can fail several times if a database, socket or network service is not ready at startup. After= controls start order; it does not guarantee that a service is ready. If the application supports readiness checks, use an appropriate health check or the application’s own retry behavior. Adding a long delay only to hide the loop does not solve a dependency problem.

Permissions and the working directory

When User= is defined, the application accesses the filesystem as that user. Directories need not only file permissions but also the x permission that allows traversal through parent directories. Inspect the path during diagnosis:

namei -l /srv/example-app/config.yml

Apply the narrowest permission correction to the specific file or directory you found. Making configuration world-readable is not a safe fix, especially when it contains credentials. Repeat the read test as the service account after the change.

Do not confuse FPM, CLI and queue workers

In a web application, PHP-FPM workers and CLI cron or queue workers are separate processes. Restarting PHP-FPM web workers does not mean that a CLI queue worker stopped for the same reason. CLI work does not consume an FPM worker unless it makes an HTTP request.

This distinction prevents you from diagnosing the wrong unit. For web request failures, inspect the FPM unit logs. For queues or scheduled work, inspect the relevant CLI process, cron entry or worker unit. If you are choosing a scheduler, the existing guide on Cron vs systemd Timers can help; a Timer job’s failure should still be investigated separately from a web service restart loop.

Example scenario

Assume that app-worker.service restarts every five seconds and the journal says that its configuration file cannot be found. The unit defines WorkingDirectory=/srv/app/current, but a symbolic link was removed after a new release was deployed. The narrow fix is to restore the valid release path and the link target; adding only RestartSec=60 would not repair the failure. This is a constructed example scenario used to explain the diagnosis.

When should you change Restart settings?

Restart=on-failure restarts a service that exits with an error code or unexpected signal. Restart=always can restart it even after a normal exit. RestartSec sets the delay between attempts. These settings can be useful for long-running services, but if the application fails in the same way at every start, repeated attempts can fill the logs and hide the root cause.

Fix the application error first. If you need to stop the loop temporarily during diagnosis, stop the service and inspect its journal instead of permanently changing the configuration:

sudo systemctl stop example-service.service
sudo journalctl -u example-service.service --since "10 minutes ago" --no-pager

If you decide to change Restart=, back up the current file or override first. For example, Restart=on-failure may be a narrower choice when the service should restart only after an error, but the correct setting depends on the application’s expected operating model. Do not change a production service without recording the purpose and rollback command.

You can save the currently rendered unit before editing it:

sudo sh -c 'systemctl cat example-service.service > /root/example-service.service.backup'

This backup records the main unit and its visible drop-ins. Also back up application configuration separately, and confirm where each file will be restored before making a data-changing edit.

After many failed attempts, systemd may limit new starts. Once the root cause is fixed, clear the failed-state counter:

sudo systemctl reset-failed example-service.service
sudo systemctl start example-service.service

reset-failed only clears systemd’s failure state. If the application still exits, the service will become failed again. Observe at least several restart intervals after starting it.

Apply the fix and verify it

If you changed the unit or an override, systemd must read the new content:

sudo systemctl daemon-reload

Then start the service and perform several checks:

sudo systemctl restart example-service.service
sudo systemctl is-active example-service.service
sudo systemctl status example-service.service --no-pager -l
sudo journalctl -u example-service.service --since "2 minutes ago" --no-pager

An active result from is-active is not enough by itself. Confirm that the application is listening on the expected port, that the required endpoint responds and that no new errors appear in the journal. For an externally accessed web application, application logs and the web server’s 4xx and 5xx records represent different layers. The existing guide on reading hosting server logs can help with that comparison.

If the problem appeared during a new deployment, rolling back to the previous known-good release may be safer than editing the unit immediately. With versioned symbolic-link directories, confirm that the old target still exists and that its application files and configuration remain compatible. If the service is stable after rollback, investigate the new release’s dependencies as a separate maintenance step.

For WordPress or Node.js applications, systemd is only the start layer. The application’s dependencies, release command and proxy settings also need review. If you are redesigning a deployment process, compare your current definition with an existing guide on PM2 and systemd deployments, and make sure two process managers are not managing the same application at once.

Frequently Asked Questions

Why does a service restart while status still says active?

status is a point-in-time view. The command may run during the interval between two attempts. Use the journal to compare PIDs and the timestamps of each start.

Can a service restart without systemd being the cause?

Yes. An application can exit by itself, a supervisor can send a signal, or the kernel can terminate it under memory pressure. Check the unit’s parent process, application logs and kernel journal.

Should I disable Restart= while troubleshooting?

Not necessarily. Stopping the service temporarily is often less disruptive than changing the unit. If you do change the setting, back up the definition and restore it after the diagnosis.

What should I check if the service works after a manual start?

Compare the manual shell with the unit’s User, WorkingDirectory, ExecStart, PATH and environment files. Also check whether the manual process is already occupying the application’s port.

Final checklist

  • Confirm the real unit name and every active override file.
  • Use systemctl status and journalctl -u to find the exit code, signal and first application error.
  • Test ExecStart, the user, working directory, environment file and file permissions under the same conditions as the service.
  • Check port conflicts, dependency readiness and memory pressure separately.
  • Do not hide the loop with only RestartSec or Restart= changes before fixing the application error.
  • Save a backup and rollback path before changing anything; if the unit changed, run daemon-reload, start the service in a controlled way and verify the journal.

Your next step is to compare the service status with its last 200 journal lines over the same time period and isolate the first real error. Once you identify it, fix the layer responsible for that error rather than changing systemd blindly.

↑