Debugging a systemd service without guessing
A repeatable path from failed unit to useful evidence using systemctl, journalctl, exit status, and the unit's real runtime context.
On this page
- Start with the unit state
- Exit statuses that come from systemd, not your program
- Read the journal narrowly
- Restarts hide the first failure
- Timeouts are usually about readiness
- Reproduce the runtime context
- Run a shell inside the unit’s sandbox
- Check what the unit depends on
- Change one variable at a time
- The whole path, in order
When a service fails under systemd, the fastest route to the answer is usually to inspect the state that systemd already collected rather than immediately editing the unit or restarting it repeatedly.
systemd records a lot about every run: why the unit stopped, how the main process exited, how many times it was restarted, and every line it wrote to stdout or stderr. This post works through that evidence in order, from the cheapest question to the most expensive, and ends with a way to reproduce the service’s exact environment in a shell.
Start with the unit state
systemctl status example.service
systemctl show example.service \
-p Result \
-p ExecMainCode \
-p ExecMainStatus \
-p NRestartsstatus is for humans: the last few log lines, the cgroup’s processes, and a one-line summary. show is for precision. It prints properties as Key=value, after every default and drop-in has been applied, which makes it safe to script against.
The important distinction is between what systemd attempted and what the process actually returned. A unit can fail because the process exited non-zero, because an executable could not be started, because a dependency failed, or because a timeout expired. Result tells you which:
Result= |
What happened | Look next at |
|---|---|---|
exit-code |
The main process exited non-zero | ExecMainStatus, then the journal |
signal / core-dump |
The main process was killed by a signal | ExecMainStatus holds the signal; coredumpctl for dumps |
timeout |
A start or stop step took too long | TimeoutStartSec, Type= and readiness |
watchdog |
The service stopped sending watchdog keep-alives | Whether the process is hung |
oom-kill |
The kernel OOM killer killed a process in the unit | MemoryMax, journalctl -k |
start-limit-hit |
It restarted too often, too quickly | The first failure, not the last |
resources |
systemd could not set up the environment | The journal; often a missing directory or user |
ExecMainCode is the kernel’s view of how the process ended: 1 for a normal exit, 2 for killed by a signal, 3 for killed with a core dump. When it is 1, ExecMainStatus is the exit status; otherwise it is the signal number.
Exit statuses that come from systemd, not your program
Here is what that looks like when the binary itself is missing:
root@web-01:~$ systemctl --failed --no-legend● example.service loaded failed failed Example service$ systemctl show example.service -p Result -p ExecMainStatusResult=exit-codeExecMainStatus=203$ journalctl -u example.service -b -n 2 --no-pager -o catexample.service: Failed to locate executable /opt/example/bin/server: No such file or directoryexample.service: Failed at step EXEC spawning /opt/example/bin/server: No such file or directory
Status 203 is systemd’s EXEC failure: the process never started, so the application’s own logs will be empty. No amount of restarting will change that.
Statuses from 200 upward are reserved by systemd for failures that happen between forking and running your program, while it is still applying the unit’s settings. A few that come up often:
| Status | Name | Usually means |
|---|---|---|
200 |
CHDIR |
WorkingDirectory= does not exist or is not accessible |
203 |
EXEC |
The executable is missing, not executable, or has a bad interpreter line |
217 |
USER |
The User= account does not exist |
226 |
NAMESPACE |
A sandboxing option could not set up its mount namespace, often a missing path in ReadWritePaths= |
You do not need to memorise the rest. systemd-analyze exit-status prints the full table for the version you are running, and the journal line before the exit usually names the failed step, as it did above.
Read the journal narrowly
Avoid beginning with an unbounded log dump. Query the unit and the current boot first:
journalctl -u example.service -b --no-pagerThen reduce the time window if the service is noisy. The goal is to correlate the unit transition with the application’s own output. A few options do most of the narrowing:
--since "-15min"or--since "09:40" --until "09:45"bound the time window.-p warningdrops everything below warning priority, for programs that log with priorities.-o catstrips the timestamp and hostname prefix, which helps when you want to read the messages rather than the metadata.-o verbosedoes the opposite and shows every field attached to a message, including_PID,_EXEand_SYSTEMD_INVOCATION_ID.
That last field is the most precise filter there is. Every run of a unit gets a new invocation ID, so you can ask for the output of exactly the most recent run, without the noise of the previous dozen restarts:
id=$(systemctl show -p InvocationID --value example.service)
journalctl _SYSTEMD_INVOCATION_ID="$id" --no-pagerRestarts hide the first failure
Restart=on-failure is good for production and bad for diagnosis. The unit you are looking at may have failed, restarted, failed differently and restarted again, and status only shows the latest run.
Check NRestarts from the show command above. If it is non-zero, go back through the journal for the first failure in the sequence; later ones are often consequences of it, such as a port still bound or a lock file left behind.
If the unit restarts more than StartLimitBurst= times within StartLimitIntervalSec= (by default, 5 times in 10 seconds), systemd stops trying and the result becomes start-limit-hit. That message describes the restart policy, not the bug. Once you have fixed the cause, clear the counter before starting it again:
systemctl reset-failed example.service
systemctl start example.serviceTimeouts are usually about readiness
A timeout result rarely means the program is slow. It usually means systemd and the program disagree about when the service is ready, and the 90-second default TimeoutStartSec= ran out while systemd waited for a signal that never came.
What systemd waits for depends on Type=:
simpleandexec: nothing beyond the process starting (execalso waits for theexecve()to succeed). These rarely time out on start.forking: the original process to exit, leaving a daemon behind. A program that does not fork under this type will time out.notify: an explicitREADY=1message oversd_notify(). A program that never sends it will time out, however healthy it is.
If a service works fine when started by hand but times out under systemd, check that Type= matches what the program actually does before you raise any timeout.
Reproduce the runtime context
A command that works in your interactive shell can still fail as a service. Check the details the shell normally hides:
- user and group
- working directory
- environment variables, including
PATH; nothing from your.bashrcis present - file permissions
- mount and namespace restrictions
- capabilities
systemctl cat example.service shows the effective unit definition, including drop-ins. systemctl show exposes the values after systemd has resolved defaults and overrides:
systemctl show example.service \
-p User -p Group -p WorkingDirectory -p Environment \
-p ProtectSystem -p ProtectHome -p ReadWritePaths -p PrivateTmpSandboxing options are the classic case. Here the service can read its configuration but not write its state:
root@web-01:~$ journalctl -u example.service -b -n 3 --no-pager -o catStarting example.service - Example service...server: open /var/lib/example/state.db: read-only file systemexample.service: Main process exited, code=exited, status=1/FAILURE$ systemctl show example.service -p ProtectSystem -p ReadWritePaths -p StateDirectoryProtectSystem=strictReadWritePaths=StateDirectory=
ProtectSystem=strict mounts the entire filesystem read-only for the service, apart from paths the unit explicitly grants. /var/lib/example is writable for you as root and read-only for the service, which is exactly the “works when I run it” trap. The idiomatic fix is StateDirectory=example, which creates /var/lib/example, gives it to the service’s user, and makes it writable inside the sandbox.
Run a shell inside the unit’s sandbox
Rather than reading properties and imagining their combined effect, you can create a transient unit with the same settings and get a shell inside it:
systemd-run --pty --wait --collect \
-p User=example \
-p WorkingDirectory=/var/lib/example \
-p ProtectSystem=strict \
-p ProtectHome=yes \
-p PrivateTmp=yes \
/bin/bashFrom there, touch, ls -l, id and env answer questions that would otherwise take several rounds of editing and restarting. Copy the properties from systemctl cat, and add or remove them one by one until the shell behaves like the service.
systemd-analyze security example.service is a useful companion: it lists every sandboxing option and whether the unit uses it. It is meant for hardening, but it is also a quick inventory of everything that might be restricting the process.
Check what the unit depends on
Sometimes the service never ran at all. If something it Requires= failed, the journal says Dependency failed for Example service and the service’s own logs are empty, because there was nothing to log.
systemctl list-dependencies example.service
systemctl --failedFix the failed dependency first. A common version of this is a service ordered After=network-online.target on a host where nothing pulls in a wait-online service, or a mount unit that failed because a disk was missing.
Change one variable at a time
Once you have a concrete failure mode, make the smallest change that tests the hypothesis. That keeps the debugging loop legible and makes the eventual fix easier to explain in version control.
For example, to test whether more verbose logging reveals the problem, add a drop-in instead of editing the packaged unit:
[Service]
Environment=LOG_LEVEL=debugIf you edit unit files directly instead, run systemctl daemon-reload afterwards. Otherwise systemd keeps using the old definition and systemctl status warns that the unit file changed on disk. systemd-analyze verify example.service catches typos and unknown directives before you restart anything.
The whole path, in order
systemctl show -p Result -p ExecMainCode -p ExecMainStatus -p NRestartsto learn how it failed.- A status of 200 or more means systemd failed before your program ran; read the step name in the journal.
journalctl -u … -b, narrowed by time or invocation ID, for the program’s own account.- If it restarted, find the first failure, not the last.
- On a timeout, check that
Type=matches how the program signals readiness. - Compare the runtime context with your shell’s, or use
systemd-runto step into it. - Test one change at a time with a drop-in.
Most service failures are decided in the first two steps. The rest exist for the ones that are not, and they work just as well at 3 a.m.