1. Stop grep matching itself without grep -v grep

    ps aux | grep nginx always lists the grep nginx process too, which is why | grep -v grep is so common. Wrap one character of the pattern in brackets instead:

    bash
    ps aux | grep '[n]ginx'

    The regex [n]ginx still matches nginx, but grep’s own command line now contains the literal text [n]ginx, which the regex doesn’t match. Keep the quotes: unquoted, the shell treats [n]ginx as a glob, and if a file called nginx exists in the current directory, grep receives nginx and matches itself again.

    Usually a dedicated tool is simpler:

    bash
    pgrep -a nginx            # PID + command line; never matches itself
    pgrep -af 'nginx: worker' # match the full command line, not just the name
    pidof nginx               # exact program name, PIDs only
    ps -C nginx -o pid,args   # ps output columns, filtered by name

    pgrep matches the process name (truncated to 15 characters) unless you pass -f. With -f, and with the grep trick, any process whose command line contains the pattern matches too — a wrapping bash -c '...' or ssh host '...' that mentions nginx will show up alongside the real process.

  2. ldd may run the binary you are inspecting

    ldd is not a parser. It is a shell script that asks the dynamic loader to trace itself — normally by running the loader with --list, and in some versions and some situations by setting LD_TRACE_LOADED_OBJECTS=1 and executing the file itself. Whatever the loader or the file does on the way there happens with your privileges: constructors, LD_AUDIT and LD_PRELOAD handlers, or the entire program. ldd(1) states the rule plainly — never use it on an untrusted executable, because it may result in the execution of arbitrary code — and adds that upstream ldd did precisely that direct-exec trick before glibc 2.27, though most distributions shipped a version that did not.

    The library list is already in the file. Read it:

    bash
    objdump -p ./downloaded | grep NEEDED   # what the man page recommends

    The same metadata, with the context you usually want next to it — RUNPATH, RPATH, SONAME, and the symbol versions a binary insists on:

    bash
    readelf -d ./downloaded
    patchelf --print-needed ./downloaded
    patchelf --print-interpreter ./downloaded   # the PT_INTERP loader the file expects

    Two caveats the manual page flags alongside its own recommendation. NEEDED lists direct dependencies only, where ldd prints the whole resolved tree; and a statically linked binary has no NEEDED entries at all, which is itself the answer if you are checking whether something carries its own libc.

    Turning sonames into paths does not require running anything. Ask the loader cache:

    bash
    ldconfig -p | grep -F 'libssl.so.3'

    For the full tree, lddtree from pax-utils walks the dependencies by reading ELF files on disk and explicitly never executes or loads the code:

    bash
    lddtree ./downloaded

    That answers “what does this need”. If the question is what it does while loading, no static tool will answer it, and the answer is to run it somewhere it cannot reach:

    bash
    unshare --user --map-root-user --net --mount ./downloaded

    An empty network namespace and a private mount table contain most of the damage, but that is blast-radius reduction, not a sandbox. Against code you genuinely distrust, use a throwaway VM. Against code you ship, ldd is fine and still the friendliest tool — the objection is only ever about trust.

  3. MemoryHigh is a slope, MemoryMax is a cliff

    A container that runs out of memory did not run out of machine memory — it hit a limit recorded under /sys/fs/cgroup, which is why free -h inside the container shows the host’s total and tells you nothing. The cgroup files are the only place the real numbers live:

    bash
    cg=/sys/fs/cgroup/system.slice/example.service
    cat "$cg"/memory.current "$cg"/memory.peak
    cat "$cg"/memory.high "$cg"/memory.max "$cg"/memory.swap.max

    The two limits behave differently, and conflating them is the source of most bad advice in this area:

    • memory.high is a throttle. Over it, the kernel reclaims and stalls the tasks in that cgroup. It never invokes the OOM killer, and the limit can still be breached under extreme pressure. It is the early-warning system.
    • memory.max is a hard ceiling. When reclaim cannot free enough, the memcg OOM killer runs — and it chooses only among processes inside that cgroup, which is why a capped container kills its own worker instead of destabilising the host.

    memory.peak records the highest memory.current the cgroup has reached, which turns “what did it actually peak at?” from an archaeology exercise into a file read. memory.events is the counter set to alert on:

    bash
    cat "$cg"/memory.events
    # low 0
    # high 1284
    # max 6
    # oom 6
    # oom_kill 2

    These values are cumulative, so read twice and diff. A rising high counter means throttling — latency, not an outage, and expected when you set a meaningful high limit. A rising max or oom counter means the ceiling was reached. A rising oom_kill counter is an incident, and on a healthy service it never moves.

    systemd maps the same knobs onto unit directives: MemoryHigh=, MemoryMax=, MemorySwapMax=, plus OOMPolicy= for what happens to the rest of the unit once one of its processes is killed, and ManagedOOMMemoryPressure= to let systemd-oomd kill a cgroup on sustained pressure before it reaches MemoryMax at all. For a pool of workers that share state, memory.oom.group=1 makes the kernel take the entire group down rather than leaving half the pool alive with its session gone.

    Set high well below max and alert on the counters between them, not on utilisation:

    bash
    systemd-cgtop --order=memory --depth=3
    systemctl show example.service -p MemoryCurrent -p MemoryPeak -p MemoryHigh -p MemoryMax

    Two limits worth keeping in mind: memory.current includes page cache, which is cheap to reclaim, so it is anonymous growth that actually kills you; and on cgroup v1 hosts the equivalents are memory.limit_in_bytes, memory.soft_limit_in_bytes, memory.failcnt and memory.oom_control.

  4. Core dumps are free until you need one

    The kernel refuses to write core dumps by default (RLIMIT_CORE starts at 0), which is why production crashes usually leave nothing behind. systemd raises that limit for everything it starts and registers itself as the kernel’s kernel.core_pattern handler, so on a systemd host dumps already work — they just land somewhere nobody looks:

    bash
    sysctl kernel.core_pattern          # |/usr/lib/systemd/systemd-coredump
    sysctl kernel.core_ulimit
    systemctl show example.service -p LimitCORE

    systemd-coredump logs the signal, command line, cgroup and a backtrace to the journal, then stores a compressed core under /var/lib/systemd/coredump/, aged out by /usr/lib/tmpfiles.d/systemd.conf. coredumpctl is the interface to all of it:

    bash
    coredumpctl list -1
    coredumpctl info 1234              # or match by command name or executable path
    coredumpctl debug 1234             # opens gdb on the core, with the right paths
    coredumpctl -o core.1234 dump 1234 # the raw core file

    Matching works on PID, COMM or EXE, plus --since. An unprivileged user only sees their own crashes, because the metadata lives in the journal — a real reason for a reporting path that runs privileged.

    Two things decide whether a dump is useful when you finally want it. Symbols: a backtrace of bare addresses means the debuginfo for that exact build is not on the host, so install the matching -dbg/debuginfo package, point DEBUGINFOD_URLS at a debuginfod server, or check which build you are dealing with via eu-unstrip -n core. Size: a multi-gigabyte JVM core is a disk-full event, so /etc/systemd/coredump.conf is where you bound it — ProcessSizeMax=, ExternalSizeMax=, MaxUse=, KeepFree= — or set Storage=none to keep the log line and discard the file.

    For a process that is alive but wedged, do not make it crash to get its memory:

    bash
    gcore -o /tmp/example 1234
    gdb -p 1234 -batch -ex 'thread apply all bt' -ex 'gcore /tmp/core'

    Managed runtimes have a better answer still: -XX:+HeapDumpOnOutOfMemoryError for the JVM, GOTRACEBACK=crash to turn a Go panic into a SIGABRT and a core, faulthandler.enable to print a Python traceback instead of a bare segfault. Containers need the plumbing — bind the host’s /run/systemd/coredump into the container, or run with --ulimit core=-1 and a host-side handler — because without one a containerised crash vanishes at the container boundary.

  5. Give AI agents privileged access with pkexec

    AI coding agents usually run unprivileged and stall when they need root — installing a package, writing under /etc, restarting a service. Instead of running the whole agent as root, tell it to escalate per command with pkexec:

    text
    When a command needs root, run it as `pkexec <command>` instead of failing or trying sudo.

    pkexec goes through Polkit, so each escalation gets its own auth prompt (desktop dialog or polkit agent), the environment is sanitized, and the action is logged. Example:

    bash
    pkexec systemctl restart nginx

    Caveat: it needs a working polkit agent — on a headless SSH session with no agent it will fail, where sudo is still the right tool.

  6. Find and kill the process using a TCP port

    When a port is already in use, find who owns it before reaching for a reboot:

    bash
    ss -ltnp 'sport = :8080'
    lsof -i :8080
    fuser 8080/tcp

    ss -ltnp shows the PID in the users:(...) column (use sudo to see other users’ processes); lsof -i :8080 lists the open sockets; fuser 8080/tcp prints just the PIDs. Then kill it:

    bash
    fuser -k 8080/tcp
    # or: kill $(lsof -t -i :8080)

    -k sends SIGKILL by default — use fuser -k -TERM 8080/tcp first if you want to let the process clean up.

  7. perf counts, perf samples — pick the right one

    perf has two modes and they answer different questions. Decide which one you are asking before you collect anything.

    Count when the question is how much: cycles, instructions, cache misses, context switches, page faults over a window long enough to mean something.

    bash
    perf stat -p "$(pidof example)" sleep 30

    Hardware counters are scarce, so when you request more events than the CPU has slots, perf multiplexes them and scales the numbers. The run reports the fraction of time the counters were actually running; below roughly 90% on a short window, treat the result as indicative only. Long window, few events, and the question never arises.

    Sample when the question is where the time went:

    bash
    perf record -F 99 -a -g --call-graph dwarf -o /tmp/prof.data -- sleep 30
    perf report --stdio --sort comm,dso,symbol --percent-limit 0.5
    perf report --stdio -g graph

    -F 99 is the conventional fixed sample rate; the default for cycles is adaptive and throttles itself, which quietly distorts a thirty-second run. --call-graph decides whether your stacks are usable at all: fp (the default) needs frame pointers, so a binary built with -fno-omit-frame-pointer yields truncated stacks; dwarf unwinds through .eh_frame at a real CPU cost and needs nothing from the compiler, which makes it the sane default on x86-64; lbr uses the branch recorder, and is x86-only and approximate.

    A browser is optional. perf report --stdio --percent-limit 1 finds the top consumers on its own, and perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg (Brendan Gregg’s FlameGraph) turns the same file into a flame graph when you want the shape rather than the ranking. perf top -p PID is the live version of the same idea.

    Two things stop this working, and both have obvious fixes:

    • Permissions. cat /proc/sys/kernel/perf_event_paranoid — distributions set it tightly on purpose. Without hardware access a software event still profiles userspace: perf record -e cpu-clock -a -g --all-user. Where the value blocks even that, you need root or a sysctl change, which is the host owner’s decision rather than yours to make in a shell.
    • Symbols. Frames shown as [unknown] or as raw addresses mean the debuginfo for that build is missing. Install the matching debuginfo packages, set DEBUGINFOD_URLS so build-ids resolve over the network, or rebuild with -g.

    Thirty seconds of samples is noise; aim for thousands of samples. To hand the profile to someone else, perf archive bundles it with the binaries and debuginfo they need to open it.

  8. Deleted files can still fill a disk

    When df says a filesystem is full but du can’t find the space, a process is usually holding a deleted file open. The blocks aren’t freed until the last file descriptor closes.

    bash
    lsof +L1

    +L1 lists open files with a link count below one, which means unlinked but still open. Restart the process, or truncate the file through its descriptor if you can’t:

    bash
    : > /proc/<pid>/fd/<fd>
  9. Ask the package manager what changed on disk

    Every package manager already stores checksums for the files it installed, which turns “has anything been modified since installation?” into one command instead of a filesystem-wide hunt:

    bash
    dpkg --verify               # Debian, Ubuntu
    rpm -Va                      # Fedora, RHEL
    pacman -Qkk                  # Arch

    All three speak the same output dialect: a field of check results, . where a check passed and ? where it could not be performed, a letter where one failed — 5 for a digest mismatch, M for mode, S for size, T for mtime — then c if the file is a conffile.

    text
    ??5?????? c /etc/ssh/sshd_config
    S.5....T.  c /etc/php.ini

    The first line is dpkg being precise about itself: it functionally checks digests and nothing else, so every other position is ?. It is also a conffile, the one file in that directory you expect to differ because somebody edited it on purpose. Skip conffiles first, and only care about T when nothing else changed. rpm compares size, mode, owner and mtime too, so rpm -V --nomtime --nosize narrows it to the digest comparison, which is usually what you want when triaging.

    None of them can find files that nobody owns. rpm -Va walks the package database, so a new binary dropped in /usr/local/bin is invisible to it; dnf repocheck --unowned-files walks the filesystem instead and reports those separately. That direction of travel is the one that finds attacker-planted files.

    Two limits worth stating plainly. pacman -Qkk checks against the package’s mtree file, which records size, mode and mtime but no checksum — a file patched in place with its length and timestamp restored passes. And all of these compare against the local database, so an attacker with root can rewrite a binary together with the database entry that vouches for it and none of this notices. That gap is what file-integrity monitoring with an off-host baseline belongs to: AIDE, Tripwire, or a scheduled rpm -Va whose output you ship somewhere else, alongside package signature verification. Either way, run it on a timer — the signal is the diff from yesterday, not today’s output.

  10. Measure durations with the monotonic clock

    Linux has several clocks, and mixing them up is a reliable way to ship a bug that only appears in production. CLOCK_REALTIME is the wall clock: NTP, settimeofday and leap seconds move it forwards or backwards. CLOCK_MONOTONIC counts from boot and is never stepped, though it does follow frequency adjustments and it stops during suspend — CLOCK_BOOTTIME includes the suspend time, and CLOCK_MONOTONIC_RAW skips adjustment entirely.

    The rule that keeps this simple: intervals, timeouts, backoff and deadlines use CLOCK_MONOTONIC; anything a human or a peer will read — log timestamps, database rows, TLS validity, leases — uses CLOCK_REALTIME, and always through one helper rather than at each call site. Most languages already default to the right thing: time.monotonic() in Python, Instant in Rust, time.Since on a value from Go’s time.Now() are monotonic, while date(1), datetime.now() and SystemTime::now() are not.

    The failure always has the same shape. A sync step rewinds the clock, code that stored a wall-clock timestamp subtracts it later, and the result is a negative interval: a retry loop that backs off further into the past, or an expiry that never fires. Journald keeps both, which is how you investigate it afterwards:

    bash
    journalctl -u example.service -o short-monotonic --since "-30min"
    journalctl -o json | jq -r '.__MONOTONIC_TIMESTAMP'

    Steps and slews also have different operational consequences. systemd-timesyncd steps the clock by default; chrony and ntpd slew, and the kernel caps slewing at 500 ppm, so a clock a second out takes about half an hour to walk back. That is the right behaviour for a busy machine and the wrong one for a failover pair, because a step hits every node simultaneously and two peers can briefly both believe they are primary. And if you ever need the wall clock virtualised: time namespaces offset CLOCK_MONOTONIC and CLOCK_BOOTTIME only, so unshare --time will not fake a timezone and a container’s CLOCK_REALTIME is the host’s.

    CLOCK_REALTIME_COARSE is the cheap one — wall-clock accurate to the timer tick, cached in a page the kernel keeps updated for you — and it is the right clock for stamping millions of events where microseconds are noise. Finally, timedatectl on a dual-boot box is worth reading once: another operating system will happily leave the RTC in local time for the Linux side to misread.

  11. Why is boot slow? Ask for the critical chain

    systemd-analyze blame sorts units by how long they took to start, but units start in parallel, so the slowest one often isn’t what delayed boot.

    bash
    systemd-analyze critical-chain

    This shows the chain of units that actually gated default.target. The time after @ is when a unit became active; the time after + is how long it took to start.

  12. Print a value in reusable shell quoting

    Bash 4.4 added the @Q parameter transformation, which quotes a value so it can be pasted back into a shell unchanged:

    bash
    var="it's here"
    echo "${var@Q}"   # 'it'\''s here'

    Handy in logs: echo "running: ${cmd[*]@Q}" shows exactly which arguments a command received.