cd ../writing

Linux load average is not CPU usage

Why a high load average can point at CPU pressure, blocked I/O, or scheduling pressure—and what to inspect next.

On this page
  1. What the number counts
  2. “1 minute” is a time constant, not a window
  3. Threads, not processes
  4. Containers see the host
  5. Why Linux counts blocked tasks
  6. Put the number in context
  7. Split runnable from blocked
  8. Find the blocked tasks
  9. Ask the kernel about pressure directly
  10. A short checklist

Load average is easy to recognise and easy to misread. It is the first number in uptime, top and every monitoring dashboard ever built, and it is regularly treated as “how busy the CPUs are”. It is not a percentage, and on Linux it is not even purely a CPU metric.

This post covers what the kernel actually counts, why Linux differs from other Unix systems, and a short sequence of commands that turns “load is high” into “this resource is the bottleneck”.

What the number counts

Every five seconds the kernel samples the number of tasks that are either:

  • runnable: running on a CPU, or queued waiting for one, or
  • in uninterruptible sleep: blocked in the kernel in state D, usually waiting for I/O.

It then folds that sample into three exponentially damped moving averages, the familiar 1, 5 and 15 minute figures. You can read them directly:

bash
cat /proc/loadavg
text
23.41 19.87 12.06 3/1287 482113

The first three fields are the averages. The fourth is the number of runnable scheduling entities right now over the total number on the system, and the last is the most recently allocated PID. Note that 3 in the fourth field: only three tasks were runnable at that instant, yet the 1-minute load is 23. The difference is the second category, and it is the whole subject of this post.

“1 minute” is a time constant, not a window

The averages are not the mean over the last minute. They are exponentially weighted, computed in fixed point with constants the kernel hard-codes:

include/linux/sched/loadavg.hc
#define FSHIFT    11              /* nr of bits of precision */
#define FIXED_1   (1<<FSHIFT)     /* 1.0 as fixed-point */
#define LOAD_FREQ (5*HZ+1)        /* 5 sec intervals */
#define EXP_1     1884            /* 1/exp(5sec/1min) as fixed-point */
#define EXP_5     2014            /* 1/exp(5sec/5min) */
#define EXP_15    2037            /* 1/exp(5sec/15min) */

Each sample, the old value decays by e^(-5/60) and the new count contributes the rest. Two consequences follow:

  • After a step change, the 1-minute figure reaches only about 63% of the new level after one minute. A machine that went from idle to a constant 16 runnable tasks shows roughly 10, not 16, sixty seconds later.
  • Events older than the window never fully disappear; they fade. A 15-minute value that is still climbing while the 1-minute value falls means the burst is over and the long average is catching up.

Comparing the three numbers tells you the direction of the load, which is often more useful than the magnitude.

Threads, not processes

The kernel counts tasks, and every thread is a task. A single Java or Go process with 200 threads blocked on a slow NFS mount contributes 200 to the load, even though ps shows one line for it.

Containers see the host

/proc/loadavg is not namespaced. Inside a container you see the host’s load average, including every other tenant on the node. If you need per-workload numbers, use cgroup pressure metrics (below) or a tool like LXCFS that virtualises /proc files.

Why Linux counts blocked tasks

On most Unix systems, including FreeBSD and illumos, load average measures CPU demand alone. Linux has counted tasks in uninterruptible sleep since 1993, when a patch was added so that a system thrashing on swap would show a high load rather than looking idle. The reasoning holds: a task waiting on a disk is demand that the system is failing to serve.

The cost is ambiguity. A load of 20 on Linux can mean twenty threads fighting for CPUs, twenty threads waiting on a slow disk, or twenty threads stuck behind a kernel lock that has nothing to do with either. Brendan Gregg’s 2017 write-up, Linux Load Averages: Solving the Mystery, traces the history and is worth reading in full.

Put the number in context

A load of 4 means something different on a two-core machine than on a 64-core host. Start with the shape of the system:

root@db-02:~
$ nproc16$ uptime14:02:11 up 41 days,  3:12,  2 users,  load average: 23.41, 19.87, 12.06

Load divided by core count gives a rough sense of saturation. Here it is about 1.5 per core and rising, since the 1-minute figure is above the 15-minute one. On its own, though, this says nothing about which resource is short. The next step is to split the number back into its two halves.

Split runnable from blocked

vmstat reports both halves of the load directly, as the r and b columns:

root@db-02:~
$ vmstat 1 4procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st3  1      0 845120  10212 9918204    0    0    14   210    2    5  6  2 91  1  02 21      0 812344  10212 9921880    0    0  1210 48211 4120 9810  4  3 41 52  01 22      0 810112  10212 9922004    0    0   980 51004 3988 9502  3  3 40 54  03 20      0 809876  10212 9922110    0    0  1102 49876 4051 9688  4  2 42 52  0

Ignore the first row; it is the average since boot. The rest are one-second samples.

  • r is runnable tasks, including the ones currently on a CPU. Only two or three here, on a 16-core box.
  • b is tasks blocked in uninterruptible sleep. Twenty-odd, which accounts for nearly all of the load.
  • id is around 40% idle CPU, and wa (time a CPU sat idle while a task waited on I/O) is above 50%.
  • bo shows roughly 50 MB/s of writes going out.

This is not a CPU problem. Adding cores would not change anything; the machine is waiting on storage.

For contrast, a CPU-bound system looks like this: r well above nproc, b near zero, id near zero, and most time in us or sy. And on a virtual machine, watch st (steal): if it is high, your runnable tasks are waiting for the hypervisor to schedule your vCPUs, and the fix is on the host side.

Pattern r b CPU columns Likely cause
CPU saturation above nproc low id ≈ 0, high us/sy Not enough CPU for the work
I/O stall low high high wa, id not zero Storage latency or throughput
Lock or NFS hang low high idle, low wa, disks quiet Tasks blocked on something other than local disk
Steal above nproc low high st Hypervisor contention

Find the blocked tasks

Once b is the suspect, find out who is blocked and where. The wchan column shows the kernel function a sleeping task is waiting in:

root@db-02:~
$ ps -eo state,pid,wchan:20,comm | awk '$1 == "D"' | head -5D    4127 io_schedule          postgresD    4131 io_schedule          postgresD    4133 io_schedule          postgresD    4140 io_schedule          postgresD    9902 rq_qos_wait          rsync

io_schedule is the plain “waiting for a block device” case. rq_qos_wait means the block layer is throttling writeback, which fits a large rsync filling the device queue while the database waits behind it. If the wait channel names something like an NFS or FUSE function, or a mutex, the disk is probably innocent.

For the full kernel stack of one task, as root:

bash
cat /proc/4127/stack

Then confirm what the device itself is doing:

bash
iostat -x 1

Look at r_await and w_await (average milliseconds per request) and aqu-sz (queue depth) for the device holding the data. Be careful with %util: on SSDs and NVMe devices, which serve many requests in parallel, 100% utilisation only means the device always had something in flight, not that it is at capacity.

Ask the kernel about pressure directly

Load average is an indirect measurement. On kernels with pressure stall information (PSI, Linux 4.20 and later), the kernel reports directly how much time tasks lost waiting for each resource:

bash
cat /proc/pressure/io
text
some avg10=71.32 avg60=64.08 avg300=38.51 total=912388120
full avg10=58.90 avg60=52.17 avg300=30.02 total=741203391
  • some is the share of time at least one task was stalled on I/O.
  • full is the share of time all non-idle tasks were stalled at once, which is time the machine got no useful work done.
  • avg10, avg60 and avg300 are percentages over 10 seconds, 1 minute and 5 minutes; total is cumulative microseconds.

The same format exists for cpu and memory. A 52% full figure for I/O over the last minute is far less ambiguous than a load of 19.87.

A short checklist

When an alert fires on load average:

  1. Compare the three averages to see whether load is rising or falling.
  2. Divide by nproc to see whether the number is even large.
  3. Run vmstat 1 and decide whether r or b explains it.
  4. If r: find the CPU consumers with top or pidstat -u 1, and check st on VMs.
  5. If b: list D tasks with their wchan, then check iostat -x and /proc/pressure/io.
  6. Alert on PSI instead next time.

The useful question is never “is load high?” but “what work is runnable or blocked, and what resource is keeping it there?” Load average tells you to ask it. The other tools answer it.