Linux load average is not CPU usage
Why a high load average can point at CPU pressure, blocked I/O, or scheduling pressure—and what to inspect next.
On this page
Load average is easy to recognise and easy to misread. It is the first number in uptime, top and every monitoring dashboard ever built, and it is regularly treated as “how busy the CPUs are”. It is not a percentage, and on Linux it is not even purely a CPU metric.
This post covers what the kernel actually counts, why Linux differs from other Unix systems, and a short sequence of commands that turns “load is high” into “this resource is the bottleneck”.
What the number counts
Every five seconds the kernel samples the number of tasks that are either:
- runnable: running on a CPU, or queued waiting for one, or
- in uninterruptible sleep: blocked in the kernel in state
D, usually waiting for I/O.
It then folds that sample into three exponentially damped moving averages, the familiar 1, 5 and 15 minute figures. You can read them directly:
cat /proc/loadavg23.41 19.87 12.06 3/1287 482113The first three fields are the averages. The fourth is the number of runnable scheduling entities right now over the total number on the system, and the last is the most recently allocated PID. Note that 3 in the fourth field: only three tasks were runnable at that instant, yet the 1-minute load is 23. The difference is the second category, and it is the whole subject of this post.
“1 minute” is a time constant, not a window
The averages are not the mean over the last minute. They are exponentially weighted, computed in fixed point with constants the kernel hard-codes:
#define FSHIFT 11 /* nr of bits of precision */
#define FIXED_1 (1<<FSHIFT) /* 1.0 as fixed-point */
#define LOAD_FREQ (5*HZ+1) /* 5 sec intervals */
#define EXP_1 1884 /* 1/exp(5sec/1min) as fixed-point */
#define EXP_5 2014 /* 1/exp(5sec/5min) */
#define EXP_15 2037 /* 1/exp(5sec/15min) */Each sample, the old value decays by e^(-5/60) and the new count contributes the rest. Two consequences follow:
- After a step change, the 1-minute figure reaches only about 63% of the new level after one minute. A machine that went from idle to a constant 16 runnable tasks shows roughly 10, not 16, sixty seconds later.
- Events older than the window never fully disappear; they fade. A 15-minute value that is still climbing while the 1-minute value falls means the burst is over and the long average is catching up.
Comparing the three numbers tells you the direction of the load, which is often more useful than the magnitude.
Threads, not processes
The kernel counts tasks, and every thread is a task. A single Java or Go process with 200 threads blocked on a slow NFS mount contributes 200 to the load, even though ps shows one line for it.
Containers see the host
/proc/loadavg is not namespaced. Inside a container you see the host’s load average, including every other tenant on the node. If you need per-workload numbers, use cgroup pressure metrics (below) or a tool like LXCFS that virtualises /proc files.
Why Linux counts blocked tasks
On most Unix systems, including FreeBSD and illumos, load average measures CPU demand alone. Linux has counted tasks in uninterruptible sleep since 1993, when a patch was added so that a system thrashing on swap would show a high load rather than looking idle. The reasoning holds: a task waiting on a disk is demand that the system is failing to serve.
The cost is ambiguity. A load of 20 on Linux can mean twenty threads fighting for CPUs, twenty threads waiting on a slow disk, or twenty threads stuck behind a kernel lock that has nothing to do with either. Brendan Gregg’s 2017 write-up, Linux Load Averages: Solving the Mystery, traces the history and is worth reading in full.
Put the number in context
A load of 4 means something different on a two-core machine than on a 64-core host. Start with the shape of the system:
root@db-02:~$ nproc16$ uptime14:02:11 up 41 days, 3:12, 2 users, load average: 23.41, 19.87, 12.06
Load divided by core count gives a rough sense of saturation. Here it is about 1.5 per core and rising, since the 1-minute figure is above the 15-minute one. On its own, though, this says nothing about which resource is short. The next step is to split the number back into its two halves.
Split runnable from blocked
vmstat reports both halves of the load directly, as the r and b columns:
root@db-02:~$ vmstat 1 4procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----r b swpd free buff cache si so bi bo in cs us sy id wa st3 1 0 845120 10212 9918204 0 0 14 210 2 5 6 2 91 1 02 21 0 812344 10212 9921880 0 0 1210 48211 4120 9810 4 3 41 52 01 22 0 810112 10212 9922004 0 0 980 51004 3988 9502 3 3 40 54 03 20 0 809876 10212 9922110 0 0 1102 49876 4051 9688 4 2 42 52 0
Ignore the first row; it is the average since boot. The rest are one-second samples.
ris runnable tasks, including the ones currently on a CPU. Only two or three here, on a 16-core box.bis tasks blocked in uninterruptible sleep. Twenty-odd, which accounts for nearly all of the load.idis around 40% idle CPU, andwa(time a CPU sat idle while a task waited on I/O) is above 50%.boshows roughly 50 MB/s of writes going out.
This is not a CPU problem. Adding cores would not change anything; the machine is waiting on storage.
For contrast, a CPU-bound system looks like this: r well above nproc, b near zero, id near zero, and most time in us or sy. And on a virtual machine, watch st (steal): if it is high, your runnable tasks are waiting for the hypervisor to schedule your vCPUs, and the fix is on the host side.
| Pattern | r |
b |
CPU columns | Likely cause |
|---|---|---|---|---|
| CPU saturation | above nproc |
low | id ≈ 0, high us/sy |
Not enough CPU for the work |
| I/O stall | low | high | high wa, id not zero |
Storage latency or throughput |
| Lock or NFS hang | low | high | idle, low wa, disks quiet |
Tasks blocked on something other than local disk |
| Steal | above nproc |
low | high st |
Hypervisor contention |
Find the blocked tasks
Once b is the suspect, find out who is blocked and where. The wchan column shows the kernel function a sleeping task is waiting in:
root@db-02:~$ ps -eo state,pid,wchan:20,comm | awk '$1 == "D"' | head -5D 4127 io_schedule postgresD 4131 io_schedule postgresD 4133 io_schedule postgresD 4140 io_schedule postgresD 9902 rq_qos_wait rsync
io_schedule is the plain “waiting for a block device” case. rq_qos_wait means the block layer is throttling writeback, which fits a large rsync filling the device queue while the database waits behind it. If the wait channel names something like an NFS or FUSE function, or a mutex, the disk is probably innocent.
For the full kernel stack of one task, as root:
cat /proc/4127/stackThen confirm what the device itself is doing:
iostat -x 1Look at r_await and w_await (average milliseconds per request) and aqu-sz (queue depth) for the device holding the data. Be careful with %util: on SSDs and NVMe devices, which serve many requests in parallel, 100% utilisation only means the device always had something in flight, not that it is at capacity.
Ask the kernel about pressure directly
Load average is an indirect measurement. On kernels with pressure stall information (PSI, Linux 4.20 and later), the kernel reports directly how much time tasks lost waiting for each resource:
cat /proc/pressure/iosome avg10=71.32 avg60=64.08 avg300=38.51 total=912388120
full avg10=58.90 avg60=52.17 avg300=30.02 total=741203391someis the share of time at least one task was stalled on I/O.fullis the share of time all non-idle tasks were stalled at once, which is time the machine got no useful work done.avg10,avg60andavg300are percentages over 10 seconds, 1 minute and 5 minutes;totalis cumulative microseconds.
The same format exists for cpu and memory. A 52% full figure for I/O over the last minute is far less ambiguous than a load of 19.87.
A short checklist
When an alert fires on load average:
- Compare the three averages to see whether load is rising or falling.
- Divide by
nprocto see whether the number is even large. - Run
vmstat 1and decide whetherrorbexplains it. - If
r: find the CPU consumers withtoporpidstat -u 1, and checkston VMs. - If
b: listDtasks with theirwchan, then checkiostat -xand/proc/pressure/io. - Alert on PSI instead next time.
The useful question is never “is load high?” but “what work is runnable or blocked, and what resource is keeping it there?” Load average tells you to ask it. The other tools answer it.