Inside the Go Scheduler: G, M, P, Work Stealing, and Why Goroutines Block Cheaply

Every Go program you have ever shipped is a swarm of goroutines multiplexed onto a handful of operating system threads. The machinery that decides which goroutine runs where, and for how long, is the Go runtime scheduler. Most of the time it is invisible. Then one day a tail latency spike appears under load, a CPU-bound goroutine hogs a core, or a Kubernetes pod ignores its CPU limit, and the scheduler stops being an implementation detail.

This post walks through how the scheduler actually works: the G-M-P model, work stealing, preemption, and how goroutine stacks make cheap blocking possible. Everything here is verifiable in the runtime source, and the mechanics have been stable for years — which is exactly why understanding them pays off every time you debug concurrency behavior.

G, M, and P: The Three Actors

The runtime models concurrency with three types. A G is a goroutine: it carries its own stack, a small state machine (_Grunnable, _Grunning, _Gwaiting, and friends), and the function it will execute. An M is an OS thread — the thing that actually executes machine code. A P is a logical processor, a scheduling context that an M must hold before it can run Go code.

The P is the part people skip over, and it is the key to the whole design. The number of P’s equals GOMAXPROCS, which defaults to the number of logical CPUs — and since Go 1.25, on Linux the runtime also respects the cgroup CPU bandwidth limit if it is lower, updating the value periodically as limits change. This is a real fix for a decade-old container problem: before that change, a pod limited to one CPU core would still see 64 logical CPUs and default to 64 P’s, leaving dozens of threads contending for throttled quota. You can read the full behavior in the Go 1.25 release notes.

Why not just map goroutines directly onto threads? Because blocking syscalls are the enemy of M:N scheduling. When a goroutine enters a blocking syscall, its M blocks with it. Instead of stalling everything, the runtime detaches the P from that M and hands it to another M (spinning one up from a pool if needed) so other goroutines keep running. When the syscall returns, the old M tries to grab a P; if none is free, it parks the goroutine on the global queue and puts itself to sleep. This is why a program with hundreds of goroutines stuck in file I/O does not necessarily burn hundreds of threads — but also why syscall-heavy workloads can cause thread churn if you look at thread counts in a profiler and panic.

Local Run Queues and Work Stealing

Each P owns a local run queue: a fixed-size ring buffer of up to 256 runnable goroutines. When code calls go f(), the new G is placed on the current P’s local queue — specifically with runqput, which by default puts it in the runnext slot, meaning it will be scheduled next. This is a deliberate locality optimization: the goroutine you just spawned is very likely to touch the same cache lines as the code that spawned it.

When the local queue is full, half of it is shuffled to the global run queue. There is also a global queue that receives goroutines from network poller wakeups and timed callbacks. The scheduling hot path is findRunnable in runtime/proc.go, and its priority order tells you most of what you need to know about fairness:

  • Check for traced or GC-related goroutines that must run first.
  • Run the P’s runnext goroutine (the one just readied, for cache locality).
  • Drain the local run queue.
  • Every 61st scheduling tick, check the global queue first — a prime-ish constant chosen so that no pattern of scheduling behavior can starve the global queue systematically.
  • Poll the network poller for ready I/O.
  • If still nothing: steal work — visit other P’s in random order and take half of a victim’s run queue, or even steal its timers.

That last step is what makes the scheduler scale. Without work stealing, a goroutine-heavy burst on one thread would leave other cores idle while queues back up. The randomized victim order keeps stealing from becoming a contention hotspot, and stealing half the queue means work redistributes geometrically fast. The 61 tick detail is worth knowing when you read scheduler traces: a goroutine that sat in the global queue for a while probably lost the periodic coin flip, not anything sinister.

Preemption: The 10ms Ceiling

Early Go had a famous weakness: a tight loop like for {} could freeze the entire program, because goroutines only yielded at function calls, where the compiler inserted stack-growth checks. Go 1.14 introduced asynchronous preemption to fix this, and the mechanism is still in place.

A background thread called the sysmon runs without a P and watches for goroutines that have been running too long. The threshold is defined in the runtime as forcePreemptNS = 10ms. When a G exceeds it, the sysmon sends a signal to the thread running it; the signal handler executes asyncPreempt, which saves all user registers and pushes the goroutine back into the scheduler. Because this can interrupt a goroutine at (almost) any instruction, the runtime has to be careful at unsafe points — and it is, which is why you rarely notice it.

There is one documented exception worth internalizing: preemption happens at safe points, and a goroutine in a tight loop of fully inlined code can still delay preemption longer than you would expect. Also, runtime.Gosched() exists as an explicit yield, and the internal goyield path is what cooperative yielding bottoms out at. But the days of needing runtime.Gosched() calls sprinkled through CPU-bound loops are gone — if you still have one in your codebase, it predates Go 1.14.

Growable Stacks and What Blocking Really Costs

Goroutines are cheap because their stacks start small and grow on demand. The initial stack is 2 KB (stackMin = 2048 in runtime/stack.go). Compare that with an OS thread, where the kernel typically reserves a megabyte. The compiler inserts prologue checks before every function call; when the stack is insufficient, the runtime allocates a larger one (doubling up to the 1 GB limit) and copies the old stack over, adjusting every pointer into it. Because Go’s escape analysis and write barriers keep precise pointer maps, moving stacks is safe — this is the same trick that makes preemption-with-stack-copy possible at all.

Shrinking also happens: if the runtime notices a goroutine is using less than a quarter of its stack during GC, it can halve it. The practical consequence is that spawning a million goroutines is feasible (a million 2 KB stacks is about 2 GB at worst, far less in practice since most never grow), but a million goroutines each holding a 64 KB buffer is not a scheduler problem — it is a memory problem that people wrongly blame on the scheduler.

Similarly, “blocking” in Go is not one thing. Blocking on a channel, mutex, or the network poller parks the G and releases the P, costing only a scheduler transition. Blocking in a raw syscall holds the M but hands off the P. Blocking in cgo, or holding an OS-level lock inside a syscall, is the expensive kind. When you see high thread counts, the syscall path is usually the culprit, not goroutine volume.

Debugging With the Scheduler in Mind

Three tools turn this theory into practice. First, GODEBUG=schedtrace=1000 prints scheduler state every second: run queue lengths per P, thread counts, and global queue depth. Long local queues with idle P’s elsewhere are a classic imbalance signature that work stealing is fighting to correct. Second, runtime/metrics exposes the same data programmatically, including /sched/latencies:seconds, a histogram of how long goroutines waited to be scheduled — the single best number for detecting scheduler-induced tail latency.

Third, execution traces via runtime/trace show the G-M-P timeline directly. Go 1.25 added a flight recorder that keeps the last few seconds of trace in an in-memory ring buffer so you can snapshot the moments around a rare latency spike without paying the always-on tracing cost. Between scheduler latency histograms and trace timelines, almost every stall that looks mysterious turns out to be a lock convoy, syscall thread churn, GOMAXPROCS mismatched to a container quota, or a GC assist storm — all visible in these tools once you know what the scheduler is doing underneath.

Wrapping Up

The Go scheduler is a work-stealing M:N scheduler with cooperative yields at safe points, asynchronous preemption after 10 ms, 2 KB growable stacks, and — as of Go 1.25 — container-aware defaults for GOMAXPROCS. None of it is magic, and all of it is readable in the runtime source on GitHub. If you operate Go services under load, spend an hour with schedtrace output and the scheduler latency histogram on your next quiet Friday. The mental model you build will pay for itself the first time a latency spike stops being a mystery.

Leave a Reply

Your email address will not be published. Required fields are marked *