Most Go services run for months without anyone touching a single GC knob, and that is exactly how the runtime designers want it. But the moment your container gets OOM-killed at 3 a.m., or a p99 latency spike lines up with GC cycles in the trace, “the GC handles it” stops being an answer. The good news: the Go garbage collector is not a black box. Its behavior follows a small cost model that you can reason about with two knobs — GOGC and GOMEMLIMIT — and a handful of allocation habits that determine whether those knobs matter at all.
This post walks through how the collector actually works, where its CPU and memory costs come from, and what you can do about both. The official GC guide covers the same ground in more depth; what follows is the working subset that pays off in production services.
Stack, heap, and escape analysis
Before the GC gets involved, the compiler tries hard to keep values off the heap entirely. Non-pointer values in local variables usually live on the goroutine stack, where allocation is a stack-pointer bump and reclamation is free — the memory dies with the stack frame. A value “escapes to the heap” when the compiler cannot prove its lifetime: its address outlives the function, its size depends on a runtime value, or a reference to it is stored somewhere that itself escapes. Escape analysis is transitive — storing a pointer into an already-escaping value forces the pointee out too.
The practical consequence: heap allocation rate, not heap size alone, is what drives GC work. Every escaped value has to be traced eventually. The compiler will tell you its decisions with the escape-analysis flag:
go build -gcflags=-m ./...
type User struct {
Name string
Age int
}
// Score does not escape: it stays on the caller's stack.
func Score(u User) int {
bonus := 10
return u.Age + bonus
}
// profile escapes: the pointer is returned, so the compiler
// must allocate it on the heap.
func profile(name string) *User {
return &User{Name: name}
}
A concurrent mark-sweep collector
Go’s collector is a tracing, mark-sweep, non-moving collector. It walks the object graph starting from roots — goroutine stacks and globals — marks everything reachable, then sweeps unmarked memory back into the allocator. Objects never move, so pointers stay stable.
Three properties matter for anyone running a service. First, the collector is concurrent: most of the mark phase runs alongside your goroutines, and stop-the-world pauses happen only briefly at phase transitions and when goroutines are suspended for root scanning. Pauses are proportional to GOMAXPROCS rather than heap size, which is why a 40 GB heap service can still have sub-millisecond GC pauses.
Second, the mark phase claims roughly a quarter of the CPU while it is active, served by dedicated background mark workers. Third, when the application allocates faster than the background workers can keep up, allocating goroutines are drafted into mark assists — your request handler literally does GC work inline, which shows up as latency on that request. Assists are the mechanism that turns “allocation rate too high” into “tail latency you can feel.”
The cost model: live heap and allocation rate
The GC guide reduces the whole system to two costs. Memory cost per cycle is the live heap from the previous cycle plus everything allocated since. CPU cost is a small fixed amount per cycle plus a marginal cost proportional to the live heap — the GC only walks live objects, never the garbage.
From that model follows the central trade-off. GC frequency is decided by how much new memory may accumulate before the next cycle must start. Run cycles rarely and you spend little CPU on marking but hold more memory; run them often and you save memory but burn CPU. In steady state, doubling the allowed heap growth roughly halves GC CPU cost — a clean, tunable dial.
GOGC: the frequency dial
GOGC is that dial. After each cycle the runtime computes a target heap size:
Target = Live heap + (Live heap + GC roots) x GOGC / 100
The next cycle begins when total heap size reaches the target. With the default GOGC of 100, a service with a 200 MB live heap (plus small stacks and globals) lets the heap roughly double to about 400 MB before collecting. Raising GOGC to 300 quadruples the headroom and cuts GC CPU cost to about a third; lowering it to 50 halves memory overhead at the price of more frequent cycles. Since Go 1.18, root memory (goroutine stacks and globals) counts in the formula, which fixed pathological tuning for programs with hundreds of thousands of goroutines.
GOGC can be set through the environment variable or programmatically via runtime/debug:
import "runtime/debug"
func init() {
debug.SetGCPercent(200)
}
The catch: GOGC is proportional, so it must be tuned for the peak live heap. A service whose live heap normally sits at 200 MB but spikes to 1 GB during a batch job forces you to pick a conservative GOGC for a condition that holds 2% of the time — wasting memory the other 98%.
GOMEMLIMIT: the ceiling
Go 1.19 added a second knob that decouples the two problems. GOMEMLIMIT caps total runtime memory (heap plus stacks plus runtime metadata; roughly Sys minus released heap in MemStats terms) and makes the GC run more often as usage approaches it. Now the batch-spike service can keep GOGC high for the common case and simply refuse to cross the limit during spikes.
import "runtime/debug"
func init() {
// Aim high in the common case, hard-cap the tail.
debug.SetGCPercent(200)
debug.SetMemoryLimit(900 << 20) // 900 MiB
}
Two details make the limit safe to use. It is soft: the runtime caps GC CPU usage at about 50% over a window of a few GOMAXPROCS-seconds, so a badly chosen limit slows the program by at most roughly 2x instead of livelocking it in endless collection — the failure mode the docs call thrashing. An indefinite stall is usually worse than an OOM kill, which fails fast and loudly. And the limit is honored even with the GC disabled entirely, which enables the aggressive GOGC=off-plus-limit configuration for containerized services where the Go process owns the whole cgroup budget.
The common recipe for a containerized web service: set GOMEMLIMIT to about 90–95% of the container limit (leave headroom for memory the runtime cannot see — cgo allocations, non-heap pages), and leave GOGC at its default or raise it. Do not set a memory limit on CLI tools or libraries where you control neither the host nor the inputs.
Finding and fixing allocation hot spots
Knobs change where the GC spends its budget; allocation discipline changes the size of the budget. Three tools cover most investigations. CPU profiles expose GC time through runtime symbols — runtime.gcBgMarkWorker for background marking, runtime.gcAssistAlloc for assists (more than about 5% cumulative means allocation is outpacing the collector), and runtime.mallocgc for raw allocation cost. Heap profiles in alloc_space mode rank allocation sites by total bytes since process start, which is the metric that maps directly to GC frequency. And GODEBUG=gctrace=1 prints one line per cycle with live heap, target, and pause times — the quickest sanity check that a server is behaving.
When a hot spot needs fixing, the wins are usually structural rather than micro-optimizations:
- Reuse buffers across requests (sync.Pool or explicitly owned arenas per request) instead of allocating per iteration.
- Pre-size slices and maps when the length is known — repeated growth re-allocates and copies.
- Replace pointers with indices where the data structure allows it; every pointer is GC work.
For workloads where marking itself dominates even at low allocation rates, the collector has two implementation details worth exploiting. It segregates pointer-free values from pointer-bearing ones, and it stops scanning a value at its last pointer field. That makes struct layout a real, if minor, lever:
// Pointer-bearing fields first: the GC stops scanning
// after the last one, skipping the scalar tail entirely.
type Entry struct {
Next *Entry
Key string
ID uint64
Cost float64
Flags uint8
}
The guide’s own caveat applies: these layout tricks obscure intent and may stop mattering in a future release. Apply them only where profiling shows marking is genuinely hot.
A note on finalizers, cleanups, and weak pointers
One area where allocation habits and the GC intersect more subtly is object lifecycle observation. Go now offers three mechanisms — finalizers, cleanups (Go 1.24’s runtime.AddCleanup), and weak pointers — and the guidance has consolidated: prefer cleanups. Finalizers resurrect their object to pass it to the finalizer function, which delays reclamation by at least a cycle and makes reference cycles immortal. Cleanups avoid resurrection, run concurrently with each other, and allow multiple cleanups per object. The evergreen rule stands regardless of mechanism: a cleanup that captures the object it is attached to prevents that object from ever being reclaimed, so the cleanup never runs. Capture only the raw resource handle — a file descriptor, a C pointer — never the wrapper. And treat all of these as safety nets behind explicit Close methods, never as the primary resource-management strategy; the GC’s timing is not under your control.
Wrapping up
The mental model is small enough to keep in your head: the GC’s cost is dominated by live heap size times GC frequency, frequency is set by GOGC, GOMEMLIMIT bounds the worst case, and mark assists are how allocation pressure becomes user-visible latency. When something looks wrong, reach for gctrace and an alloc_space heap profile before touching the knobs — in most services the fix is one hot allocation loop, not a tuning session. When the knobs are the right tool, the container recipe of a soft memory limit with sensible headroom plus default GOGC covers the overwhelming majority of deployments.