Go vs Rust vs Free-Threaded Python: Memory Models Compared

Every concurrent program rests on a promise most developers have never read: the language’s memory model. It is the contract between your code, the compiler, and the CPU that determines what one thread is allowed to see when another writes. When it holds, concurrency feels straightforward. When you step outside it — a data race here, an unsynchronized flag there — failures appear that look impossible: torn snapshots under load, mutexes that spin at 100 percent CPU, benchmarks that pass on every core count and corrupt data on the one machine in the fleet with a different CPU architecture.

Three ecosystems have converged on the same underlying model — happens-before edges, acquire/release semantics, sequential consistency as the safe default — while making very different trade-offs about who enforces it. Go puts the race detector after the fact. Rust makes data races a compile error and asks you to pick memory orderings yourself. CPython is ripping out the global lock that made the whole question moot, and discovering what that promise was actually worth. Understanding all three teaches you something no single one can: which parts of “it works” are guarantees, and which are luck.

Why reordering breaks “obviously correct” code

Two things stand between the program you wrote and the instructions that execute. The compiler reorders or eliminates operations it believes are equivalent — sometimes hoisting a load out of a loop because it can prove no write in the loop touches that address, which is false the moment another thread writes it. The CPU does its own reordering at runtime: stores sit in per-core store buffers, and on weakly-ordered architectures like ARM the effects of independent stores can become visible to other cores in a different order than the program issued them.

The Go Memory Model names the result directly: a data race is two goroutines accessing the same variable concurrently, at least one writing, without a synchronization operation ordering them — and it’s not “undefined values,” it’s undefined behavior. The spec’s wording is famously blunt: reads of variables that race with writes may observe any value in the variable’s type range. Not the previous value, not a torn half-write. Any value. The compiler is allowed to assume races never happen and optimize accordingly.

The canonical broken example is a boolean flag used as a ready signal:

var ready bool
var msg   string

go func() {
    msg = "hello"       // write 1
    ready = true        // write 2
}()

for !ready {            // no sync edge: data race
    runtime.Gosched()
}
fmt.Println(msg)        // LEGALLY may print ""

Nothing orders write 1 against the read of msg. The fix is to move the data through a synchronization edge rather than alongside it:

done := make(chan struct{})

go func() {
    msg = "hello"
    close(done)         // close happens-before the receive observes it
}()

<-done
fmt.Println(msg)        // guaranteed "hello"

The happens-before rule doing the work: a send (or close) on a channel happens before the corresponding receive completes. Same for mutexes — an unlock happens before any subsequent lock of that mutex — and for sync.Once, WaitGroup, and goroutine creation/join. Everything written before the edge is visible after it. That single idea is the entire model; the rest is bookkeeping.

Go’s bet: shared memory plus a race detector

Go’s design philosophy is that goroutines and channels make concurrent code easy to write correctly, but the language doesn’t prevent races — it detects them. go build -race instruments every memory access with happens-before tracking (an algorithm derived from ThreadSanitizer) and reports any two racing accesses with full stack traces of both. It’s the most practical race-finding tool in any ecosystem, and the reason Go teams can run shared-memory designs in production without living in fear.

The cost is that the detector is a test-time tool: 5-10x CPU and memory overhead means it’s for CI and soak environments, not production. And it only finds races that actually execute — a race in a rarely-hit error path can survive years of 80 percent code coverage. The discipline that works: services run under the race detector in CI on every commit, plus a periodic race-enabled soak under production-like traffic patterns, because race manifestation depends on interleavings you can’t reach with unit tests alone.

The Go memory model document includes an advisory that’s easy to miss: if you must use sync/atomic to fix a race, treat it as a signal to redesign the synchronization strategy instead. Atomics in Go are deliberately minimal — there is no AtomicBoolean.CompareAndSet with ordering parameters, no explicit fence API — because the language’s position is that atomic operations are for the runtime and a handful of synchronization primitives, not a general programming tool. Compared to what follows, that restraint is a feature.

One more Go-specific point that trips people up: the happens-before guarantees are exactly what the spec says they are, nothing more. A 64-bit write is not guaranteed atomic on 32-bit platforms without sync. A mutex-protected variable stays protected only if every access goes through the mutex — one code path that reads it directly reintroduces the race. The model is tight but narrow, and the narrowness is what makes it checkable.

Rust: preventing races in the type system — and then picking a fence

Rust attacks the same problem one layer earlier. Ownership means every value has exactly one owner; the borrow checker enforces that aliasing and mutation never coexist; the Send and Sync marker traits extend the rule across threads. A &T reference is only shareable across threads if T: Sync, and a value can only be moved to another thread if T: Send. The result: data races in safe Rust are a compile error. Not a warning, not a detector report — the program does not build. That’s a strictly stronger guarantee than Go’s, and the reason Rust needs no race detector (though the tooling exists, and unsafe blocks still need it).

The price is that atomics become your problem. Rust’s atomic types follow the C++11 model — the same one C++ and effectively every modern systems language adopted — and every atomic operation takes an explicit Ordering parameter. The std::sync::atomic module docs are the practical reference; the three levels you actually choose between:

  • Relaxed — atomicity with no ordering. Two threads incrementing a counter will never tear a value, but nothing else is guaranteed: a thread that reads the counter has no guarantee about what else it sees.
  • Acquire/Release — the workhorse. A Release store paired with an Acquire load on the same variable creates a happens-before edge: everything written before the store is visible after the load observes it. This is how lock-free “publish” works.
  • SeqCst — all SeqCst operations across all threads agree on a single total order. The most conservative and most expensive option; the right default when you can’t prove a weaker ordering suffices.
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;

static CONFIG_READY: AtomicBool = AtomicBool::new(false);
// In real code: OnceLock<Config> — a plain static mut here keeps the
// example focused on the ordering pair.
static mut CONFIG: Option<Config> = None;

fn load_config() {
    let cfg = parse_config_from_disk();      // expensive, single-threaded
    unsafe { CONFIG = Some(cfg); }
    CONFIG_READY.store(true, Ordering::Release);   // publish: everything above is visible...
}

fn config_ready() -> bool {
    CONFIG_READY.load(Ordering::Acquire)     // ...to any thread that sees this
}

Swap either half to Relaxed and the guarantee evaporates: a thread could observe CONFIG_READY == true while seeing a half-initialized or entirely absent CONFIG. This is the single most common real-world atomics bug — the data is protected by an ordering the flag doesn’t actually provide. The compiler will not catch it; the race detector may not, either, because whether it manifests depends on timing. The rule of thumb: Relaxed is only correct when the atomic value carries no information about other memory. Counters qualify. Flags that publish data never do.

There’s a hardware angle too. x86 is the friendliest architecture you’ll ever run on — loads have acquire semantics and stores release semantics built in. ARM, which now runs a significant share of production workloads (Graviton, Apple Silicon, mobile), does not: it permits ordering violations x86 never exhibits. Code with a mis-ordered atomic can pass every test on a developer laptop and fail intermittently on ARM production hardware. If you deploy to ARM, test on ARM; aarch64 CI runners are cheap insurance against a class of bug that otherwise only appears as a 3 a.m. page.

CPython: when the memory model was the GIL

For most of its life, CPython had the simplest possible answer to all of this: one lock. The Global Interpreter Lock meant only one thread executed Python bytecode at a time, so Python-level data races on object state were impossible by construction. The GIL was the memory model — so coarse it was invisible. Nobody thought about happens-before because there was exactly one thread that could observe anything at a time.

PEP 703 — “Making the Global Interpreter Lock Optional in CPython” — removes it. Accepted in 2023, it shipped as an optional build configuration (the free-threaded build, python3.14t) alongside the traditional GIL build, and per the PEP 779 acceptance criteria, the free-threaded build has now reached official support in the 3.14 release cycle. Per-object locks replace the global one, biased reference counting keeps the uncontended fast path single-threaded, and immortalization freezes refcounts for objects like small integers that every thread touches.

The migration cost lands on extension authors and on code with hidden assumptions. C extensions that manipulate Py_REFCNT directly, or that stash mutable state in module globals without locks, were “safe” only because the GIL made every bytecode atomic. Under the free-threaded build those are real races. The ecosystem response — per-object locking in the runtime, thread-safety audits across the extension ecosystem — is essentially the whole Go and Rust story playing out in reverse: a language that never needed an explicit memory model discovering what breaks when the implicit one goes away.

The practical takeaway for Python teams: free-threaded builds are opt-in and remain so, most pure-Python code is unaffected, but any deployment considering python3.14t needs to audit its dependencies the way a Go team audits for races — because the guarantee they were relying on was never theirs, it was the interpreter’s.

The traps that survive every memory model

Three failure patterns account for most real-world concurrency bugs, in all three ecosystems:

  • Publishing without an edge. Initialize an object, store a pointer to it, and let another thread read — with a plain store or a relaxed atomic. The reader can see the pointer but not the initialization. In Rust this is the relaxed-store-instead-of-release bug; in Go it’s the boolean flag from earlier; in post-GIL Python it’s any module-level cache populated lazily. The fix is always the same: the publication must go through a happens-before edge, whether that’s a Release store, a channel, or a lock.
  • Assuming coherence is ordering. It’s widely believed that because cache coherence protocols keep all cores’ views consistent, reordering “can’t happen on real hardware.” Coherence guarantees a single order of writes to each individual location. It says nothing about the order between two different locations — two independent stores can become visible to other cores in different orders. This is exactly the gap Dekker’s algorithm exposes, and why “it’s coherent so fences are unnecessary” is wrong.
  • Context switches as false synchronization. A rare race that “never reproduces” often disappears when a debugger, logger, or extra CPU load perturbs timing. Debuggers serialize execution; production parallelism doesn’t. The fix is never “add a sleep” — it’s finding the missing happens-before edge.

A field guide

  • Writing Go services: share memory through channels and mutexes by default; let the race detector find what slips through; treat any urge to reach for sync/atomic as a design smell worth a second look.
  • Writing Rust: stay in safe code and the compiler enforces everything. The moment you write Ordering::Relaxed, ask whether the atomic value is used to make decisions about other memory — if yes, you almost certainly need Acquire/Release. Default to SeqCst when unsure.
  • Deploying to ARM: test there. x86 hides ordering bugs that ARM exhibits.
  • Running Python: the GIL’s protection was the interpreter’s, not yours. Audit C extensions and shared global state before adopting free-threaded builds, and keep an eye on the PEP 703 adoption timeline as the ecosystem catches up.

Memory models reward the engineers who read them before production does. The happens-before edge is the atom of every guarantee in every one of these systems — find it for each piece of shared state in your program, or add one, and the impossible bugs stop being possible.

Leave a Reply

Your email address will not be published. Required fields are marked *