Skip to content
jieqi's archive
Go back

C++: Atomics and Memory Ordering

Contents

std::atomic<T> is a wrapper that makes operations on a T indivisible across threads, and gives you a knob to control ordering relative to other memory.

Two separate guarantees, worth keeping apart.

1. Atomicity: No Torn or Interleaved Access

int counter = 0;           // thread A and B both do counter++
                           // ends up < 2000000. Lost updates.

std::atomic<int> counter{0};
counter.fetch_add(1);      // exactly 2000000. Never loses one.

counter++ on a plain int is load, add, store. Two threads interleave and one overwrites the other. fetch_add is one indivisible operation.

Two threads, a million increments each:

expected      : 2000000
plain int     : 1097940   (lost 902060)
atomic<int>   : 2000000

Over 900,000 increments vanished, and the number is different on every run.

The bug hides under optimization

That measurement is at -O0. At -O2 the plain int prints 2000000 every run, because the loop collapses into one read-modify-write:

ldr  w9, [x8, _plain@PAGEOFF]
add  w9, w9, #244, lsl #12      ; = 999424
add  w9, w9, #576               ; = 1000000 total
str  w9, [x8, _plain@PAGEOFF]

A million increments became one add. The race window is now a single instruction pair, so the bug stops reproducing — hidden, not fixed.

The compiler may do this because a data race is undefined behaviour, not merely a lost update. Once it may assume the race cannot happen, asking which interleavings occur is asking about a program the standard has stopped describing.

Tearing

A plain read can also observe a half-written value. An atomic read never does: the old value or the new one, never a mixture.

An aligned int won’t tear on x86-64 or ARM64, so this never shows on the machine you tested. The standard guarantees nothing for any type, and anything wider than a machine word tears readily.

Core Memory Ordering Relations

Sequenced-before — a strict, asymmetrical, intra-thread relationship where one evaluation happens before another within the same thread.

Synchronizes-with — an inter-thread relationship where an atomic release operation on a variable matches with an atomic acquire load on the same variable, sharing visibility.

Happens-before — a combined relation. If one operation is sequenced-before or synchronizes-with another, it happens-before it, guaranteeing state visibility across threads.

The three compose, and that composition is the point:

int data = 0;                       // plain int, NOT atomic
std::atomic<bool> ready{false};

// thread A
data = 42;                                      // (1)
ready.store(true, std::memory_order_release);   // (2)

// thread B
while (!ready.load(std::memory_order_acquire))  // (3)
    ;
assert(data == 42);                             // (4) guaranteed

Read it as a chain:

(1) sequenced-before  (2)      same thread
(2) synchronizes-with (3)      release pairs with acquire on `ready`
(3) sequenced-before  (4)      same thread
------------------------------------------------
(1) happens-before    (4)      so `data` is visible, and 42 is guaranteed

Note what carries the payload. ready is the only atomic; data is an ordinary int that no thread ever guards. The release/acquire pair on one variable drags everything sequenced-before it into visibility. That is the whole mechanism — one atomic publishes an arbitrary amount of ordinary memory.

It is also load-bearing rather than decorative. Swap both operations to memory_order_relaxed and the atomicity of ready is unchanged, but the pairing is gone, so nothing synchronizes-with anything and data becomes a plain data race. ThreadSanitizer sees the difference immediately:

release/acquire  ->  clean over 200000 iterations
relaxed          ->  WARNING: ThreadSanitizer: data race
                     Read of size 4 ... by thread T2

Understanding Acquire-Release and Relaxed

Acquire-release is a pattern where a release store and an acquire load, when the load observes the store, form a synchronizes-with edge, which makes everything the producer did before the release happen-before everything the consumer does after the acquire.

// Thread A
data = 42;                       // payload
ready.store(true, release);      // flag

// Thread B
while (!ready.load(acquire)) {}  // spin
observed = data;                 // read payload
  CORE 0   producer          I   CORE 1   consumer
                             N
  STORE BUFFER               T   STORE BUFFER
  +----------------------+   E   +----------------------+
  |      -- empty --     |   R   |      -- empty --     |
  +----------------------+   C   +----------------------+
                             O
  PRIVATE L1                 N   PRIVATE L1
  +----------------------+   N   +----------------------+
  | data        0   [SH] |   E   | data        0   [SH] |
  | ready   false   [SH] |   C   | ready   false   [SH] |
  +----------------------+   T   +----------------------+
                             ·
                            MESI

Watching It Happen

The visual below runs this exact program one micro-event at a time.

Each core shows two things. The store buffer is where a write lands first — a retired store sits there, visible only to its own core, before it reaches cache. The private L1 below it holds the values other cores can actually observe, each tagged with its MESI state: SH shared, MOD modified, IN invalid.

The whole question is the order in which core 0’s two buffered writes drain, and what core 1 can see in between.

Toggle the mode to compare:

Same code shape, same atomic variable. The only difference is whether an edge was ever built.

Previous Post
C++: Error Handling Pt 1 - Error Codes
Next Post
C++: Lambda Basics