std::atomic<T> is a wrapper that makes operations on a T indivisible across threads, and gives you a knob to control ordering relative to other memory.
Two separate guarantees, worth keeping apart.
1. Atomicity: No Torn or Interleaved Access
int counter = 0; // thread A and B both do counter++
// ends up < 2000000. Lost updates.
std::atomic<int> counter{0};
counter.fetch_add(1); // exactly 2000000. Never loses one.
counter++ on a plain int is load, add, store. Two threads interleave and one overwrites the other. fetch_add is one indivisible operation.
Two threads, a million increments each:
expected : 2000000
plain int : 1097940 (lost 902060)
atomic<int> : 2000000
Over 900,000 increments vanished, and the number is different on every run.
The bug hides under optimization
That measurement is at -O0. At -O2 the plain int prints 2000000 every run, because the loop collapses into one read-modify-write:
ldr w9, [x8, _plain@PAGEOFF]
add w9, w9, #244, lsl #12 ; = 999424
add w9, w9, #576 ; = 1000000 total
str w9, [x8, _plain@PAGEOFF]
A million increments became one add. The race window is now a single instruction pair, so the bug stops reproducing — hidden, not fixed.
The compiler may do this because a data race is undefined behaviour, not merely a lost update. Once it may assume the race cannot happen, asking which interleavings occur is asking about a program the standard has stopped describing.
Tearing
A plain read can also observe a half-written value. An atomic read never does: the old value or the new one, never a mixture.
An aligned int won’t tear on x86-64 or ARM64, so this never shows on the machine you tested. The standard guarantees nothing for any type, and anything wider than a machine word tears readily.
Core Memory Ordering Relations
Sequenced-before — a strict, asymmetrical, intra-thread relationship where one evaluation happens before another within the same thread.
Synchronizes-with — an inter-thread relationship where an atomic release operation on a variable matches with an atomic acquire load on the same variable, sharing visibility.
Happens-before — a combined relation. If one operation is sequenced-before or synchronizes-with another, it happens-before it, guaranteeing state visibility across threads.
The three compose, and that composition is the point:
int data = 0; // plain int, NOT atomic
std::atomic<bool> ready{false};
// thread A
data = 42; // (1)
ready.store(true, std::memory_order_release); // (2)
// thread B
while (!ready.load(std::memory_order_acquire)) // (3)
;
assert(data == 42); // (4) guaranteed
Read it as a chain:
(1) sequenced-before (2) same thread
(2) synchronizes-with (3) release pairs with acquire on `ready`
(3) sequenced-before (4) same thread
------------------------------------------------
(1) happens-before (4) so `data` is visible, and 42 is guaranteed
Note what carries the payload. ready is the only atomic; data is an ordinary int that no thread ever guards. The release/acquire pair on one variable drags everything sequenced-before it into visibility. That is the whole mechanism — one atomic publishes an arbitrary amount of ordinary memory.
It is also load-bearing rather than decorative. Swap both operations to memory_order_relaxed and the atomicity of ready is unchanged, but the pairing is gone, so nothing synchronizes-with anything and data becomes a plain data race. ThreadSanitizer sees the difference immediately:
release/acquire -> clean over 200000 iterations
relaxed -> WARNING: ThreadSanitizer: data race
Read of size 4 ... by thread T2
Understanding Acquire-Release and Relaxed
Acquire-release is a pattern where a release store and an acquire load, when the load observes the store, form a synchronizes-with edge, which makes everything the producer did before the release happen-before everything the consumer does after the acquire.
// Thread A
data = 42; // payload
ready.store(true, release); // flag
// Thread B
while (!ready.load(acquire)) {} // spin
observed = data; // read payload
CORE 0 producer I CORE 1 consumer
N
STORE BUFFER T STORE BUFFER
+----------------------+ E +----------------------+
| -- empty -- | R | -- empty -- |
+----------------------+ C +----------------------+
O
PRIVATE L1 N PRIVATE L1
+----------------------+ N +----------------------+
| data 0 [SH] | E | data 0 [SH] |
| ready false [SH] | C | ready false [SH] |
+----------------------+ T +----------------------+
·
MESI
Watching It Happen
The visual below runs this exact program one micro-event at a time.
Each core shows two things. The store buffer is where a write lands first — a retired store sits there, visible only to its own core, before it reaches cache. The private L1 below it holds the values other cores can actually observe, each tagged with its MESI state: SH shared, MOD modified, IN invalid.
The whole question is the order in which core 0’s two buffered writes drain, and what core 1 can see in between.
Toggle the mode to compare:
- release/acquire — the payload drains first and the flag second, because release forbids a prior store from becoming visible after it. When core 1’s acquire load finally observes the flag, the edge forms and
datais guaranteed to be 42. - relaxed/relaxed —
readyis still atomic, so the flag itself is never torn. But nothing orders the drain, so the flag can overtake the payload: core 1 leaves the spin loop and reads a staledata == 0.
Same code shape, same atomic variable. The only difference is whether an edge was ever built.