Today I want to talk about split-brain. The name sounds dramatic, and the problem is dramatic too. One system splits into two halves, and each half thinks it is the boss. Both accept writes. The data slowly goes in different directions, and there is no clean way to merge it back.
How it happens
A crash almost never causes it. A crash is the easy case: the node is dead, so it writes nothing. For split-brain you need everyone alive, but wrong about the world.
There are two common ways to get there.
A network partition. Two nodes back each other up. The link between them dies. Each one thinks the other is dead, and both make themselves the leader.
A paused process. The leader freezes for some time (a long GC pause, a VM migration). The cluster picks a new leader. Then the old one wakes up and keeps working. It does not know it is old. Nobody sends you a message saying “you are stale”.
The classic cases
Database pairs. After a bad failover you have two primaries. Both take writes, and replication cannot fix it.
Leader-based clusters. After a partition the old leader keeps answering. Clients get stale reads, or writes that look accepted but later get lost.
Many data centers. During a partition both sides keep taking writes. When the network is back, you have two histories of the same data.
Distributed locks. A lock times out while its owner is paused. A second worker takes it. Now two processes hold the same lock.
Different setups, same problem: two writers with an old idea of who owns the data.
How systems protect themselves
Quorum. Nothing important happens without a majority. This is why clusters have 3 or 5 nodes: only one side of a partition can win.
Terms. In Raft every message has a term number. If a leader comes from an old term, the others just refuse it.
Fencing tokens. Every claim gets a number that only grows. The storage checks this number at the moment of the write. So a stale writer is stopped where it writes, not where it decides.
STONITH. The brutal option: power off the other node, so it cannot write anything.
My surprise: you do not need a cluster
When I read about split-brain in Designing Data-Intensive Applications, I thought: interesting, but not my problem. My side project runs on one server with one Postgres. Nothing to split.
Then I re-read my own postmortems and found this problem twice.
First time: a deploy and a background rebuild both thought they owned the same build folder. The rebuild said “success”, but that report was already stale, and the worker trusted it more than the disk. For 50 minutes Google saw 404 on every page.
Second time: one reader raced with themselves. A heartbeat and a tab close both did read-then-insert on the same progress row. Both saw no row, both did INSERT, and the second one failed with error 23505.
Both fixes use the fencing idea, just very small. For the first one, a timestamp that the wait step must beat. For the second one, a unique index, so the database rejects the second writer, and we catch the error and merge. No Raft, no quorum. Same main idea: check freshness at the place where you write.
If you get this in an interview
“Explain split-brain” is a classic system design question. The textbook answer is two leaders in a cluster. A stronger answer: any two writers with an old idea of who owns the data, and it can happen on one server. Give the definition, name the defenses (quorum, terms, fencing tokens), and if you have a real story, tell it. A story beats a definition every time.
Where did a distributed systems problem bite you in a system you thought was too small for it? I would love to hear your story.
I build TextStack, a reading app for developers who learn AI engineering. Both postmortems from this post are public in the repo: github.com/mrviduus/textstack, folder docs/incidents. Nothing here is made up.
TextStack is starting closed testing now, and I need a few early testers. If you read technical books like DDIA and want early access, leave a comment here or write me on dev.to.
Leave a Reply