Back to blog
Series · Day 21
Distributed Systems in 30 Days
View all lessons →
Agent split-brain

Day 21: Split-Brain Is the Default State in Multi-Agent Coding

Run two coding agents on the same file and here's what actually happens: both ship a fix, both pass CI, and the merge tool just... picks one. Silently. No vote, no flag, no diff review — just whichever PR landed second. That silent pick is a decision nobody made on purpose, and it'll come back as a bug with zero trail back to why.

The incident

Picture this: two agents get assigned the same flaky retry bug in the payment client. Agent A rewrites the backoff logic with jitter. Agent B — same model, different sample, slightly different prompt — rewrites the exact same function with a circuit breaker instead. Both pass CI. Of course they do: the test suite checks 'does the retry eventually succeed,' not 'which strategy did the team agree on.' The merge tool breaks the tie by picking whichever branch landed second, overwrites the other without a trace, and ships it. Three days later the circuit breaker trips under a load pattern the jitter version would've shrugged off. Nobody can explain why the fix 'disappeared' — git history shows a clean merge, not a decision that was ever made.

Why this is a distributed systems concept, not a tooling bug

Classic split-brain needs a network partition — two nodes lose contact, each assumes it's the leader, both write, and you reconcile once the partition heals. Multi-agent coding doesn't have that excuse. The agents can see the same repo, the same ticket, the same CI logs. No partition anywhere. They diverge anyway, because the fork isn't in the network — it's in the sampling. Hand two LLM calls the same prompt at a non-zero temperature and you get two different internal answers to what 'the fix' even is. The split is epistemic, not topological: each generation is its own branch of belief about what correct means, and there was never a moment the two agents actually agreed.

Connect back to the series

The primitives from earlier lessons still hold — they just move to a different layer:

  • ▹Leader election (who gets to act) becomes: which agent's output is authoritative when two touch the same file.
  • ▹Last-write-wins becomes the default failure mode, not a deliberate strategy — 'whichever branch merged second' is LWW with no clock and no owner.
  • ▹Vector clocks, which let you detect that two writes are causally concurrent rather than ordered, map to: you need a way to flag 'these two diffs both touch function X with no shared base reasoning' before merge, not after.
  • ▹Quorum reads/writes map to: don't accept a patch as 'correct' until some arbiter — human or agent — has confirmed it, not just that tests are green.

Where teams get this wrong

The most common mistake: treating 'both branches pass CI' as proof either one is safe to merge. Passing tests proves the branch agrees with the test suite. Full stop. It says nothing about whether the branch agrees with the other branch, or with the design intent anyone actually had in mind. Two implementations of the same fix can both look correct and still embody contradictory assumptions — retry vs. circuit-break, optimistic vs. pessimistic locking, sync vs. async error handling. CI is a correctness oracle for one branch in isolation. It is not an arbitration mechanism between branches. Skip arbitration and you've quietly handed the 'who's right' call to merge-tool coin flips, or to whichever PR someone happened to click first.

The fix is an authority model, not better merging

You can't out-engineer this. Stochastic agents will keep forking no matter how clever your prompts get. What you can do is decide — before the divergence happens — who or what gets to arbitrate. Three concrete options, borrowed loosely from consensus protocols:

  • ▹Designated reviewer agent: one agent (or one model call with a stricter prompt) is the sole arbiter for a given file or module — functionally a leader-elected authority for that shard of the codebase.
  • ▹Deterministic tie-break rule: when two agents touch the same file, apply a fixed, inspectable rule (e.g. 'the agent that owns the originating ticket wins,' or 'the diff with smaller blast radius wins') instead of merge-order luck.
  • ▹Escalation to a human: if the arbitration rule can't resolve it confidently (both diffs touch the same function, no clear owner), route to a person instead of auto-merging. This is your fallback quorum — don't silently promote one branch when the arbiter itself is uncertain.

What this looks like in practice

Here's a minimal arbitration rule you can bolt onto a two-agent pipeline this week: detect file-level overlap before merge, and force a decision instead of letting the merge tool make one for you by default.

yaml
# .agents/arbitration.yml
# Runs after both agents open PRs, before auto-merge is allowed.

on_overlap:
  detect: changed_files_intersect   # same file touched by >1 agent PR
  action: block_automerge

arbitration_rule:
  - if: pr.author == file_owner(path)   # CODEOWNERS-style mapping
    then: allow
  - if: pr.ticket_id == originating_ticket(path)
    then: allow
  - else: escalate_to_human

escalation:
  channel: slack://#eng-review
  payload: [diff_a, diff_b, shared_test_results, conflicting_assumptions_summary]

Takeaway

Stop trying to make agents agree. Non-deterministic sampling guarantees they won't — not sometimes, every time, forever. Design instead for the moment they diverge: decide the arbiter, the tie-break rule, and the escalation path before two agents ever touch the same file. Do that, and when split-brain hits — and it will, maybe on your very next overlapping task — you get a decision trail instead of a silent coin flip.

Flashcards
Check yourself

Extend your knowledge

  • ▹Re-read your notes on leader election and quorum from earlier in this series, and rewrite each in terms of 'which agent's PR wins' instead of 'which node accepts writes.'
  • ▹Audit your current multi-agent pipeline: pick one file that's commonly touched by more than one agent task, and write down who the arbiter is today. If the honest answer is 'merge order,' that's your gap.
  • ▹Try CODEOWNERS-style file ownership (natively supported on GitHub, and on GitLab as part of its code owners feature) as your first deterministic tie-break rule — it's the lowest-effort way to turn 'whoever merges second wins' into 'the designated owner wins.'
  • ▹Next time two agents both touch the same function, don't just merge the winner — read the loser's diff. The disagreement itself is often the most useful code review you'll get all week.
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Agent split-brain” — trade-offs, decisions, or the story behind it.