Liveness vs Progress Detection
Day 20: Your Heartbeat Can't Tell Alive From Useful
A heartbeat tells you exactly one thing: the process hasn't died. It says nothing about whether that process is doing anything worth the compute you're paying for. In multi-agent systems, that gap is where the expensive failures hide — in plain sight, dashboard all green.
The zombie that passed every check
I've watched an agent sit in a loop, re-reading the same file and re-planning the same step, for hours. Every tick it called a tool. Every tool call returned. Every heartbeat went out on schedule. The orchestrator dashboard stayed solid green the whole time. The bill did not. A crashed agent gets noticed in minutes, because someone's sitting there waiting on its output. A looping agent that keeps heartbeating can run until someone happens to check the token spend — nothing in the system was ever designed to mark it dead.
Naming the conflation
Earlier in this series we covered failure detectors — Chandra and Toueg's unreliable failure detector model, phi-accrual's suspicion score instead of a binary alive/dead verdict. Both are good answers to a specific question: has this process crash-stopped? They assume that anything not crashed is doing its job. That assumption held fine for the systems they were built for — a replica either serves requests or it doesn't, a node either participates in consensus or it's partitioned out. It doesn't hold for an LLM agent. An agent can be fully alive, fully responsive to pings, and fully useless, because 'alive' and 'making progress on the task' are now two properties that can split apart — and classical liveness detection was never built to tell them apart.
Walking the actual failure
Here's the concrete version: an agent is told to update a config file based on a previous step's output. That output is malformed in a way that never throws an error — say, it points at a file section that doesn't exist. The agent reads the file, doesn't find what it's looking for, decides to re-read 'just to be sure,' doesn't find it again, re-plans the same step, calls the read tool again. Every one of those calls is a real network round trip, real tokens, a real heartbeat tick. The orchestrator's liveness probe sees a process responding, tool calls succeeding, heartbeat on time. Green across the board. Nothing in that signal path distinguishes this from an agent making steady progress on a long task.
Why this isn't just a harder version of an old problem
A hung RPC or a crashed process is catchable by timeout, because when the work is genuinely stuck, the signals actually stop — no response, no CPU movement, silence. A stuck reasoning loop doesn't go quiet. It keeps producing exactly what a naive liveness probe is built to look for: CPU activity, outbound calls, even partial tokens streaming back. This is the part that doesn't transfer from classical distributed systems for free. In the old model, 'the signal stopped' was a reliable stand-in for 'the work stopped.' In an agent loop, signal and work have come uncoupled — the agent can be extremely busy while making zero forward progress, a failure mode that simply didn't exist for crash-stop processes.
- ▹Liveness detection asks: is the process responding right now?
- ▹Progress detection asks: is the state of the task any different from the last time I looked?
- ▹An agent can pass the first check forever while permanently failing the second.
The fix: a second, orthogonal signal
You don't fix this by tuning the heartbeat harder — shorter intervals, stricter timeouts, backoff and retry. All of that still measures responsiveness and nothing else. What you need is a second signal, checked independently: evidence of progress. Concretely, something that moves monotonically when real work happens and sits still when the agent is looping — a new artifact written, a state delta landed in a shared store, a checkpoint counter that only ticks up on a successful, distinct step. The heartbeat proves the lights are on. The progress check proves someone's home.
Design sketch: two timers, two responses
In practice that means running two independent checks against the same agent, on different clocks, with different consequences when each one fails.
# Liveness check: fast, cheap, binary
def liveness_check(agent):
return agent.last_heartbeat_at > now() - LIVENESS_TIMEOUT # e.g. 30s
# Progress check: slower window, compares state not responsiveness
def progress_check(agent, window=PROGRESS_WINDOW): # e.g. last 5 checkpoints
recent = agent.checkpoint_history[-window:]
distinct_states = {c.hash for c in recent}
return len(distinct_states) > 1 # state actually changed
# Orchestrator loop
if not liveness_check(agent):
restart_agent(agent) # process is gone, cheap to replace
elif not progress_check(agent):
escalate_for_human_review(agent) # process is fine, work is stuck
The split in responses matters as much as the split in detection. A liveness failure is cheap and mechanical to fix — kill it, restart it, you lose at most one in-flight step. A progress failure is not something a restart fixes, because the agent isn't broken, the plan is. Restarting it usually just reproduces the same stuck loop with a fresh heartbeat. That case needs a human, or at minimum a different agent running a different prompt, to figure out why it's stuck.
The general principle
None of this is specific to agent loops — it's the shape of the problem anywhere 'running' and 'working' can come apart. A queue consumer that's ack'ing messages without processing them. A sync job that's connected but stalled on a retry loop. The rule for this whole course, not just today: if your system can be alive and useless at the same time, you need two detectors measuring two different things — not one detector tuned harder.
Extend your knowledge
- ▹Re-read Chandra & Toueg's 'Unreliable Failure Detectors for Reliable Distributed Systems' (1996) with this lesson in mind — pin down exactly which assumption (crash-stop, no Byzantine behavior, silence implies failure) breaks for LLM agents.
- ▹Audit one agent you run in production today: do you have any signal besides heartbeat and logs that proves forward progress? If not, that's this week's fix — a checkpoint counter is often a one-line addition.
- ▹Check how the orchestration framework you already use (LangGraph, Temporal workflows, or your own agent runner) exposes step/state history — is that history queryable as a progress signal, or is it only ever used for replay and debugging?
- ▹Read up on watchdog timers in embedded and real-time systems — the classical ancestor of 'liveness isn't enough,' and a useful contrast for how much cruder the agent version of this problem looks by comparison.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.