Reasoning vs Routing in ReAct Agents
Day 19 — Is Your ReAct Agent Actually Reasoning, or Just Narrating a State Machine?
Your roadmap deck says the agent "reasons" before it acts. Pull up an actual trace. At every branch point, was there ever more than one legal move on the table? If the answer is no, you didn't build a reasoning agent — you built a state machine that narrates itself in English. That gap isn't academic. It decides what you test, what you debug, and what you're allowed to promise the people funding this thing.
The PhoenixDX ticket-triage trace
Here's a real trace, pulled straight from a support-ticket triage agent we run internally at PhoenixDX. The task: decide what to do with a ticket where the customer says the invoice total doesn't match what they expected.
Thought: The user says the invoice total doesn't match what they expected. I should check if this is a known billing issue before doing anything else.
Action: search_kb(query="invoice mismatch")
Observation: 3 KB articles found, none match the exact discrepancy pattern described.
Thought: No KB article covers this specific case, so a templated reply would be guessing. This needs a human with billing system access.
Action: escalate_to_human(reason="billing discrepancy, no KB match")
Observation: Ticket escalated to billing-tier2 queue.Read the Thought lines by themselves and they sound exactly like reasoning — weighing evidence, ruling an option out, justifying a call. Now look at what the system prompt actually lets the agent do at each step. The moment a ticket lands, search_kb is mandatory — not chosen, the only move on the board. After a KB miss, the schema exposes exactly two actions: reply_with_template, locked behind a confidence score the agent has zero control over, and escalate_to_human. Confidence comes in under threshold, so reply_with_template was never really available. Two branch points, one legal action at each. The Thought didn't decide anything. It narrated a conclusion the tool schema had already reached.
The tell: delete the Thought and rerun it
Here's the fastest way to expose this. Strip every Thought line out of the trace, feed the exact same Observations back through the loop, and watch whether the Action sequence changes. Run that on the trace above and nothing moves — search_kb still fires first because it's hardcoded as the entry action, and escalate_to_human still fires on a KB miss because reply_with_template was never really on the table. The model could've written "Thought: I like turtles" at every step and the agent would behave identically.
So that's the test, in one line: if deleting the Thought text doesn't change what the agent does, the Thought wasn't doing reasoning work — it was narration bolted onto a decision the action space had already made. Genuine reasoning looks different. The Thought is causally load-bearing: change its conclusion and the next Action actually changes, because more than one option was legally available and the model had to pick based on specifics of the situation in front of it.
Why this connects to action-space design
This is the flip side of the tool-design lesson from a few days back. When we talked about shaping actions tightly instead of leaving them loose, the pitch was reliability — fewer malformed calls, fewer hallucinated parameters, fewer dangerous side effects. Nobody said the quiet part out loud: every time you tighten the action space for reliability, you shrink the room where real reasoning can happen. The work quietly shifts from "write a good Thought prompt" to "constrain the menu so hard that whatever the model writes in Thought, there's only one door it can walk through." That's a legitimate engineering move. It's just not the same thing as building a reasoning agent — and conflating the two is exactly how teams end up shipping a routing layer, calling it "agentic reasoning" in the deck, and then can't explain why the agent "reasoned" its way to the same answer on ten different inputs that should have required ten different judgment calls.
Before/after: genuine ambiguity vs. FSM in disguise
Same ticket-triage task, two different action space designs.
BEFORE — 12 loosely-specified actions (genuine ambiguity, Thought does real work):
search_kb, read_account_history, check_payment_status,
check_subscription_tier, query_billing_system, draft_reply,
send_reply, escalate_to_tier2, escalate_to_billing_team,
apply_credit, close_ticket, flag_for_review
At any point the model could legally call several of these in
different orders, combine them, or skip some entirely. The Thought
has to decide: which signals to gather, in what order, and whether
the evidence justifies a credit vs. an escalation. Different tickets
genuinely produce different sequences.AFTER — 3 tightly-typed actions (FSM in disguise):
gather_context(ticket_id) // internally runs the lookups, fixed order
resolve(action: "reply" | "escalate" | "credit", payload)
close(ticket_id)
The schema now forces gather_context first, always, with no choice
of what to look up or in what order. resolve() only accepts three
enum values, each gated by hardcoded thresholds elsewhere in the
prompt. The Thought now just announces which bucket the ticket fell
into — a classification label dressed up as deliberation.The "after" version is probably the right call for a production support system — it's more reliable, easier to test, and a lot safer around money (apply_credit running loose was a real risk in version one). Building the 3-action version isn't the mistake. Calling it "the agent reasons about the best resolution path" is. What it actually does is classify, then route.
The diagnostic test for your own agent
- ▹Pull 5-10 real traces from your agent's logs.
- ▹Mark every Thought→Action boundary as a branch point.
- ▹At each branch point, check the tool schema and the prompt: how many actions were actually legal right there — not how many exist in the schema globally, but how many were reachable given the current state and gating logic.
- ▹One option reachable? Label it 'routing' — the Thought is narration, not reasoning.
- ▹Two or more options reachable, and the choice actually depended on details specific to this input (not a fixed rule like 'confidence < 0.7')? Label it 'genuine reasoning'.
- ▹Tally the ratio across your traces. When we ran this on our own ticket-triage traces, it skewed hard toward routing. That's not a failure — it's information you can now act on deliberately.
Name it honestly
None of this is an argument against constrained action spaces. Tight, FSM-like agents are often the right engineering call when a wrong action is expensive — money moving, irreversible API calls, a reply that goes straight to a customer. The real failure here is organizational, not technical: call a routing layer "agentic reasoning" in a roadmap doc and you've set expectations the system can't meet. You make failures harder to diagnose, because now someone's hunting for a reasoning bug inside what's really a lookup table. And you lose the ability to have an honest conversation about where you genuinely need an open-ended action space versus where you deliberately closed one off for safety. Run the diagnostic. Label each branch point honestly. Design the next version on purpose, not by accident.
Next: Day 20
Tomorrow flips this around: an action space that's genuinely open-ended, where the model really is choosing among many unconstrained options. That's a different reliability problem entirely, and the fixes that work for a tightly-typed router don't carry over.
Extend your knowledge
- ▹Run the strip-the-Thought test on your own agent's last 10 production traces this week and tally the routing-vs-reasoning ratio.
- ▹Re-read your team's roadmap or pitch doc for the word 'reasoning' — check each use against the diagnostic from this lesson.
- ▹Revisit the original ReAct paper (Yao et al., 'ReAct: Synergizing Reasoning and Acting in Language Models') with this routing/reasoning distinction in mind — notice which of their examples have genuinely open action spaces versus narrow ones.
- ▹Look at your tool schema's enum-typed parameters specifically — those are usually where routing hides, since an enum caps the real option count by construction.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.