Back to blog
Series · Day 19
Engineering Leadership in 30 Days
View all lessons →
Urgency vs system risk

Why this matters

Your prioritization framework is lying to you. It's not ranking what will break the system — it's ranking who shouted loudest and most recently. On a team where agents run unattended and failures stay silent until they're catastrophic, that's not a minor miscalibration. That's where your next 3am page comes from.

3am, the pager, and the ticket nobody finished

The page lands for an inference cluster quietly dying: retry storms piling up, p99 climbing, an autoscaler thrashing because nobody touched its concurrency ceiling after the last model swap. None of this was on anyone's roadmap. Meanwhile, half-finished in the backlog sits last week's 'urgent' ticket — a CEO-requested feature, fast-tracked past everything else, for a demo. The feature shipped on time. The autoscaler review that got bumped to make room for it did not. Those two facts are not a coincidence.

Rewind: the ask that jumped the queue

A week earlier: the CEO asks for a new capability ahead of a customer call. Reasonable ask. Real urgency. Real authority behind it. The team pulls two engineers off infra hardening to ship it. In the room, every signal points the same way — a named stakeholder, a dated deadline, visible consequences for saying no. Nobody frames it as a tradeoff against system risk, because the system risk isn't making any noise yet. It never does, until it does.

What the loud ask displaced

What got bumped was a review of autoscaling thresholds and retry/backoff settings after a model version upgrade quietly changed the latency profile of every request. Nobody asked for this work. No deadline, no name attached, no demo riding on it. It was also the thing holding the floor up — the capacity assumptions the entire inference layer depended on. That's the nature of maintenance work: it never looks urgent, right up until it's the only thing that matters.

The mechanism: urgency measures leverage, not need

Here's the mechanism underneath all of this. Urgency, as it shows up in a backlog, measures who has the standing to be heard right now — not what the system actually needs. A CEO's request is urgent because a CEO made it. A capacity regression is 'not urgent' because nothing's on fire yet and nobody owns noticing it. These are two different signals, uncorrelated at best. And because the loud one eats the team's attention and context-switching budget, it's often inversely correlated with the quiet one getting done. Every queue has a default arbiter of urgency. If you never name one on purpose, it's whoever asked most recently and most forcefully.

The break

Here's what the incident actually looked like underneath. The model swap three weeks earlier had pushed token-generation latency up by a meaningful margin. The autoscaler's scale-out rules were still tuned to the old latency profile, so under real traffic it under-provisioned. Retries piled up. Retries looked like more traffic. The autoscaler chased a number that kept moving out from under it. The fix itself took under an hour once someone actually looked. The real cost was four hours of degraded service during a traffic peak, an engineer yanked off a different launch to debug it live, and an uncomfortable call with the customer whose demo had looked great the week before and whose production traffic was now timing out. The CEO's request got its demo. The business paid for it a week later, in a currency nobody had priced into that decision.

The reframe: two queues, two owners

The fix here isn't to start saying no to executives, and it isn't to score the backlog harder. It's to notice that 'what's loud' and 'what's load-bearing' are two different queues, and they need two different owners. What's loud — stakeholder asks, customer escalations, exec requests — belongs to stakeholder management: someone whose actual job is negotiating scope, timeline, and tradeoffs with the person asking. What's load-bearing — capacity assumptions, eval regressions, a config quietly expiring, an agent fleet's unmonitored failure modes — needs an owner whose job is noticing silence, not responding to noise. Put both jobs on the same person under deadline pressure and the loud queue wins every single time, because it's the only one asking to be noticed.

  • ▹Stakeholder-routed work has a name attached, a deadline, and someone who'll follow up if it slips.
  • ▹System-routed work has no name attached, no deadline — and it only follows up by failing.
  • ▹On AI teams specifically, system-routed work includes things like model or version migration side effects, eval suite drift, context-window creep, and agent retry or cost blowups — none of it pages anyone until it's already an incident.
  • ▹The real question isn't 'which ticket scores higher.' It's 'who's accountable for this category of risk even when nobody's asking about it.'

A rule for Day 19

Before you re-rank the backlog, ask one question: who's the arbiter of urgency for this ticket, and is that the right arbiter for this kind of risk? If the answer is 'whoever asked,' fine — for a UI tweak. Dangerous for anything touching capacity, cost, or the correctness of a system running without a human in the loop. Name the arbiter out loud, in planning. If nobody's named, the arbiter defaults to the loudest voice in the room — and in most orgs, that voice isn't sitting anywhere near the infra risk.

yaml
# Ticket metadata we now require before re-ranking anything
ticket:
  requested_by: "loudest-stakeholder | nobody-its-just-failing"
  urgency_arbiter: "stakeholder-manager | reliability-owner"
  displaces: "<ticket id this bumps, named explicitly>"
  system_risk_if_delayed: "none | degraded | load-bearing"

Here's the one concrete thing that changed in how this EM runs planning meetings: no ticket gets re-ranked above another without saying out loud what it's displacing and who owns the risk of that displacement. It's a thirty-second sentence — 'this bumps the autoscaler review, which is load-bearing, and I'm the one accountable if that slips' — but saying it forces the tradeoff into the open instead of letting it happen by default.

Flashcards
Check yourself

Extend your knowledge

  • ▹Audit your current backlog: for the top five tickets you re-ranked this quarter, can you name what each one displaced? If not, that's the gap.
  • ▹Pull your incident postmortems — PagerDuty, Rootly, whatever you run — for the last two quarters. How many trace back to maintenance work that got bumped for a stakeholder ask?
  • ▹If you run inference infrastructure, check whether capacity and autoscaling review triggers automatically on model version changes, or whether it depends on someone remembering.
  • ▹Look at product-ops writing on 'hidden backlogs' and invisible maintenance work — it's a useful public framing that overlaps with the stakeholder-queue vs. system-queue split here.
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Urgency vs system risk” — trade-offs, decisions, or the story behind it.