Skip to content
Managed Services

On-Call Rotations That Don't Burn Out Your Team

Vigil Engineering Team · Jul 20, 2026 · 6 min read
on-callburnoutincident-responseteam-healthmanaged-services
Share:

On-call doesn’t burn people out. Bad on-call burns people out. And most on-call is bad in ways that are entirely structural — which means they’re fixable without hiring anyone or asking your engineers to be more resilient.

The engineer who quits after six months of the pager isn’t weak. They’re rational. They looked at a rotation that woke them up for nothing, gave them no authority to fix anything, and got worse every quarter, and they concluded — correctly — that it wasn’t sustainable. The fix isn’t a wellness Slack channel. It’s the rotation itself.

Burnout Is a Structural Problem, Not a Toughness Problem

Every conversation about on-call burnout eventually reaches “we need to build resilience” or “on-call is just part of the job.” Both are ways of putting the problem back on the person.

The honest version: on-call is sustainable when the pages are real, rare, and actionable, and when the person paged has the context and authority to actually resolve them. It’s unsustainable when it’s noisy, frequent, and disempowering. Those are properties of the system, not the engineer. You can fix a system.

Here are the four structural failures that produce burnout, and what each one actually needs.

Failure One: The Rotation Is Too Small

A three-person rotation means every engineer is on call every third week, forever. There is no recovery period long enough to matter. Life gets planned around the pager permanently — no uninterrupted holiday, no evening you can fully switch off, no week where you’re not either on call, just off call, or about to be on call again.

The rule of thumb: a healthy rotation gives each person roughly three to four weeks off between shifts. That usually means six or more people in the rotation. Below that, you’re not running a rotation — you’re running a slow-motion resignation letter for your most senior people, because they carry the heaviest load and are the most expensive to replace.

At an early-stage company you may simply not have six people who can hold the pager. That’s not a failure of will; it’s a headcount reality. It’s also the clearest signal that on-call is a function you should be augmenting from outside rather than grinding your small team through.

Failure Two: The Pages Aren’t Real

This is the big one. If your on-call engineer is woken three times a night by alerts that self-resolve, the problem was never the rotation length. It’s the signal. A team receiving thousands of alerts a week where only a fraction need action has an on-call experience that would exhaust anyone on any schedule.

Noisy on-call is corrosive in a specific way: it trains the responder to distrust the pager. After enough false alarms, they check alerts slowly and skeptically — and then the one that actually matters gets a late response. Burnout and slower incident response come from the same root. Fix the signal and you fix both at once.

You cannot rotation-schedule your way out of bad alerting. The prerequisite for a humane on-call is a page that’s worth waking up for. Getting there is continuous tuning work, not a one-time cleanup.

Failure Three: The Responder Can’t Actually Fix Anything

Being paged for a problem you have no power to resolve is uniquely demoralising. The engineer gets the alert, diagnoses the issue, and then discovers they lack the access, the runbook, or the authority to remediate — so they escalate, wait, and watch. They carry the stress of the incident without the agency to end it.

Sustainable on-call requires that the person holding the pager can, in the large majority of cases, resolve the incident themselves. That means real access, current runbooks, and the standing to act at 3 AM without waiting for approval from someone who’s asleep. If your on-call is mostly “wake up, diagnose, escalate, wait,” you’ve given people the anxiety of responsibility without the relief of resolution.

Failure Four: Nobody Owns the Rotation Itself

Rotations decay. Someone leaves and the remaining people silently absorb their shifts. Compensating time off gets promised and never taken. Handoffs are a hurried Slack message with no context, so the incoming engineer walks into a live issue blind. There’s no review of who’s actually carrying the load or whether it’s getting heavier.

On-call health is itself a thing that needs an owner — someone tracking rotation size, page volume per shift, time-to-resolve, and whether the burden is trending up or down. On most teams nobody owns this, so it degrades quietly until someone burns out and it becomes visible as a resignation. By then it’s expensive.

What Humane On-Call Actually Requires

Put the four fixes together and the picture is clear. Sustainable on-call needs:

  • Enough people in the rotation for real recovery between shifts — or an external tier carrying the pager so your small team doesn’t have to.
  • Signal, not noise — pages that are real, rare, and actionable, maintained continuously.
  • Real authority to remediate — access, runbooks, and the standing to act without waiting.
  • An owner for the rotation’s health — someone accountable for whether the burden is sustainable, watching the numbers, not just reacting to burnout after it happens.

Notice what none of these are. None of them are “hire more resilient engineers” or “add a meditation benefit.” Burnout is designed into a bad rotation, and it can be designed out.

Or Don’t Carry the Pager at All

The deepest version of the fix is to question whether your engineers should be on the primary rotation in the first place. For a small team, the honest answer is often no. The load of holding a 24/7 pager doesn’t scale down gracefully to five or ten people — the maths simply doesn’t allow for humane recovery windows, and every hour spent on 3 AM triage is an hour not spent building.

This is the case for on-call as a managed function rather than a burden your team absorbs. An external tier holds the primary pager, tunes the signal so pages are real, remediates the routine incidents, and escalates to your team only for the genuine product-level decisions that actually need them. Your engineers get to build, and they still sleep.

Vigil by IOanyT carries the primary pager as a managed service — we hold the rotation, keep the signal clean, remediate at 3 AM, and escalate to your team only when a real product decision needs a human on your side.

Your engineers build. We watch. Nobody loses sleep.

See how outcome ownership works →

Start with a free infrastructure assessment →

V

About the Author

Vigil Engineering Team

See outcome ownership in action

Your infrastructure deserves more than a dashboard. Schedule a demo to see how Vigil handles the monitoring — and the 2 AM pages.