On-call rotation design — what we learned building one
Six rules we landed on after running on-call for a 4-person ops team for two years — primary/secondary handoff, weekend rotation length, the «follow the sun» trap, and how to compensate for it.
- By
- WardenPoint team
- Published
- Jun 4, 2026
- min read
- 3
The six rules
After two years running on-call for a 4-person ops team — and after our share of "Joe is on for the third weekend in a row because the schedule import broke" mistakes — these are the rules we ended up with.
1. Rotation length matches your incident cadence
The default for small teams is "weekly rotation." That's the wrong default. If incidents arrive twice a quarter, weekly rotation means most weeks of on-call are silent and the operator is fresh when an incident finally hits — great. But if incidents arrive twice a week, weekly rotation produces burnout because every shift is a war.
We landed on: rotation length × incident probability ≈ 1 painful shift per cycle. That means:
- 4 people, 2 incidents/week → 3-day rotation
- 4 people, 1 incident/quarter → 2-week rotation
- 8 people, daily noise → 1-day rotation
Tune by feedback, not by industry default.
2. Primary plus secondary, never solo
The temptation with a 4-person team is to skip secondary on-call and just run a 4-week cycle of solos. Don't. Secondary is the difference between "the on-call person is in the bathroom and the page goes unanswered" and "the page goes to the second person 5 minutes later." It also halves the psychological weight — knowing someone else has your back is enormous, even if they almost never need to step in.
3. Weekend rotation is shorter, not the same
Weekends are 2 days, weekdays are 5. If your rotation flips every Monday, you get 5 weekday shifts and 2 weekend shifts — but the WEEKEND shifts cost more, because they cut into family time and the operator can't run errands. Compensate:
- Either rotate weekends separately (Friday hand-off, Monday hand-off)
- Or make weekday rotation 5 days and weekend rotation 2 days as two separate schedules
- Or run a 4-day rotation that "lands" on different days each week and lets the weekend pain rotate naturally
We use the third pattern. It's the only one that doesn't make Mondays into "the long weekend is over" pain points.
4. The "follow the sun" trap
If your team is multi-timezone, the textbook answer is "follow the sun" — pass on-call to whoever's awake. In practice this breaks two ways:
- Context loss at every handoff. "Why is this database alarm firing?" becomes a 20-minute investigation every 8 hours because the next operator doesn't have the running mental model. Mid-incident handoffs are particularly bad.
- Schedule complexity explodes. Now you have three rotations, each in their own timezone, plus an overlap protocol when an incident lasts longer than one operator's shift.
If you're sub-15 people, run a single primary across all timezones, accept that someone has bad-hour shifts, and rotate them around. If you're 30+, follow-the-sun works because the overhead amortizes. In between is the worst zone.
5. Handoff has a script
Every handoff has a 5-minute checklist:
- What's currently firing? (state of the system at handoff)
- What was firing recently and got resolved?
- Any known maintenance windows in the incoming shift?
- Any "ignore for now" alerts that the outgoing person flagged?
- Did the outgoing person eat dinner? (you'd be surprised)
Skipping handoff is the single most expensive mistake. It produces "alarms I don't recognize" and "I thought the other person was on" page failures at the worst possible time.
6. Pay for it
If the company isn't paying for on-call (separate from regular salary), the schedule will degrade. People will swap out of the bad shifts and into the easy ones. Volunteering for weekends will be done by the most senior — until that senior burns out.
The amount doesn't have to be enormous. A token weekend bonus + paid time-off equivalent to the on-call hours is enough. The point is that the schedule has economic weight, so the team optimizes the schedule instead of optimizing who has to suffer.
What WardenPoint helps with
Rotations live in the dashboard as declarative chains. iCal export so the operator can subscribe in Google Calendar / Apple Calendar. Override handling that doesn't require deleting and recreating the whole rotation. And — the part we needed most — the audit log shows who was actually on-call when an alert fired, not just who was supposed to be.
Keep reading
Related posts

Migrating from PagerDuty: a step-by-step playbook
A practical sequence we walk customers through to leave PagerDuty without dropping a single alert — services, schedules, escalation policies, integrations and the dual-fire sanity window.

Telegram VoIP for ops: bypass-DND alerting that recipients don't hate
Why a Telegram VoIP call is the cheapest reliable way to wake on-call without per-message carrier fees, what it costs at scale (spoiler: nothing) and the failure modes you have to wire around.

PHP FFI was the wrong tool: a child-process bridge for Telegram VoIP
How a GLib background thread corrupted the Zend heap in production, why we moved ntgcalls into a child process speaking JSON over pipes, and the two deadlocks plus one 6x latency win on the way to one-second Telegram VoIP calls.