Rotating on-call across timezones
A 6-person team on a 6-week rotation looks fair on paper. Then you actually count what each person absorbed, and one of them did roughly twice the work. This essay is about why that happens, what to measure instead of weeks-of-rotation, and the operational fixes that hold.
Posted 22 May 2026.
The on-call rotation that wasn't fair
Picture a backend team of six engineers across three timezones: two in London, two in New York, two in Bangalore. They run a weekly on-call rotation. Each engineer is on-call one week in six. On paper this looks symmetric: everyone does the same number of weeks per year, everyone is treated identically.
Then you actually count what each person absorbed in a typical week, and the picture changes.
Most pages fire between 9am and 11pm Pacific Time, because that's when the largest fraction of users are awake and using the product. So the Bangalore engineer's on-call week is mostly nighttime alerts. The New York engineer's week is mostly working-hours alerts plus a handful of evening pages. The London engineer's week is in between.
The Bangalore engineer doesn't do fewer pages. They do roughly the same number. They just absorb them at a much higher per-page cost: sleep interruption, partner woken, harder to context-switch back to sleep, longer recovery the next day. Their on-call week is significantly worse than the New York engineer's on-call week, by any honest measure. And the rotation policy, by treating weeks-of-rotation as the unit, makes this invisible.
This is what most distributed on-call rotations look like. The unfairness isn't in the rotation; it's in the choice of what to measure.
Why weeks-of-rotation is the wrong unit
Weeks-of-rotation answers the question “how often does each person take a turn?” That's the wrong question for fairness.
The right question is “how much actual cost does each person absorb?” The answer to that is some function of:
- How many pages they take
- What time of day, in their local time, each page arrives
- How disruptive each page was (sleep-window vs working-hours vs evening)
- Whether weekends or local holidays were involved
- How long the rotation week was (some on-call rotations include a handover day or two)
Of those, the only one captured by “weeks of rotation” is the last bullet, and even that imperfectly. The other variables, which collectively dominate the actual experience, are completely unrepresented.
A team that runs a fair weeks-of-rotation policy is doing the right thing at the surface level, but optimising a metric that almost nothing of consequence depends on. That's not better than no policy. It's slightly worse, because the team can point to the policy and feel that fairness has been addressed.
What to measure instead
The single most useful unit is minutes paged outside working hours, per person, per quarter.
Working hours are defined per-person, in their own local time. “Outside working hours” includes evenings, weekends, and sleep windows, with sleep weighted highest. The aggregate is computed at quarter-level (not week-level) because weekly variance is too high for a fair signal, and quarter-level is short enough that you can act on it before the next quarter starts.
Why this unit?
It captures the asymmetry weeks-of-rotation misses. The Bangalore engineer's on-call week generates many more out-of-hours minutes than the New York engineer's, even when they take the same number of pages. The metric makes that visible.
It's quantifiable from the data your paging tool already produces. PagerDuty, Opsgenie, Splunk On-Call, FireHydrant: all of them log timestamps. Mapping those timestamps to local time per engineer and bucketing into in-hours / out-of-hours is a script. You don't need a new tool to start.
It corresponds to something real. An engineer's actual experience of on-call is more strongly correlated with out-of-hours minutes than with any other single variable. Pages during the workday are mostly free; pages at 3am are mostly expensive. The metric pays attention to the costly category.
It's actionable. Once you can see that Person A is at 240 out-of-hours minutes per quarter and Person B is at 80, you can have a specific conversation about how to bring those closer.
A second useful unit is sleep-window pages per quarter. This is a subset of the first, isolated because sleep-window pages have an outsized cost relative to evening or weekend pages. A 3am page costs more than a 9pm page; a manager who only looks at the aggregate misses the sleep-window concentration. Tracking sleep-window pages separately makes the worst category legible.
Why follow-the-sun isn't a fix
The standard response to “the Bangalore engineer is bearing too many night pages” is “let's do follow-the-sun: the Bangalore engineer only takes pages during their working hours, and we hand off to the New York engineer at end-of-day.” This sounds clean and is usually wrong.
Three problems with follow-the-sun.
Handoff cost. A handoff happens at the boundary between two engineers' working hours. Each handoff requires the outgoing engineer to brief the incoming engineer on whatever's in-flight. For a 24/7 product with three regional engineers, that's three handoffs per day, every day. Each handoff takes 10 to 30 minutes. Three handoffs per day × 30 minutes × 5 weekdays = 7.5 person-hours per week of pure handoff overhead. For a six-person team, that's >10% of the on-call engineers' working time.
Mid-incident handoff is much worse than scheduled handoff. An incident in progress at 5pm London is messy to hand off to New York. The London engineer has the context; the New York engineer has the duty. Either the London engineer stays on past their working hours (which defeats the point), or the incident gets handed off mid-flight (which is risky and expensive). Follow-the-sun makes handoff cost a function of incident timing, which you can't control.
It re-introduces unfairness in a different shape. If pages cluster around peak-traffic Pacific hours, the New York engineer's follow-the-sun shift includes most of the day's pages. The Bangalore engineer's shift includes few. Now Bangalore is doing too little on-call work and New York is doing too much: a different unfairness, but still unfair.
Follow-the-sun is the right pattern for some teams. Specifically: teams where pages are roughly uniformly distributed across the day, where handoff cost is low (e.g. status-monitoring-only on-call, not incident-response), and where the engineers strongly prefer to never be paged out of hours. It is not a general fix for the underlying fairness problem in night-heavy paging.
Holidays and the “every Christmas” problem
The other asymmetry that on-call rotations rarely address: holidays distribute very unequally across a globally-distributed team.
The US-based engineer's holiday week (Thanksgiving, plus the late-December / New Year stretch) lines up with the team's lightest traffic week of the year. The Indian engineer's holiday week (Diwali, regional holidays) often doesn't line up with any traffic dip; it just happens to be on-call somebody else's job that week. The result: the Indian engineer typically takes more on-call holidays per year than the American engineer, even on a strict weeks-of-rotation policy.
The fix is to track on-call holidays per person per year as a separate variable, and to manually balance it. Specifically: if the rotation would land an engineer on-call during one of their local holidays, swap with someone who'd otherwise be on-call during a non-holiday week. This requires manual intervention; no automated rotation policy gets it right by default, because the rotation engine doesn't know which days are holidays in which jurisdictions.
A team that doesn't do this swap finds, after two or three years, that one engineer has done all the Diwali on-calls and another has done all the Christmas off-calls. That's a fairness asymmetry that compounds over years rather than weeks, and it's the kind of thing that quietly drives engineers to leave.
What about tools? PagerDuty, Opsgenie, FireHydrant
The standard on-call tools have improved here. Some of them report per-engineer page statistics by hour of day. PagerDuty's analytics dashboard can show this if you set it up. Opsgenie has similar reporting. The signal exists; teams just rarely look at it.
Two honest limits.
First, none of these tools compute fairness across the dimensions above by default. They show you the raw data; the team has to do the work of mapping page timestamps to per-engineer local times and computing the asymmetry. For a 6 to 10 person team it's a 1 to 2 hour quarterly exercise. Nobody does it consistently, for the same reasons covered in the early-meeting essay: the data is invisible until somebody does the work to make it visible.
Second, the on-call data only captures actual pages. It doesn't capture “I was on-call all weekend and nothing fired, but I was tethered to my laptop the whole time.” The cost of being on-call is more than the cost of the pages that fire. Most teams haven't found a good way to measure the standby cost; the closest proxies are explicit on-call standby compensation (which some teams pay) or a strict policy that on-call rotations don't fall on holidays/special occasions.
The operational fixes that actually hold
The reasons above suggest four operational moves, in priority order.
1. Measure out-of-hours minutes per engineer, per quarter. Export the data from your paging tool. Map timestamps to each engineer's local time. Bucket into working-hours / evening / weekend / sleep-window. Look at the asymmetries. Do this every quarter, not every week. The exercise is what produces the fairness data; the rest follows from having the data.
2. Address sleep-window concentration first. If the data shows one engineer absorbing many more sleep-window pages than the others, that's the highest-priority asymmetry to address. The fix isn't usually changing the rotation: it's reducing the absolute number of sleep-window pages by fixing the underlying alert. Many sleep-window pages are spurious or non-urgent and should be either silenced or batched into business-hours review.
3. Manually rebalance holidays. Once a year, look at next year's rotation against everyone's local holiday calendar. Swap out engineers from local holiday weeks. This is 30 minutes of work per year and removes a slowly-compounding fairness problem.
4. Make the data visible to the whole team. The biggest single change a team can make is not any specific operational fix: it's making the quarterly out-of-hours-minutes-per-person number visible to everyone on the team. Once that number exists publicly, the conversation about what to do is much easier, because the asymmetries are no longer hidden inside one engineer's experience.
The order matters. Without measurement (move 1), the other moves can't be targeted. Without sleep-window prioritisation (move 2), you can rebalance other categories and still leave the worst category untouched. Without holiday rebalancing (move 3), the year-on-year asymmetry compounds. Without visibility (move 4), the team negotiates blind and the rotation drifts back to its original state.
The thing you can do this week
If your team isn't ready to adopt anything new, the minimum-viable version of this is one CSV export and 90 minutes:
- Export the last 90 days of pages from your paging tool, with timestamps in UTC.
- For each engineer, convert each page timestamp to their local time.
- Bucket each page into working-hours (defined per engineer), evening (6pm to 11pm local), sleep-window (11pm to 7am local), or weekend.
- Sum the time spent in each bucket per engineer. (Estimate 15 minutes per page if you don't have actual incident-duration data; precision is less important than visibility.)
- Compute out-of-hours minutes per engineer.
- Look at the spread.
Almost every team that does this discovers a spread of 2× or more between the engineer at the top and the engineer at the bottom. The conversation about what to do becomes much easier with the numbers in hand. The hardest part is doing the exercise once; after the first time, it's a 30-minute job to repeat each quarter.
The thing you can do this week, with FairlyRemote
FairlyRemote was built primarily for meeting fairness rather than paging fairness, but the underlying model is the same: track out-of-hours load per person, surface the asymmetries to the team, let the team negotiate from data. The sample team shows what per-person out-of-hours load looks like in a UI. The on-call equivalent isn't shipped yet but follows the same shape.
If on-call rotation fairness is your immediate problem, your paging tool's analytics are probably closer to the data you need. The recommended sequence is: use the paging tool's data for on-call analysis, use a fairness-aware tool like FairlyRemote for the meeting-load side, and treat the two as related but distinct workstreams.
A note on what this isn't
This essay is about a structural problem in distributed on-call, not about individual blame or evaluation. The engineer absorbing more out-of-hours minutes isn't “more dedicated”; the one absorbing fewer isn't “less committed.” The data is descriptive, not evaluative. The whole point of measuring it is to redistribute structural cost, not to rank engineers.
The data is also not a measure of engineer wellness or productivity. Many factors affect on-call experience that the metric doesn't capture: page severity, follow-up cost, time-of-week effects, partner/family context. The metric is a starting point for conversation, not a verdict.
See out-of-hours load on a sample team
You and four teammates across four timezones, with a week of meetings already in place. No signup, no card.