What is fairness scoring?
An explainer of the four-dimension model that powers FairlyRemote, why scheduling tools that don't model sleep produce unfair meeting times, and what the numbers actually predict (and don't).
Posted 4 May 2026.
Fairness scoring measures how much a meeting time affects each person, not just whether they are free. Each candidate hour is scored against the member's local working hours, evening time and sleep window, then weighed against what they have already absorbed recently. The result is one impact label per slot, from Low to High, so a team can compare times by who carries them rather than by who happens to be available. The arithmetic stays under the hood: FairlyRemote shows people words and counts, never a personal score out of 100.
The problem fairness scoring is trying to solve
If you've worked on a distributed team across 3+ timezones, you already know the feeling. The team agrees to rotate the awkward standup time. Three months later, somehow, the standup is still at 8am London, meaning Tokyo joins at 5pm and San Francisco joins at midnight. Nobody decided this; it just happened. The meeting has drifted, and the impact is landing on whoever has the fewest scheduling conflicts to push back with.
This drift is the fairness problem, and it is structural rather than anybody's fault: the five reasons a recurring meeting keeps landing on the same person hold on almost every distributed team. Most scheduling tools all answer the same question: is this slot free? They don't answer the more useful question: given everyone's working hours, sleep schedule, social commitments, and what they've already given up this month, where would this meeting have the least impact?
Fairness scoring is the name for an algorithm that answers the second question.
How is meeting fairness measured?
There's no perfect single metric for “is this meeting fair.” Real-world meeting fairness has four components that move independently, so a useful model scores each independently and rolls them up.
1. Accommodation
Hours outside a member's normal working window. Early mornings, late nights, lunch hour. The scheduler maps each local hour to a point cost: business hours score 0, lunch (12-1pm local) scores 0.5 (mildly inconvenient, not a sacrifice), early morning (6-9am for a Standard chronotype) scores 3, very early (before 6am) scores 5.
The Accommodation dimension is the easiest part to model and the part most schedulers handle correctly. The remaining three are where most schedulers fall down.
2. Flexibility
Each person has a different capacity to absorb scheduling inconvenience. A member with three young kids has a more rigid evening schedule than a single member without dependants. A member with a chronic health condition might have a non-negotiable mid-afternoon nap. A model that treats everyone identically gets these wrong.
The Flexibility dimension is operationalised as a self-declared profile: Strict (low capacity to absorb inconvenience), Standard (the default), or Flexible (more capacity to absorb extra). The profile sets that member's monthly accommodation budget, which we'll come back to.
3. Social time
Evening hours 6-9pm local are a category most schedulers miss entirely. They're outside working hours, but they're also outside “genuine off-hours”. They're the family-dinner / commute-home / personal-life window. A meeting at 7pm local isn't sleep-deprivation, but it is a real impact most people don't volunteer for repeatedly.
FairlyRemote treats this band as Reluctant: scheduled-able, but with a 1.5× multiplier into the monthly budget. Repeat meetings here show up loudly in the per-member record.
4. Rest quality
This is the dimension that surprises most buyers. Meetings during sleep windows are not the same as meetings outside working hours. They're a different category of ask, and a useful model treats them separately.
FairlyRemote's deep-sleep window is 1-4am local, and the cost is 10 points (the maximum). The window doesn't shift by chronotype (biology is biology), and the model treats it as a hard limit. Meetings here are not scheduled by accident.
Light sleep (11pm-1am, 4-5am) is 8 points. Scheduled-able, but flagged as the most expensive option after deep sleep.
Why don't scheduling tools measure fairness?
Mostly because they were built for a different job: external booking, one-off polls, or finding a free hour inside one office. We set out which tool fits which job in scheduling tools for distributed teams, compared.
Two reasons.
First, the inputs are harder. Modelling sleep requires per-member chronotype declarations and a published curve. Modelling social time requires accepting that 7pm is different from 11am even though both are outside the 9-5. Most schedulers were originally built for one-time external bookings, where the user explicitly chooses their available slots and the model never has to reason about “was this fair?”
Second, the output is harder to display. A scheduler that surfaces “this slot is at 3am for one member” has to design a UI that makes it clear, makes it actionable, and doesn't punish the booking-flow ergonomics. Most schedulers chose the simpler thing: just show available slots, let the user reason about cost themselves. That's fine for an external booking link to a freelancer; it's a poor fit for an internal recurring team meeting.
Monthly accommodation budgets
A single meeting being scored fairly is necessary but not sufficient. The deeper question is whether cumulative meeting load over a month is fairly distributed.
FairlyRemote tracks each member's accommodated minutes against a monthly budget: absolute, not ratio-based. The defaults are 10.5 / 15 / 22.5 points/month for Strict / Standard / Flexible profiles. Each scheduled meeting deducts from the relevant member's budget. The dashboard reports this in words rather than numbers: who is carrying more, who has room, and how many meetings landed outside each person's hours.
The crucial split: voluntary vs structural components. The structural component is the unavoidable timezone tax: the cost of being the eastmost or westmost person on a team that has to meet sometimes. The voluntary component is the choices people make to absorb extra inconvenience for the team's benefit. Tracking these separately means a person who lives furthest east isn't penalised forever for living there; their structural component is identified and the voluntary part is what drives the weighted-vote multiplier.
Decay: why old accommodations don't haunt anyone
If fairness history accumulates without forgetting, two bad things happen. People who took the awkward calls last year carry an inflated voice in this year's votes. People who took them yesterday carry the same voice as people who took them six months ago. Both are wrong.
FairlyRemote's history decays exponentially at 0.7 per week, a half-life of about 13.6 days. Today's meeting weighs 1.0; one from a week ago weighs 0.70; from two weeks ago, 0.49; from four weeks ago, 0.24; from 90 days ago, roughly 1 per cent. The team's recent record dominates current decisions without forgetting last month's pattern, and the record looks back 90 days so a long run of early calls cannot quietly age out of the system's memory.
Where the history actually lands
The history does not weight anybody's vote. Every ballot counts the same. What it changes is the ranking the team is choosing from: each candidate slot is scored with the fairness history already inside it, so a 6am that keeps falling on the same person ranks below one that spreads the impact. By the time a shortlist reaches the team, the accommodation record has already done its work.
That is a deliberate design choice. Weighting ballots makes a vote harder to reason about (why did my choice lose to fewer people?) and invites gaming. Ranking the options instead keeps the decision legible: the team picks from a list that is already fair, and the pick is plainly the team's.
Individual ballots aren't visible to peers: a database access rule restricts each ballot row to the person who cast it, and the proposal UI shows only counts. Separately, a board admin can set the fairness dashboard to identity-shielded, where each person's load appears under a stable label like “Member A1B2C3” instead of a name. That shields identities in the UI, but it isn't proof against someone determined to work out who is who.
Why you never see the number
Every fairness model produces a number somewhere. The design question is whether to show it to people, and we decided not to.
A visible score out of 100 turns a scheduling aid into a leaderboard. People compare, they ask why they are a 47 and their colleague is a 72, and a tool meant to spread the early calls starts reading as a performance rating. Nothing about the arithmetic justifies that weight.
So the number stays where it is useful: inside the ranking. What people see is what they can act on. Their own view says how many meetings they attended, how many landed outside their hours, how many touched their evening or sleep window, and which specific meetings those were. The group view says who is carrying more, who has room, and nothing else about anybody. No rings, no gauges, no personal score.
What the model does not predict
This is the part most explainers gloss over. Fairness scoring is a heuristic suggestion for scheduling decisions. It is explicitly not:
- An employment-decision input. Carrying more early calls one month doesn't mean a member is unhappy, underperforming, or about to quit.
- A wellness or clinical metric. We don’t measure sleep quality; we model when meetings would intersect typical sleep windows. (Full scope on the about page.)
- A productivity score. Meeting accommodation is not a proxy for value.
- A definitive answer to “is this fair?” It's an input. The team makes the actual decision.
The model is auditable on purpose. If a teammate asks “why does it say I'm carrying more?”, the dashboard's What's behind this list names the actual meetings that counted, with the local time each one landed at. That transparency is the difference between a fairness tool and a black-box productivity-monitoring tool.
How fairness scoring fails
Three honest limits:
It can't see meetings outside the system. If a team also meets in Slack huddles, drop-by Zoom calls, or hallway conversations, those don't get scored. The accuracy depends on the team logging their meetings, or syncing a calendar that reflects them.
It treats time-as-cost, not time-as-value. A meeting that's awkward but high-value (a customer escalation) has the same fairness cost as one that's awkward and low-value (a status meeting that should've been an email). The score doesn't help you decide which meetings should exist; it only helps with timing.
It assumes self-reported chronotypes are accurate. If someone declares Standard but actually wakes at 5am, the model is wrong about them. The defaults are conservative on purpose; in practice, members tend to refine their chronotype within a few weeks.
Where to go from here
If you want the scoring rules in detail, see the how it works reference page.
If you want to see it on a real team, the sample team loads instantly with four members across four timezones, a fortnight of real meeting history already in the ledger, no signup.
If you're evaluating it for your team, the how it works page walks through the full flow.
Try fairness scoring on a real team
You and four teammates across four timezones, with a week of meetings already in place. No signup, no card.