Volume Per Labor Hour Measures How Busy a Site Is, Not How Well It Runs
Jeff Bugbee · August 17, 2026
- Workforce analytics
- Labor standards
- Labor analytics
- UKG Pro WFM
Every multi-site operation I’ve worked in has a scorecard, and somewhere on it is volume per labor hour. Units per hour, transactions per hour, cases per hour, visits per hour — the noun changes, the arithmetic doesn’t. Sites get ranked on it, regional calls open with it, and in a few places money attaches to it.
And in every one of those operations there is at least one site manager near the bottom of that ranking who is, in fact, running a tight operation. They know it. Their regional director half-knows it. Nobody can prove it, because the metric that says otherwise is the metric everyone agreed to use.
The problem isn’t that the number is wrong. It’s that it’s three numbers wearing a trench coat.
What is volume per labor hour actually measuring?
Look at the ratio honestly. The numerator is demand — customers who walked in, cases the DC was asked to move, transactions the POS rang. A site manager influences that at the margins and controls almost none of it. The denominator is hours, which they do control, but only partly: the schedule has to cover operating hours whether or not anyone shows up.
So a metric sold as “productivity” mostly measures how much work arrived per hour the doors were open. Rank sites on it and you have, to a first approximation, ranked them by demand density. High-traffic locations look excellent. Low-traffic locations look lazy. Neither conclusion is supported by the data.
Why does the fixed coverage floor break the comparison?
Because every site carries hours that exist regardless of volume. Someone opens. Someone closes. Someone counts the drawer. Policy or safety says at least two people are on the floor at all times. Meal and rest coverage has to come from somewhere. Those hours are a function of operating hours, not of demand.
Here is the arithmetic, with illustrative numbers to make the shape visible:
- Site A: 12,000 units, 1,000 hours → 12.0 units per hour
- Site B: 4,000 units, 500 hours → 8.0 units per hour
Site B looks 33% less productive, and that is exactly how it will be described on the call. Now subtract a coverage floor of 180 hours per week from each — the same at both sites, because both are open the same hours:
- Site A: 12,000 ÷ 820 variable hours → 14.6
- Site B: 4,000 ÷ 320 variable hours → 12.5
The gap drops from 33% to 14%. More than half of the apparent performance difference was never performance; it was a fixed base spread across a smaller denominator. The smaller the site, the more the coverage floor dominates — which is why low-volume locations sit at the bottom of these rankings year after year and no amount of coaching moves them.
Isn’t that what earned hours are for?
Partly — and this is the second thing hiding in the ratio. Raw volume treats every unit as equivalent work, and it never is. A basket of six items and a basket of sixty are one transaction each. A full-case pick and an each-pick are one line each. Delivery, curbside, and in-store consume very different labor minutes per order.
Weighting volume by earned minutes fixes this: convert each volume stream through its own time factor so the denominator of your comparison is work, not events. That is what a labor standards model is for, and why I’d rather rebuild those standards from a client’s own process timestamps than inherit documentation written by someone who left in 2019.
But even a perfectly weighted earned-hours number still fuses two capabilities that call for completely different interventions.
The two numbers that should replace it
Process speed — earned-to-actual on variable work. Of the hours spent on demand-driven tasks, how many did the standard say the work should take? This answers are we performing the work efficiently, and the levers are method, training, layout, equipment.
Staffing fit — scheduled and worked hours against the demand curve, at interval. Did the hours land when the work arrived? This answers are we deploying the hours well, and the levers are forecast quality, schedule generation, shift structure, and whether managers trust the auto-scheduler enough to leave it alone.
A site can be genuinely excellent at the first and terrible at the second. Volume per labor hour averages them into a single figure that tells you a site has a problem without telling you which problem, and the coaching that follows is a guess.
The interval point deserves its own sentence, because daily numbers conceal it entirely. A site can be staffed correctly on a day total and mis-staffed in every hour of that day — overstaffed all morning, drowning from 4pm — and the daily ratio still looks fine. A productivity metric computed daily cannot see the most common and most expensive staffing failure there is.
What this looked like in the field
A national specialty retailer wanted labor standard forecasting stood up and, underneath it, an honest read on which stores were actually underperforming. The existing ranking was a volume-per-hour league table that had been stable for years — the first clue, since operations rarely stay that consistent unless the metric is measuring something structural.
The work was mostly modeling decisions, not analysis: designate which hours are coverage versus variable, derive time factors per volume driver from the stores’ own process timestamps, weight demand by earned minutes, then evaluate each store as a residual against comparable peers rather than the network average.
The ranking changed materially, and more usefully it split into two rankings — stores whose process was slow, and stores whose hours were in the wrong places. Those are different problems with different owners, and the old single number had been quietly blending them for years. This is the same discipline as keeping validated savings separate from modeled savings: the value is in refusing to let one figure carry two claims.
When NOT to bother
- Single-site, or sites that are genuinely identical. If volume, operating hours, and format are comparable, the raw ratio is a fine relative signal. The decomposition earns its cost when the network is heterogeneous.
- You’re trending one site against itself. That already controls for the coverage floor and most of the mix, as long as the format hasn’t changed. Cross-sectional comparison is where the metric breaks.
- Your standards aren’t derived yet. Earned-minute weighting requires time factors you trust. Applying vendor benchmark standards to make the metric “fairer” just swaps one unexamined assumption for another.
- Nobody will change what the scorecard measures. If the ranking is tied to incentive plans that can’t move this cycle, build the decomposition as a diagnostic and don’t pretend it’s a replacement. A better metric nobody acts on is an expensive report.
This metric persists not because people are naive about it, but because it’s the only one available from data everyone already has. Constructing the honest version requires deciding which hours are coverage, what a unit of work is worth, and which sites are genuinely comparable. Those decisions are specific to your operation, which is precisely why no product ships them for you.
If your site ranking has been suspiciously stable for years, tell me what’s in the numerator — I’ll tell you straight whether you’re measuring performance or measuring demand.
Common questions
- What is wrong with using volume per labor hour as a productivity metric?
- It fuses three different things into one ratio: how much demand arrived, how much fixed coverage the site has to run regardless of demand, and how efficiently the work was actually performed. Only the third is performance. Because the fixed coverage floor is roughly constant while volume varies, low-volume sites score structurally worse — that part of the gap is arithmetic, not effort.
- How do you measure site productivity fairly across different volumes?
- Subtract the coverage floor before you compute anything, and weight volume by earned minutes rather than counting raw units, so a transaction at one site is comparable to a transaction at another. Then report two numbers: earned-to-actual hours on variable work for process speed, and interval-level staffing fit against demand. Rank sites on residual performance against comparable peers, not on the raw ratio.
- Can UKG Pro WFM produce this out of the box?
- Pro WFM holds the ingredients — punches, schedules, job transfers, and, with labor standard forecasting configured, earned hours against volume drivers. What standard reporting will not do for you is decide which of your hours are coverage versus variable, how to weight a mixed volume stream, or which peer group a site belongs in. Those are modeling decisions about your operation, so they get built.