In an illustrative week of high-volume hiring, roughly 3,200 applications are never opened by anyone. They are not rejected, because no decision is taken on them at all: a screening team gets through about 800 files before the week runs out, and which 800 is settled by an application timestamp. The queue holds cashiers, agents, pickers, and baristas, and in retail and BPO hiring it also holds customers, because the person the process ignores on Tuesday may be standing at the counter on Saturday. Every open seat, meanwhile, is an uncovered shift somewhere real, so the funnel is timed in days, and the margin a good hire adds is thin enough that no stage of the process can afford to be artisanal.
Hourly selection is usually run as an administrative problem: move the queue, cover the shifts, keep the requisitions green. It is better understood as selection science under a throughput constraint. The constructs that predict who serves customers well and who is still on the roster a quarter later are measurable here the way they are anywhere else in applied psychology; what volume changes is that the funnel must capture them at a pace, and on devices, that no traditional testing program ever planned for, while treating applicants the way a consumer brand treats customers. Fast, measured, and civil: any two of those are easy.
The scarce input in high-volume hiring is screening time
A week of hourly hiring has a shape you can sketch from memory: a few thousand applications arriving against a few hundred hourly seats, most of them submitted from a phone, many in the gap after one shift ends and before the next begins. Volume does not change what predicts performance in those seats; it changes which stage of the process runs out first. Treat the funnel as a pipeline in which every stage has a weekly capacity, and it becomes obvious where the throughput actually stops. Figure 1 runs that idea over an illustrative week, and the number that binds everything else is the 800 files a screening team can genuinely read against the 4,000 that arrive, a ceiling set by how many hours recruiters and hourly managers have left after running a floor. The pipe narrows at the top, where the reading happens, before selectivity ever gets its turn.
Four applications in five therefore land in a residue nobody works. Triage by proxy fills the gap: first come, first read; keyword matches on availability; distance from the site; a familiar employer name in a work history. Under load the most consequential sort in the funnel is performed by arrival order, and speed of response gets treated as merit because it is the one variable the team can see for everyone.
A pipeline has one binding constraint at a time, and effort spent anywhere else is absorbed by it. Reading capacity moves only with recruiter headcount, and it moves linearly: twice the files read takes roughly twice the hours, in an operation whose margin per hire is thin by construction. Assessment capacity does not behave that way, because delivering a battery to the four-thousandth applicant asks no more of the organization than delivering it to the fortieth. Two adjacent stages of one funnel obey different laws of growth, and at volume that difference decides which of them the program should be designed around.
The unread portion of the queue is not inert. Applications age, candidates take other work, and the residue rolls forward behind a fresh inflow, so the practical rejection is administered by the calendar. No record marks those files as declined, which leaves the largest single sorting event in the funnel with no evidence that it occurred and nothing anyone could review. The features that do get read were adopted because a human can process them in seconds. Their power to forecast who will still be serving customers in October was never part of the choice. Legibility becomes the selection criterion by default, and it holds that job for as long as reading is the constraint.
The reflex fix, adding interviews, moves the shortage without relieving it, since interview hours are the scarcest reading of all, and what makes interviewing defensible at any volume is structure, a case the seasonal cousin of this problem has already worked through. The boundary matters: campus hiring at scale is a compressed season with a deadline and an off-season in which to recalibrate, while hourly hiring is rolling and always-on, with no annual pause in which to rebuild the funnel. Every redesign therefore ships onto a moving funnel, with that week's applicants passing through the version being replaced.
One stage in Figure 1 matches the inflow: the battery. Delivered from software to whatever device the applicant already holds, an assessment measures the whole queue in the time a screening team reads a fifth of it, and it applies one standard to everyone rather than a proxy to most. Moving the first sort onto the applicant's phone, though, is a psychometric decision before it is a logistical one, and the evidence on what survives the move turns out to be unusually clear.
The phone is the venue, and the evidence splits by construct
In frontline hiring, the phone is where the applicant pool lives. Candidates increasingly complete assessments on mobile devices (King, Ryan, Kantrowitz, Grelle, & Dainis, 2015), and an hourly applicant with no laptop at home is the ordinary case. Any assessment strategy for these roles is therefore a mobile assessment strategy, whatever the design documents assumed, and the useful question is which measurements survive the venue.
The reassuring half of the evidence covers the measures an hourly screen leans on most. Illingworth, Morelli, Scott, and Boyd (2015), studying internet-based, unproctored assessments in operational use, tested whether non-cognitive measures keep their measurement properties when candidates complete them on mobile devices. Equivalence held at every level psychometricians test for it: configural, metric, scalar, and latent-mean, meaning phone and desktop versions behaved as one instrument, from the structure of what it measures down to the score levels it produces. For personality and judgment, the device is a venue, and the venue is psychometrically safe.
King and colleagues (2015) reached the matching conclusion from the test-taker's side of the screen: the equivalence evidence for non-cognitive measures held up, and candidate reactions to mobile testing were comparable to reactions on other devices. In a funnel where the applicant is also a customer, that second finding carries real weight: the phone sitting does not degrade the experience it delivers the measurement through.
Equivalence at that depth is a stronger license than it sounds. The levels are nested, each a stricter condition than the last, so clearing the strictest of them means a phone score and a desktop score can enter one table: a single norm, a single cut, a single report, with the device recorded as a delivery detail and never as a factor in interpretation. Without that evidence a program drifts toward workarounds nobody wants to defend, separate thresholds by device or a desktop requirement that filters the pool by what people happen to own. An hourly battery's untimed core is free of that problem.
Timed cognitive testing behaves differently, and the difference is well mapped. Arthur, Doverspike, Muñoz, Taylor, and Carr (2014), examining high-stakes, remotely delivered testing in operational use, found that candidates who tested on mobile devices scored lower on cognitive measures than candidates on other devices, with mobile sittings already a meaningful share of test traffic. The follow-up framework from Arthur, Keiser, and Doverspike (2018) places the deficit precisely: it shows up in operational settings, while randomized, low-stakes laboratory studies fail to reproduce it. That pattern points away from the hardware alone and toward a pairing: self-selection, who ends up taking a timed test on a phone and under what circumstances, and the device's information-processing demands, the scrolling, the cramped screen, and the input burden that a running clock converts into lost points.
The mechanisms fail in different ways, and the difference is operational. Self-selection means that in live hiring nobody assigns a device: the candidate picks one, and the picking is bound up with circumstance, with where the person is sitting, what the household owns, how much of the evening is free, and how badly this particular job is wanted tonight. A mobile-versus-desktop comparison drawn from operational data is therefore in part a comparison of unlike groups. Device demands work on the item as delivered, on a stem that does not fit the screen, a response set reached by scrolling, and a clock running through both. Randomizing the device in a low-stakes study strips out self-selection and softens the delivery burden, and the laboratory version of the question comes back clean. Neither reading cancels the other, and neither is repaired by better handsets.
Figure 2 sorts the evidence by construct, which is the form a delivery policy can act on. Untimed, non-cognitive measurement can go to the whole queue on any device the applicant chooses; the equivalence findings license exactly that. Timed ability testing, where the role genuinely needs it, gets device-aware design: a stated device policy, a timed section offered at a moment the candidate can control, or ability testing attached to a later, controlled stage the way verification designs already attach a confirmatory sitting. Blanket phone bans fail the population; pretending the deficit away fails the measure; device-aware design does neither.
Construct-specific equivalence is also an access finding, read from the other direction. The device a candidate reaches for is tangled up with circumstance, so a funnel that in practice demands a desktop and a free hour excludes on logistics before it ever measures anything. A design that runs the untimed core anywhere and confines the clock to controlled moments takes the venue as it finds it. The telemetry has to keep up as well: completion and drop-off tracked by device, so that a delivery failure surfaces as the delivery problem it is and never masquerades as low ability.
Short by design, not by amputation
Everything about the venue argues for brevity. Sittings happen on breaks, attention is rationed between obligations, and long forms get abandoned at rates every funnel operator has watched in their own drop-off numbers. The psychometric literature grants the brevity and disputes the usual method of getting it. Kruyen, Emons, and Sijtsma (2013), reviewing what happens when fixed tests are cut down, found that shortening degrades reliability, the repeatability of the score, and with it decision consistency, the likelihood that a pass-fail call would come out identically on another sitting, and that short forms are routinely used for decisions their length cannot support.
At hiring stakes the degradation lands where an hourly funnel notices it least: near the cutoff, among candidates whose scores sit inside one another's error bands, which at thousands of applications a week is a crowd. A cut-down form keeps making the old decisions with less instrument behind them, and nothing in the score report announces the change.
Items are not interchangeable filler, which is the assumption a cut-down form runs on. A scale carries its measurement unevenly across the range it covers, so trimming for time removes information that was never evenly distributed to begin with, and the reliability given up is not proportional to the minutes saved. Kruyen and colleagues found short forms being put, routinely, to decisions their length could not support. That is the condition an hourly funnel reaches fastest, since the pressure to shorten runs highest exactly where the volume of decisions runs highest.
The defensible route to a short sitting is adaptive delivery, which shortens by selecting items for the candidate instead of administering one truncated form to everyone; the psychometric case is the subject of the companion briefing on adaptive testing in hiring. The equivalence evidence has already drawn the short battery's map: the untimed, non-cognitive core, personality and judgment measured at proper depth, runs on any device, and a timed reasoning section, where the role calls for it, waits for a controlled moment. An hourly hiring assessment that fits inside a break does not have to be a worse instrument than the hour-long form it replaced; it has to be built differently.
Brevity has to be taken from the right place. A defensible short battery drops constructs the role does not use before it thins the scales it keeps, so what survives is measured at a depth that can carry a decision. Its budget covers the whole sitting: instructions, device setup, and the number of screens between the first tap and the last all draw on the one break the candidate has. And it keeps a light check on how the responses were produced, because an unproctored sitting on a device nobody controls is the condition verification designs were written for.
Speed-to-fill wins the week and loses the quarter
Hourly funnels are managed on time-to-fill for an understandable reason: a vacancy is visible today, on a schedule with a hole in it, while churn arrives later, as a monthly rate on a dashboard nobody staffs against. So the funnel optimizes the visible number, in labor markets where frontline churn is chronically high, and every reopened seat re-enters the queue at the top of Figure 1. From inside the week that produced it, filling fast and refilling constantly can look like operational excellence.
The criterion an hourly program should answer to is ninety-day survival alongside service behavior: whether the person the funnel chose is still on the roster, and serving customers well, a quarter in. Early quitting is a selection outcome with measurable pre-hire signals, and the evidence for that claim, including which measures carry it, is the subject of the companion briefing on early-attrition prediction. An hourly funnel is where those signals matter at the largest scale, because fills are so frequent that any per-hire improvement in survival repeats itself weekly.
Ninety days works as a criterion for reasons beyond convention. It clears the initial training window, so the survivors are people the operation has actually watched on the floor, and it lands back inside the same operating cycle that produced the cohort, which means the screen can be revised while it still resembles the screen under review. Longer criteria make better science and worse feedback, and high-volume hiring cannot wait a year to learn something it needs on Monday.
Survival also reads better as a curve than as a rate. Exits clustered before training completes point at scheduling, onboarding, and the accuracy of what the funnel promised about hours; exits after it point back at the screen. Survival on its own rewards inertia, so service behavior belongs beside it, since a person who stays and serves badly is not the outcome the program was built for. Volume repays the operator here: the property that makes the funnel hard to run also delivers cohorts large enough, month after month, to check the screen against its own claims.
Churn's claim on the P&L is structural, and Figure 3 traces it as the loop it is: an open seat forces coverage; a replacement absorbs sourcing, onboarding, and training hours; an early exit reopens the seat with the training never repaid; and the vacancy lands back in the funnel. No single line item carries the whole loop, which is part of how it escapes review. The values are deliberately absent here: they belong to the companion briefing on the cost of a bad hire, which works the numbers for a single role, where an hourly operation runs the loop at fleet scale.
The loop also begins upstream of day one. Some accepted offers never clock in, and a rolling funnel simply absorbs the no-shows into next week's queue, even though offer-to-start leakage responds to the funnel's own behavior: its speed from application to offer, and its clarity about schedules and start dates. A program scored on ninety-day survival counts those exits, traces them to their stage, and treats the pattern as telemetry. A program scored on time-to-fill alone never records that they happened.
A stage design the funnel can run while it runs
The evidence adds up to a stage design, compact enough to state in full.
- Untimed, non-cognitive measurement goes first, to the whole queue, on any device. The equivalence evidence licenses it (Illingworth et al., 2015; King et al., 2015), and a platform built phone-first makes the first sort a measured sort, so screening time goes to candidates the measurement has already argued for.
- Timed ability runs on a device-aware policy or not at all. Where a role justifies a cognitive measure, the operational mobile deficit is a delivery fact to design around (Arthur et al., 2014): the policy announced, the moment under the candidate's control, and the screen and the clock treated as part of the test until the design removes them (Arthur, Keiser, & Doverspike, 2018).
- The program's criterion is ninety-day survival with service quality. Fill dates are an input; the survival curve, read by site and by funnel stage, is the output the program answers for, and it is the number that connects hiring to the churn loop in Figure 3.
- Applicants are customers, and the funnel is brand exposure. The applicant-reactions evidence, and what it implies for assessment design, is reviewed in the companion briefing on candidate experience during assessment; in high-volume hourly hiring it binds hardest, because the people in the queue walk back in as customers.
The standards do not scale away. The validation and documentation expectations of the profession's Principles apply at any volume (Society for Industrial and Organizational Psychology, 2018), and a rolling funnel, unlike a seasonal one, is never between cohorts, so monitoring is standing work done while the funnel runs: completion by device, drop-off by stage, score drift, survival by cohort. For a multi-site enterprise frontline program, that standing telemetry is also how headquarters sees the funnel at all. How a talent function turns this into its default machinery, owners, cadence, and documentation included, is where this series closes, in the briefing on the evidence-based hiring operating model.
High-volume hiring never stops being a queue: it re-forms every Monday, over a funnel that will never have enough readers. What a measured funnel changes is what the week leaves behind: cohorts chosen on constructs instead of timestamps, tracked to ninety days, in a loop that finally runs the other way, each week's survival data sharpening the next week's screen. The queue will be back on Monday. The floor it feeds no longer has to empty at the old rate.
Where 5Profiler stands
5Profiler's delivery is phone-first because hourly funnels are: batteries run on the device the application came from, in a sitting that fits a break between shifts, kept short by adaptive testing rather than by cutting measurement away. Untimed measures go to every applicant on any device, and timed sections wait for the controlled moments the delivery policy names, so a running clock is never applied to a sitting whose conditions nobody chose. Where a role justifies that timed section, the policy is stated before the sitting opens: which part runs against a clock, on what device, and at a moment the candidate schedules.
References
- Arthur, W., Jr., Doverspike, D., Muñoz, G. J., Taylor, J. E., & Carr, A. E. (2014). The use of mobile devices in high-stakes remotely delivered assessments and testing. International Journal of Selection and Assessment, 22(2), 113–123.
- Arthur, W., Jr., Keiser, N. L., & Doverspike, D. (2018). An information-processing-based conceptual framework of the effects of unproctored internet-based testing devices on scores on employment-related assessments and tests. Human Performance, 31(1), 1–32.
- Illingworth, A. J., Morelli, N. A., Scott, J. C., & Boyd, S. L. (2015). Internet-based, unproctored assessments on mobile and non-mobile devices: Usage, measurement equivalence, and outcomes. Journal of Business and Psychology, 30(2), 325–343.
- King, D. D., Ryan, A. M., Kantrowitz, T., Grelle, D., & Dainis, A. (2015). Mobile internet testing: An analysis of equivalence, individual differences, and reactions. International Journal of Selection and Assessment, 23(4), 382–394.
- Kruyen, P. M., Emons, W. H. M., & Sijtsma, K. (2013). On the shortcomings of shortened tests: A literature review. International Journal of Testing, 13(3), 223–248.
- Society for Industrial and Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). Industrial and Organizational Psychology, 11(Suppl. 1), 1–97.