The first mandate for algorithmic transparency in hiring shipped with a test of itself built in: the bias audits it required had to be published, so anyone could go and count them. A year into enforcement of New York City's Local Law 144, a research team went counting. Wright and colleagues (2024) recorded the compliance of 391 employers and found 18 posted audit reports and 13 posted transparency notices. The mandate had produced, in its first measured year, a public record thin enough to read in an afternoon.
The near-empty record means less than it seems to, and that ambiguity is itself the study's sharpest finding. The law lets each employer judge whether its own tools fall within scope, so a missing posting is unreadable: it may mean no covered tool is in use, or that the employer decided its screening software sits outside the definition, or that the law is being ignored. All three produce the identical public silence. Wright and colleagues named the condition null compliance and traced it to the statute's dearth of enforcement mechanisms.
The measurement changes what transparency talk in hiring is about. Transparency stopped being a virtue vendors advertise and became a legal artifact with an outcome you can count, and the first count is sobering. Meanwhile, the explanation duties accumulating in other law (meaningful information under the GDPR, transparency and oversight duties for high-risk systems under the EU AI Act) ask for something the psychometric tradition already knows how to produce: a statement of what was measured, why it is job-relevant, what the score means, and how to contest it. That is a report, not a model dump.
Algorithmic transparency now has a denominator
Local Law 144 of 2021 attaches to the use of an automated employment decision tool an annual bias audit conducted by an independent party, a published summary of its results, and advance notice to the candidates the tool will evaluate (New York City Council, 2021). What each duty covers, and how the regime compares with its counterparts elsewhere, is the territory of this series' map of AI hiring regulations. This briefing is about what the mandate produced once it was live.
Wright, Muenster, Vecchione, and colleagues (2024) treated the mandate's first year as an empirical question. Their study, presented at the ACM Conference on Fairness, Accountability, and Transparency, distributed the fieldwork across 155 investigators, each recording whether specific employers had made public the artifacts the law requires: the audit summary and the candidate notice. Figure 1 holds the resulting counts on a single scale, and the single scale is most of the message. Beside the bar for employers checked, the bars for posted audits and posted notices are barely wide enough to print.
The count, read carefully, is not an indictment of any employer on the list, and the study names no violators; neither does this article. The point is structural. A transparency statute whose scope each regulated party interprets for itself, and whose breaches no office is resourced to find, converts a duty into an invitation. Some employers may have concluded in good faith that no tool they use meets the definition. Others may have reached the opposite conclusion and stayed silent anyway. The design guarantees that the public record cannot distinguish them, which is exactly the condition the phrase null compliance was coined to capture.
Disclosure regimes borrow their enforcement theory from other corners of regulation: put the fact where the public can see it, and readers will do what inspectors cannot. For the mechanism to run, a reader has to exist, notice an absence, and have somewhere to take the complaint. The Wright fieldwork shows what the mechanism looks like when volunteers supply it instead of the statute: a distributed roster checking employer pages one at a time, producing the first public measurement of the mandate's outcome. A transparency rule that depends on unpaid vigilance to notice its own failure has, in effect, no enforcement at all, and its published numbers will look like Figure 1's until something else changes.
What the rules actually ask employers to explain
New York's regime is unusually countable; the duties elsewhere are broader and less prescriptive about form. Under Article 22 of the GDPR (Regulation (EU) 2016/679), decisions based solely on automated processing that carry legal or similarly significant effects trigger safeguards, including the possibility of human intervention, and the regulation's information and access provisions entitle the person to meaningful information about the logic involved in such decisions. What those words license is narrower than the "right to explanation" often attributed to them in commentary, and the next section returns to the gap. The wider set of rights that surrounds assessment data (access, erasure, minimization) belongs to the companion briefing on candidate data privacy.
The EU AI Act (Regulation (EU) 2024/1689) approaches from the other direction: it classes AI systems used in employment as high-risk and attaches transparency and human-oversight obligations to that class. Set the regimes beside one another and the convergence is striking. No regime in the set demands model weights, source code, or training data. Each asks, in its own vocabulary, who was evaluated by what, whether outcomes differ by group, what the person was told beforehand, what the score meant, and who can intervene. Those are questions about measurement and its governance, and the testing profession has carried the corresponding obligations for its own reasons: the Standards oblige test users to document their instruments and to give test takers information about what scores mean and how they will be used (AERA, APA, & NCME, 2014). The explanation duties now accumulating in law are, in shape, a psychometric file.
The convergence carries a practical corollary for whoever owns compliance. Teams tend to assume that algorithmic accountability begins with access to the vendor's model, and stalls when the vendor declines. The duties as written run on artifacts an assessment program already controls: records of what each tool measures, the evidence for job-relevance, group-level outcome monitoring, the notice text, the review route. An employer can assemble every one of them without ever seeing a weight, which means vendor secrecy, whatever else it costs, is no alibi for an empty file.
An explained model can leave justification untouched
The demand for explainable AI in hiring usually arrives as an engineering requirement: open the black box, surface the features, rank their importance. Selbst and Barocas (2018), in a Fordham Law Review analysis of why explainable machines appeal to us, argued that the appeal runs together a question about mechanism with a question about justification. Inscrutability, the property that makes a model's operation hard to follow, is one obstacle; whether the basis of the decision is defensible is another; and an explanation can dissolve the first while leaving the second untouched. A model can be opened, its features listed, its logic traced, and the demonstration can still say nothing about whether any of it should bear on a hiring decision.
The distinction survives full disclosure, which is what makes it more than a complaint about deep learning. Plenty of scoring in hiring funnels involves no neural network at all; a weighted rubric or a linear screen can be read in a spreadsheet. There the inscrutability problem never arises, and the justification question stands exactly where it stood: why these variables, at these weights, for this role. Neither mystery nor its absence settles whether the basis was sound.
The hiring translation is direct. A feature-importance chart over an automated interview score explains mechanism: these inputs moved the number. It justifies nothing, because job-relevance is not a property a model can confer on its own inputs. Justification is validity evidence, the demonstrated relationship between what an instrument measures and how people perform in the role, which is the ground covered by this series' flagship review of what predicts job performance and by the measurement foundations on our science page. How far vendor claims in this market run ahead of published evidence is audited in the companion briefing on AI-scored interviews; the pattern documented there is the Selbst and Barocas warning in commercial form, explanation offered where evidence should be.
The distinction settles what a compliant explanation should contain. An employer that answers a candidate's query with a saliency map has satisfied curiosity about mechanism and said nothing about soundness. An employer that answers with the constructs measured, the evidence connecting them to the role, and the meaning of the score has addressed the question the law is reaching for, whether or not the underlying system is complex. Explainable AI in hiring, pursued seriously, turns out to be mostly documentation the psychometric tradition already prescribes.
In the candidate's hands, information cuts both ways
It would be convenient if more disclosure were simply better for candidates, and the experimental evidence declines to cooperate. Langer, König, and Fitili (2018) provided applicants with procedural information about a technology-mediated selection process and found the provision double-edged: information raised how open and transparent the procedure appeared while it could simultaneously worsen other reactions to the process. Telling people more about an automated procedure changes what they attend to, and some of what it makes salient they do not welcome. The authors' conclusion was about design: what matters is not whether information is given but which information, framed how. That argues for craft in the text itself: treat the candidate-facing disclosure as a designed artifact, drafted, tried on people who resemble applicants, and revised when a disclosure produces alarm in place of understanding. Organizations pilot assessment content as a matter of course and release the words around it untested, though the words are the part every candidate is guaranteed to read.
The framing principle comes from the candidate-reactions research this series reviews in its briefing on candidate experience: applicants read every disclosure through job-relatedness, and information connecting the procedure to the job lands differently from information that merely describes the technology. On that logic, and on the testing Standards' account of what test takers are owed, a usable candidate explanation answers the questions a candidate actually holds: what was measured, why it maps to this job, what the score means, what happens next, and how to ask for review. The last answer borrows its shape from due process, a stated route to a human decision-maker, of the kind the companion briefing on integrity flags and false positives builds for its own high-stakes case.
Transparency in hiring now runs in both directions, and the disclosure conversation should assume it. The candidates reading these notices increasingly bring their own AI to the process; what that does to each assessment format, and which policies survive contact with it, is the subject of the companion briefing on candidates using AI. The notice layer described below is the natural place for an employer to state its policy on that use, which makes the notice a two-way contract rather than a warning label.
One architecture can answer every jurisdiction
An organization facing this accumulation of duties can respond statute by statute, or it can notice that the duties stack. Figure 2 arranges the response as layers: a governance audit at the top, an employer notice in the middle, a candidate-level explanation at the base, each answering a different legal duty and producing a different artifact for a different audience. The work is in building each layer for the audience it serves, and then letting one set of artifacts do duty across jurisdictions, because the statutes differ far more than the evidence they ask for.
The audit layer is New York's shape taken seriously: independent, periodic, published, whether or not a statute compels it for a particular tool in a particular city. What the Wright count exposed was not a flaw in auditing but the fate of unenforced publication duties; an organization that commissions the audit for its own governance rather than for a regulator has removed the incentive problem, because the client for the work is itself. The candidate explanation at the base is the layer the other two cannot substitute for. An audit tells a regulator the tool was examined; a notice tells a population what will happen; only the explanation tells one person what her score is grounded in, and it is the artifact the testing Standards have in effect specified all along.
The notice layer in between is thinner but load-bearing in its own way, and it can promise only what the systems beneath it deliver: which tools run, what data they take in, at which stages, and under what protections the data is held, operational facts of the kind set out on our security page. Writing the notice from those facts, and keeping it current as tools change, beats writing it aspirationally and discovering the gap when a candidate or an auditor tests it.
Each layer also needs an owner with a calendar. The audit layer sits naturally with whoever governs assessment; the notice sits with the team that changes the funnel, because a notice goes stale the day a new tool is piloted; the candidate explanation sits with whoever selects instruments, since it can only be written from an instrument's documentation. Algorithmic transparency survives reorganizations exactly as well as those assignments do, and an architecture with named owners has a property the New York record conspicuously lacked: someone whose job it is to notice the gap.
A mandate nobody checks is a request
The Wright result generalizes beyond New York, and organizations should absorb it as a lesson about their own internal rules as much as about statutes. Algorithmic transparency, on this evidence, behaves like any other output of governance: it appears where someone owns it and checks it, and nowhere else. A mandate whose scope is self-judged and whose observance nobody measures produces null compliance wherever it lives, and a corporate policy that assessment tools should be documented, owned by nobody and reviewed never, fails in the identical way: silence that cannot be told from compliance. The fix in both cases is machinery, a named owner, a cadence, a check that runs on schedule. An external mandate at least gets counted eventually, as New York's was; an internal one can stay unmeasured for years. The Wright design is cheap to point at oneself: sample the funnel on a schedule, record which tools carry current evidence and current notices, and circulate the result where leadership will read it.
For employers, the practical stance is to want the artifacts for reasons no statute supplies. Treat the audit layer as evidence hygiene and it yields the documentation that a discrimination claim, a procurement review, or a skeptical candidate will eventually request anyway; the logic of running assessment on standing evidence, collected as a program instead of assembled in the week a request lands, is developed in this series' briefing on standards-based assessment programs. The test of the posture is simple. An organization that would keep auditing if the mandate were repealed has solved its own enforcement problem. One that would stop has learned what its compliance was made of.
Building AI hiring transparency around validity evidence
The implications for an organization running or buying algorithmic assessment sort into a sequence. First, inventory. Most funnels contain more scoring than their owners reckon with: résumé parsers, knockout questions, ranked search, scored assessments, scored interviews. Each regime defines its covered category differently, and the definitions are collected in the regulatory-map briefing; the requirement is simply to know, before anyone asks, which tools in the funnel would meet any regime's automated-decision definition and what evidence exists for each.
Second, build the layers once and reuse them. An audit file assembled to the strictest standard an organization faces will satisfy the rest; a notice written for genuine clarity discharges duties that ask for less; a candidate explanation that names constructs, evidence, and meaning answers the GDPR's information provisions, the AI Act's transparency expectations, and the testing Standards at once. Jurisdictional difference is real, but it lives mostly in thresholds and filing details, and hardly at all in the underlying artifacts.
Third, put validity evidence at the center of every artifact, because it is the one ingredient the transparency conversation cannot synthesize. A bias audit of an instrument that measures nothing job-relevant is a fairness certificate for noise; an explanation of such an instrument is a well-lit tour of an empty building. The order of operations matters. Validity evidence comes first, then the transparency layers that carry it outward, and the sequence is also the economical one, since documentation of sound measurement mostly exists by the time anyone regulates it.
Fourth, pre-write the candidate explanation for every instrument before it goes live. Figure 3 assembles the anatomy of that artifact, the answers this article has already inventoried, arranged as a candidate would meet them. The middle element should state the score's uncertainty as well as its level, the case for stating bands is made in this series' briefing on measurement error. An explanation drafted once per instrument deploys to every candidate without per-decision authoring, and its existence changes the character of what arrives afterward: a candidate who received a real explanation asks follow-up questions, while one who received boilerplate files complaints.
AI hiring transparency will accumulate more duties, and jurisdictions will go on differing over thresholds, definitions, and filing formats. The constant underneath is older than any of the statutes: an explanation is only as strong as the measurement it explains. Organizations that run validated instruments get their transparency artifacts at the margin, as descriptions of work already done. Organizations that bolt explanation onto unvalidated scoring will find that the exercise exposes exactly what they might have preferred it conceal. An explained score that measures nothing explains nothing.
Where 5Profiler stands
A 5Profiler report is written as a candidate-level explanation from the first line: it names the constructs an instrument measured, states why each maps to the role under review (role-referenced scoring), and shows what every score means, so the artifact a hiring team files is one a candidate, an auditor, or a regulator could read. You can inspect the shape directly in a sample report. Assessment built this way stacks into the three layers this article describes without extra authoring: the report answers the candidate, its documentation answers the audit, and the notice can promise only what the platform actually shows.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
- European Parliament, & Council of the European Union. (2016). Regulation (EU) 2016/679 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data (General Data Protection Regulation). Official Journal of the European Union, L 119.
- European Parliament, & Council of the European Union. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act).
- Langer, M., König, C. J., & Fitili, A. (2018). Information as a double-edged sword: The role of computer experience and information on applicant reactions towards novel technologies for personnel selection. Computers in Human Behavior, 81, 19–30.
- New York City Council. (2021). Local Law 144 of 2021, N.Y.C. Administrative Code § 20-870 et seq.
- Selbst, A. D., & Barocas, S. (2018). The intuitive appeal of explainable machines. Fordham Law Review, 87(3), 1085–1139.
- Wright, L., Muenster, R. M., Vecchione, B., Qu, T., Cai, P., Metcalf, J., & Matias, J. N. (2024). Null compliance: NYC Local Law 144 and the challenges of algorithm accountability. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24).