Buying AI screening software still looks like buying any other recruiting tool. A shortlist of vendors, a demo each, a feature grid, a reference call, a negotiation on price per seat.
The questions that decide whether the purchase holds up are not on that grid. They surface later, in a legal review of a rejected applicant's complaint, in a finance conversation about why the invoice tripled in a hiring surge, or in the moment someone asks what the score of 72 actually meant and nobody in the room can answer.
You are not buying a screening tool. You are buying a set of claims you will have to defend.
Why does this purchase get harder to defend after you sign?
The rules moved while the market matured, and they are still moving. California's regulations on automated decision systems took effect in October 2025, requiring employers to keep four years of records including dataset descriptors, scoring outputs and audit findings. Illinois' amendment to its Human Rights Act took effect on 1 January 2026, adding a plain-language notice duty. Meanwhile the EU's high-risk obligations for employment AI were deferred to 2 December 2027, and Colorado repealed and replaced its AI Act outright.
Buying against a fixed rulebook is therefore not an option. What you can buy is a vendor who can produce evidence on request, in whatever form a regulator or a claimant's lawyer eventually asks for it.
Most buyers do not test for that. Capterra's 2025 Tech Trends survey found that 90% of regretful HR software buyers were likely to purchase based solely on vendor-provided information, and that 64% were the only person responsible for the decision.
So the checklist below is not a feature list. It is twelve questions, each of which should be answered with a document rather than a sentence.
What the tool claims to measure
Every AI screening vendor makes a predictive claim. Very few state it precisely enough to be tested, and a claim you cannot state is a claim you cannot defend.
- What outcome does the score predict, and how was that outcome measured? The answer you want names a job performance measure: supervisor ratings, sales attainment, retention at twelve months. The answer that should worry you is "recruiter agreement" or "hiring manager satisfaction with the shortlist," which measures agreement with your existing decisions rather than the quality of them.
- Show me the validation study: sample size, job families, and date. A validity coefficient established on 400 software engineers tells you very little about screening warehouse supervisors. Ask specifically whether the tool was validated on roles resembling yours, and whether that study has been repeated since the model was last retrained.
- What happens to accuracy as our data diverges from the training data? Ask about retraining cadence, who decides when it happens, and whether the model will be tuned on your historical hiring decisions. Training on your own past hires is often sold as personalization. It is also the most reliable way to automate whatever your organization was already doing wrong.
Takeaway: if the vendor cannot name the outcome variable in one sentence, stop the evaluation there.
Bias-audit documentation
An audit is only as good as its scope, and scope is exactly what summary documents leave out.
- Produce the most recent bias audit, with its date. Annual auditing has been required for automated employment decision tools in New York City since 2023, and compliance is thin. A study of 391 employers found that only 18 had posted a bias audit report, a rate of around 5%. A vendor who has one, unprompted, is already unusual.
- Is the audit disaggregated by role, or aggregated across all of them? This is the single most consequential question on the list. Stanford researchers analyzing four million applications across 156 employers found that a tool whose published audits showed no measurable bias produced substantial disparities once results were separated by position. Thirty percent of Black applicants had applied to at least one role showing adverse impact under the four-fifths rule. Aggregation hides exactly the pattern you are auditing for.
- Who conducted it, and what did they have access to? Independence matters, but access matters more. An auditor working from output data alone can tell you the impact ratios. They cannot tell you what drove them. The gap is real: New York's Comptroller reviewed the same 32 posted audits the city's enforcement agency had cleared with one issue, and identified at least 17 potential violations.
Takeaway: ask for the audit's methodology section, not its conclusion. If it aggregates across roles, treat it as unaudited.
Data, retention and what the candidate is told
This is the section that determines whether your legal team can answer a subject access request or a document demand without calling the vendor.
- What does the vendor do with our candidate data? Specifically: is it used to train models serving other customers, is it pooled into benchmark products, and does either continue after termination? Whatever the answer, it belongs in the contract rather than in the sales deck.
- What records does the tool generate, and how long does it keep them? California now expects four years of retained scoring outputs and audit findings. A vendor whose default purge runs at twelve months has quietly made you non-compliant, and the configuration to change it is often on a higher tier.
- What is the candidate told, when, and by whom? Illinois requires notice in terms a person can actually understand. Decide whether that notice comes from the vendor's automated emails or from your own templates, because if nobody decides, the answer is that it does not go out at all.
Takeaway: write your retention period into the contract as a number. Defaults change on the vendor's schedule, not yours.
Commercial terms and the exit
Finance will ask three questions about this purchase. It helps to have asked them first.
- How does pricing behave in a hiring surge? Per-seat pricing is predictable and rewards you for hiring more with the same team. Per-candidate or per-completed-interview pricing is cheap in a quiet quarter and punishing in a volume push, which is precisely when you need the tool. Capterra found cost overruns and excessive complexity tied as the leading product-related causes of buyer regret, each cited by 37%.
- What does integration actually involve? Name the ATS and the version, establish which direction data flows, and get in writing who builds it, who maintains it when the ATS updates, and what a broken sync does to a live requisition. "Integrates with your ATS" is a category, not a commitment.
- What happens on the day we leave? Ask for the export format, whether scores and rationale come with it or only candidate records, how long you have to retrieve them, and what happens to the four years of documentation you are still legally required to hold after the contract ends. Exit terms are cheap to negotiate before signature and impossible afterward.
Takeaway: price two scenarios, not one: your normal quarter and your worst surge. The gap between them is the real number.
Where to Start
Send the twelve questions to your shortlist in writing, before the demos rather than after. The demo shows you the product the vendor has built. The written answers show you the company you would be buying from.
Then read the answers for form as much as content. Note which questions came back with an attached document, which came back with a paragraph, and which came back with an offer to "cover that on a call." That distribution is your risk register, and it will predict the relationship better than any reference call.
Take the three weakest answers to legal and finance before you negotiate, not after. The point of doing this early is not to disqualify vendors. It is that everything on this list is negotiable while you still have a signature to give, and none of it is negotiable once you have given it.
FAQs
- Do we need a bias audit if we are not in New York City or Illinois? Probably, though not as a filing obligation. California's rules already direct courts and agencies to weigh the quality, scope and recency of bias testing when assessing a discrimination claim, which means the absence of an audit becomes evidence of its own. Audit because it is your defense, not because a statute names you.
- Is a vendor's own audit good enough, or do we need an independent one? A vendor audit covers the model. It cannot cover how your team configured it, which roles you applied it to, or where you set the cutoffs, and that is where impact usually appears. Treat the vendor's audit as a precondition for buying and your own as a recurring operational task. We cover the six-step process for running one separately.
- How much of this applies to a small team buying a single screening tool? All of it, and it takes less time than you would expect. Questions 1, 5, 8 and 12 carry most of the weight. A small team's real exposure is not regulatory scale, it is that nobody internally owns the tool once the person who bought it changes roles.
- What is the most common mistake in this process? Evaluating the model and ignoring the configuration. Two employers running identical software with different cutoffs, different job families and different reviewer overrides will produce different impact ratios. The tool is roughly half of what you are buying. The other half is a set of decisions your team has not made yet.
Related Articles