When my newsroom first experimented with an AI recruiter scanning CVs, I was skeptical. The appeal is obvious: faster screening, apparent consistency, and the promise of removing human bias. But after months of watching resumes disappear into opaque models and seeing surprising shortlists emerge, I learned a few hard lessons. If you're considering letting an AI recruiter scan job applicants’ CVs, there are specific bias checks and contract clauses you should insist on before any code touches a real person's data.
Ask for the model’s provenance and training data summary
I always start by asking not just “what model?” but “what was it trained on?” A resume-screening model trained primarily on resumes from one industry, geography, or demographic cohort will reproduce those patterns. Ask the vendor to provide a high-level summary of the training data sources (e.g., public job boards, internal HR data, synthetic augmentation) and whether any protected characteristic proxies were used.
Concrete questions to ask:
Demand transparency about features the model uses
AI models can rely on obvious signals — job titles, keywords — but also on proxies we might not expect, like gaps in employment, university names, or even email domains. I insist on a clear list of the input features the model uses and any engineered features derived from them.
Run concrete bias checks and insist on public metrics
Metrics change the conversation. I ask vendors to provide bias audits that include disaggregated performance metrics across protected groups (where legally permissible) and proxies like gendered names, non-Anglo names, older applicants, international experience, and care gaps.
If the vendor resists, treat that as a red flag. You can also request an independent third-party audit — and require the right to publish a sanitized summary of the findings with your brand.
Insist on a human-in-the-loop policy and escalation paths
Automating all screening is a mistake I won't make again. Insist on a clear human-in-the-loop (HITL) requirement in the contract: final hiring decisions must remain with human recruiters, and there must be thresholds where human review is mandatory (e.g., when the model rejects but a hiring manager requests reconsideration).
Ask for explainability and contestability features
Applicants and recruiters should be able to understand why the model made a recommendation. I require explainability outputs that are human-readable, like top contributing features for each decision (e.g., “advanced because: 6+ years experience; rejected because: missing degree requirement”).
Contract clauses I always include
When it comes to the legal side, vague assurances aren’t enough. Here are practical contract clauses I push for, phrased as negotiable asks you can adapt.
| Clause | Purpose | Sample language ideas |
|---|---|---|
| Data provenance & documentation | Know what trained the model | Vendor shall provide a dataset provenance report and training summary within 30 days of contract start. |
| Bias audit & remedial plan | Detect and fix disparities | Vendor will conduct quarterly bias audits and supply disaggregated metrics; if disparity > X, vendor will implement remediation within Y days at no additional cost. |
| Explainability & outputs | Make decisions contestable | For each recommendation, vendor will provide top N contributing factors and a machine-readable rationale. |
| Human-in-the-loop (HITL) | Preserve human oversight | Model outputs are advisory only; final selection decisions must be performed by designated human staff. Vendor will support override logging. |
| Right to audit & independent review | Verify vendor claims | Customer may commission an independent auditor annually; vendor will cooperate and provide necessary documentation under NDA. |
| Data protection & retention | Protect applicants’ personal data | Applicant data will only be used for agreed purposes, retained for maximum Z months, and deleted upon request or contract termination. |
| Non-discrimination indemnity | Shift risk | Vendor indemnifies customer against legal claims arising from discriminatory outcomes proven to be caused by the vendor’s model. |
| Versioning & change control | Avoid silent model drift | Any model updates affecting decisions require 30 days’ notice, and re-testing on representative datasets; emergency patches must be documented. |
Practical bias tests you can run yourself
If you can’t wait for an audit, here are quick checks that reveal glaring problems. I ran these on a pilot and they exposed surprising sensitivities.
Negotiate SLAs, transparency, and termination rights
You should negotiate service-level agreements that cover accuracy thresholds and response times for appeals and incidents. I also insist on a termination-for-convenience clause tied to unacceptable disparate impact — meaning if the tool demonstrably harms hiring equity, I can pause or terminate quickly without punitive penalties.
AI can help sort floods of applicants, but without deliberate checks and enforceable contractual protections, it risks automating discrimination. Ask pointed questions, demand transparency, and put clear, testable obligations into the contract. Insist on human judgment where it matters, and build the right to audit and terminate into any agreement. Those measures turned our pilot from a risky experiment into a tool that speeds hiring while protecting fairness — and they'll give you the same kind of confidence before you let any algorithm touch candidates’ careers.