No pull request gets merged on my teams without a review. No release ships without tests. Nobody on the team argues with this. We learned the hard way what happens when code goes to production unchecked.
So here's a question I keep asking engineering leaders. When did you last review the software deciding who gets to join your team?
Most of them go quiet. The applicant tracking system came from a vendor. HR bought it. It ranks, filters, and scores candidates before a human ever reads a name. And nobody on the engineering side has ever looked inside it.
I hired 56 engineers over four years in one role. I read a lot of resumes. I also know how many good people I would have missed if a filter had quietly dropped them before they reached my inbox. I'd never have known. Neither would they.

We Already Know What Untested Hiring AI Does
This isn't a thought experiment. Amazon ran the experiment for us years ago.
According to Reuters' 2018 report, Amazon built a tool to score job candidates from one to five stars, "much like shoppers rate products on Amazon." The team trained it on resumes submitted to the company over a 10-year period. Most of those resumes came from men.
The model learned its lesson well. It penalized resumes containing the word "women's," as in "women's chess club captain." It downgraded graduates of two all-women's colleges. Amazon disbanded the team.
Here's the part engineers should sit with. Nobody wrote a line of code saying "reject women." The bias came from the training data. The model did exactly what it was built to do... copy past decisions. If your past decisions were skewed, your model is skewed. And it runs at a scale no biased human recruiter ever managed.
Amazon had world-class machine learning engineers and still built something broken. What do you think is running inside the cheapest screening add-on your HR platform upsold you last year?
The Vendor Is Not a Shield Anymore
For years the comfortable assumption went like this. The vendor built it. The vendor carries the risk. We only clicked "enable."
A federal court in California put a dent in this idea. In Mobley v. Workday, a job applicant alleged Workday's AI screening tools discriminated against candidates over 40. In spring 2025, the court preliminarily certified a nationwide collective action covering "all individuals aged 40 and over" who applied through Workday's platform and were denied employment recommendations.
The court's reasoning matters more than the headline. According to the Davis Wright Tremaine summary, vendors of AI tools face possible direct liability under federal anti-discrimination law "if their tools function as gatekeepers in hiring decisions." The claim is disparate impact. Nobody has to prove anyone intended to discriminate. The outcome is enough.
I'm not sure about this part: I haven't confirmed where the Workday case stands today, so treat the 2025 certification ruling as the last point I verified, not the final word.
If the vendor is a gatekeeper, so is the employer who switched it on. Both names end up on the paperwork.

The Laws Exist. The Enforcement Doesn't. Yet.
New York City passed the first serious attempt at this. Local Law 144 took effect on July 5, 2023. Employers using automated employment decision tools for NYC jobs need an annual independent bias audit, and they must publish the results.
Sounds strong. Then the New York State Comptroller audited the city's enforcement. According to Warden AI's coverage of the December 2025 audit, the city's consumer protection department surveyed 32 employers and found one instance of non-compliance. The Comptroller's auditors looked at the same set and flagged at least 17 potential violations.
One versus seventeen. Same companies.
Europe is slower still. The EU AI Act classes AI used in recruitment and employment decisions as high-risk. Those obligations were due in August 2026. The Digital Omnibus, published in the Official Journal on July 24, 2026, pushed standalone high-risk systems to December 2, 2027.
I've watched plenty of leaders read news like this as a reprieve. Weak enforcement. Delayed deadlines. Relax.
I read it the other way. The rules are written. The lawsuits are already moving. The only thing missing is the regulator who shows up. When the regulator arrives, "we never looked" won't sound like a defense. It will sound like an admission.
Treat Hiring AI Like Production Code
Engineering teams already own the discipline this problem needs. We test. We monitor. We roll back. We ask what the system does with bad input. Somehow none of it gets applied to the one system deciding who sits next to you on Monday.
Here's what I'd do if I walked into your company tomorrow.

Map every automated decision
Write down each place software touches a candidate. Resume parsing. Knockout questions. Ranking. Video interview scoring. Chatbot screening. Most teams find more than they expected. You own every one of them, whether you configured it or not.
Ask the vendor for the audit
Ask for their most recent independent bias audit. Ask what data trained the model. Ask how they measure impact by age, sex, and race. If they wave you off with "proprietary," you have your answer. You wouldn't accept a closed-source payment library with no test results. Don't accept it here.
Run your own impact numbers
Pull a year of applicant data. Compare pass-through rates at each stage across groups. You don't need a data science team for a first pass. A spreadsheet and an honest afternoon will show you where the funnel narrows for one group and not another.
Keep a human on every rejection path
A tool which ranks candidates for a human to review is one thing. A tool which silently rejects people nobody ever sees is another. If a candidate gets cut, a person should be able to see why and overrule it.
Put it on the incident list
If a model starts rejecting a group at a higher rate, treat it as an outage. It has an owner, a severity, and a fix date. Not a line in next year's HR review.
Your Team Is Your Output
We spend enormous effort making sure our systems don't hurt users. Candidates are users too. They're the people who might have become your best engineer, if the filter had let them through.
My research on bad bosses found 99.5% of people have worked for at least one. I spend a lot of my time on what happens after someone joins a team. Hiring is where the story starts. Every filter you run decides who gets a chance to lead, to mentor, to become the manager people remember for the right reasons.
When a person makes a bad hiring call, someone notices eventually. A colleague pushes back. A candidate complains. When software makes the same call, it repeats it thousands of times without a sound. Nobody pushes back on a filter they never see.
So pick one step in your hiring pipeline this week. Find out what the software does there, who tested it, and what the numbers show. If nobody has an answer, you've found your most important untested system.
Would you ship code like this?