I spend my days building document AI for regulated environments. That work has one defining characteristic: accuracy is not a feeling. It is a measured property of the system, defined at the field level, evaluated against ground truth, and gated before anything moves to production. If a model cannot demonstrate, on a held-out dataset, that it produces correct output at a defined threshold, it does not ship. The metric exists. The reference exists. The decision is documented.
The important part is not the specific metric. It is that the system has a defined notion of correctness, a measurable reference, and a decision threshold tied to both.
This is a necessity in regulated work. When the output of a system feeds a critical process, the question "how do you know it works" cannot be answered with a vendor demo or an internal hunch. It has to be answered with numbers that hold up to external scrutiny.
Because I also have a background in HR technology, I find myself looking at a domain I know well from a different angle. Organizations are increasingly using AI inside core HR operations: payroll anomaly detection, compensation benchmarking, and leave management. Recruiting raises its own questions, including the complications around AI-generated applications, so I am setting it aside here. That makes me wonder how organizations evaluate these tools with the same rigor I have to apply to mine.
Not all HR systems are the same. A screening tool, a performance model, and a retention score all require different definitions of correctness. But in all cases, that definition has to exist.
What does "95%" actually mean?
Tools are often procured based on high-level vendor claims, frequently something in the neighborhood of "95% accuracy." From an audit perspective, that number on its own carries very little information. A few questions worth asking:
- Is that 95% measured on the end result, or per single step? If a three-step pipeline operates at 95% accuracy per stage, the compound effect means the final output is meaningfully less reliable.
- How is "correctness" actually defined? In document AI, a field is either correct or it is not. In HR, who defines the ground truth for a "reliable performance signal"?
- Where do the 5% errors fall? Are they randomly distributed, or concentrated in specific employee categories? Aggregate accuracy can hide systematic bias.
Loud failures and silent failures
Tools built on language models are probabilistic by design. The same input does not always produce the same output, and errors are spread across the input space rather than concentrated in reproducible bugs. This is not a flaw to be patched out. It is a structural property of how these systems work, and it is the reason aggregate accuracy can be deeply misleading.
A loud failure is obviously nonsense, like a free-text response where a single numerical value was expected, or a category that does not exist in the system. These are manageable because they are visible.
A silent failure looks correct. The performance tool produces consistent scores that systematically under-rate one group by a small margin. The compensation benchmark drifts because it learned from stale market data. These survive in production for years because they do not trigger alerts.
A system can also be directionally right but numerically wrong. A risk score that says "30%" but corresponds to something very different in reality, will quietly distort decisions at scale.
This is why aggregate accuracy is rarely enough. A system with mostly silent failures is a compliance time bomb.
Trust alone is rarely enough
My suspicion is that in many organizations, the answer to "where do your metrics come from?" is "we trust the vendor's documentation." HR is not unusual in this respect; the same gap exists across many AI applications. But in any other regulated workflow, it would be flagged as a missing control. The standard expectation is independent verification: your own ground truth and your own methodology.
Even if a system performs well during procurement, the more difficult question is what happens over time. Data changes, organizations change, and models are updated. Without ongoing validation, it becomes difficult to know whether the system still performs as expected.
And even if the model performs well, there is a second layer of evaluation that is harder to see: how the system is actually used. Are humans overriding it? Ignoring it? Or over-trusting it? Those behaviors often matter as much as the model itself.
An invitation, not a verdict
I am asking this question genuinely. I would like to learn how organizations are approaching this seriously:
- What does your internal validation actually look like in practice?
- How do you distinguish loud failures from silent ones?
- How do you verify a vendor's claims against your own real-world data?
If you are working on this, I would genuinely like to hear how you approach it. The answers are useful well beyond HR.
05.05.2026
