This Take 5 report from HFS Research, in partnership with QualiZeal, is for quality engineering leaders, CIOs and CTOs, and risk and compliance leaders evaluating how to test and govern AI systems running in production.
AI has moved into production faster than enterprises can ensure it works safely, fairly, or accountably.
Enterprise AI adoption is outpacing the infrastructure designed to govern it. AI applications are in production, agentic systems are making decisions autonomously, and the compliance clock is running. However, the testing and governance frameworks needed to validate these systems still rely on quality frameworks built for deterministic software to evaluate probabilistic AI. In the last two years, the quality engineering (QE) function has absorbed new scope, including AI model validation, data quality, security, and regulatory compliance, without the tools, methods, or budget to do it properly.
HFS Research, in partnership with QualiZeal, surveyed 101 enterprise quality and technology leaders in the US to examine whether organizations are properly equipped to validate AI systems in production and what a genuinely different approach to QE for AI looks like.
AI is already in production, but the controls for it are still in pilots
Fifty-seven percent of enterprises run generative AI in some form of production. Thirty-six percent have agentic systems live. However, formal AI validation frameworks are still in pilots.
Enterprises are running AI without a person or function that actually holds the line
No one owns AI quality. Thirty-seven percent of enterprises run AI in production with no named function accountable for quality. Twenty-nine percent have distributed ownership across teams. Eight percent admit no one is accountable.
Most see AI testing as different, but almost no one has built for it
Only 15% have built testing approaches for AI. Fifty-two percent are testing probabilistic AI systems with frameworks designed for deterministic software.
Enterprises can prove their AI is accurate, but not that it is trustworthy
Eighty-eight percent measure model accuracy. Yet only 47% measure bias and fairness, 37% measure explainability, and just 35% can provide regulatory compliance evidence.
QE gets the mandate to own everything that makes AI trustworthy, but not the budget or talent to actually deliver
The QE function surveyed has taken on new scope. However, only 24% received proportional resources.
The Bottom Line: AI that cannot be audited, explained, or proven fair is an open liability, and it widens with every sprint. Organizations that cannot control how their AI behaves in production are not managing risk. They are deferring the reckoning, and they can’t do it indefinitely.





The failures are already on the record. In 2026, OpenAI’s AI models broke out of a test environment and autonomously hacked into Hugging Face. In 2026, a Cursor agent deleted a company’s entire production database in nine seconds and then fabricated a justification when asked why. A Replit agent did the same after the owner declared a code freeze and told it in writing not to touch production.
Each failure shares the same structure: an AI system operating autonomously, in contexts it was never validated for, with no governance infrastructure to detect it and no audit trail to explain it afterward. That is precisely what this research documents at scale and across 101 enterprises today.
QE for AI must be treated differently than QE for legacy systems. The organizations that ignore this will spend the next 18 months explaining incidents. The ones that build for it will be explaining how they avoided them.
Appoint a named owner for AI quality
Distributed accountability is no accountability. Designate one function: QE, risk, or a dedicated AI governance team that holds the line.
Replace accuracy with a trust scorecard
Build a measurement framework that includes safety, bias and fairness, explainability, and regulatory compliance evidence. Accuracy is table stakes, not assurance.
Fund the QE mandate
If QE scope now covers AI validation, budgets must keep pace. If they don’t, it will create a structural liability. Close the gap or outsource it.
Register now for immediate access of HFS' research, data and forward looking trends.
Get StartedIf you don't have an account, Register here |
With the exception of our Horizons reports, most of our research is available for free on our website. Sign up for a free account and start realizing the power of insights now.
Our premium subscription gives enterprise clients access to our complete library of proprietary research, direct access to our industry analysts, and other benefits.
Contact us at [email protected] for more information on premium access.