Dive Brief:
- Nearly half of almost 1,500 AI code engineers, developers and leaders have shipped AI-written code that later failed in production, according to a survey published Wednesday by SmartBear, a software quality and testing provider. Among the cohort experiencing AI code failure, 69% said they still have confidence that AI-written code behaves as intended.
- SmartBear said this data mismatch is indicative of “blind trust” in AI code that escalates as it moves up the corporate ladder. Nearly three-quarters of leaders have complete confidence or a lot of confidence that AI code works as intended, compared to 52% of practitioners.
- “There’s a disconnect between using and relying on AI and its overall lack of quality or the prevalence of bugs,” Fitz Nowlan, VP of AI and architecture at SmartBear, told Channel Dive. “There's also a disconnect here between the leadership at organizations saying to drop AI in for ROI gains and to solve problems, and the practitioners who maybe initially were just as excited about it, but now are realizing the total cost.”
Dive Insight:
In practice, an AI coding failure can look like a broken button shipped onto a company website, or a failed data save while working on a user interface, according to Nowlan. The sheer volume of AI-shipped code makes oversights difficult to track.
Almost half of respondents reported not being able to explain how AI contributed to a bug or incident. Once a team had already shipped a failure, that number climbed to 73%.
“Initially, I think the thought was, ‘AI is as good as I am when I’m having my best day,’ or ‘AI is more or less perfect or infallible,’” Nowlan said. “And then kind of the realization hits that, ‘OK, AI makes mistakes just like we do.’”
Companies are increasingly using AI to write code. About two-thirds of respondents said AI writes or accelerates 41% or more of their code, up from 43% of respondents in SmartBear’s January survey. Humans simply can’t keep up.
“You can’t have a machine-powered process producing code, and then only humans doing the review and only humans doing the application testing,” Nowlan said.
Companies are already seeing the effects of AI code outpacing human review. Among the respondents, almost half were very or extremely concerned that their application quality is suffering due to AI coding.
More than half of respondents experienced application quality issues in the past year because their testing couldn’t keep up with development, resulting in revenue loss, outages and negative customer experiences.
Nowlan emphasized the importance of having an AI-powered tool on the code review and testing side, though AI code testing tools are also not impervious to making mistakes.
“We're at the phase where we know AI is very powerful,” he said. “There’s no denying that, but you can't just have blind trust in it. You can't put on your sleeping mask and then go to sleep and let AI do everything. You have to have humans in the loop.”