Wyzer Logo ← Back to Blog

Who is accountable for the quality of what AI writes into your specifications?

An AI-generated requirement can be fluent, technically structured, and completely wrong, and still pass every skim-level review before sign-off. Wyzer's co-founder argues that accountability for AI output, not ownership of AI itself, is the question automotive engineering hasn't answered yet, and the vendors' own terms of service have already decided where that accountability lands.

By Peter Virk, Patrick Bartsch August 6, 2026 8 min read

An engineer working on a laptop in the open driver's door of a white car inside an EMC test chamber

Photo by ThisIsEngineering on Pexels

A requirement can be beautifully written, technically structured, and completely wrong, and nothing in how it reads will tell you that. Peter Virk, WYZER's co-founder, spent a career in safety-critical automotive engineering, as CTO of Lotus and in senior roles at JLR and Forseven, before he built a company around specification analysis. He puts the risk more bluntly than most people selling AI tools would dare: "AI can be convincingly wrong." Not slow. Not vague. Wrong in a way that survives a glance, survives a skim read, and sometimes survives a review.

Virk was one of four industry leaders interviewed for the latest instalment of Ennis & Co's "Torque and Truth" research, developed in partnership with Pybus Recruitment from conversations with fifty senior automotive leaders. The other three voices in that piece, Paddy McGillycuddy, Managing Director at JLR, Jeanette Ward, Director of HR at GardX Group, and Nigel McMinn, Managing Director at Pybus Recruitment, were talking about AI adoption broadly: marketing, recruitment, customer experience. Virk was talking about engineering content specifically, and his answer is the one that should worry anyone who treats a specification as a document rather than as an instruction that other systems and people will act on without asking questions.

The gap between assigning AI and being accountable for it

Ennis & Co's original research, a year earlier, found that AI had become every automotive leader's priority and almost nobody's strength. The follow-up conversation found the same organisations further along and no less exposed: more people using the tools, few of them seeing the deeper benefit, because using AI and being accountable for what it produces turned out to be two different disciplines.

McGillycuddy's account of JLR is instructive here, not because it is unusually cautious but because it shows what accountability looks like when it is taken seriously. AI sits on the CEO's and the main board's agenda. The chief growth officer folds it into his monthly narrative. Every commercial teammate is being skilled up, with quarterly competitions for the best uses. None of that is about capability so much as about making sure the people closest to the output are the people answerable for it. Jeanette Ward runs the same principle at GardX with a different mechanism: the good, the bad, and the ugly of the company's AI use travels up and down the organisation deliberately, so ownership is shared rather than sitting with whichever team happened to adopt the tool first.

Virk states the same idea from the engineering side, and states it as a warning rather than a policy: "You can't lead AI from a PowerPoint presentation. You have to use it." Accountability that stops at a strategy slide is just delegation without understanding.

Why AI-authored content fails plausibly, not obviously

The specific failure mode Virk described, which is not hypothetical. Engineering teams are now producing requirements, specifications and code with AI assistance, and some of what comes back has outdated, unrelated context and technologies quietly worked into the text. Nothing looks amiss. The formatting is correct. The terminology is fluent. However, it takes an experienced engineer to read it and think: hang on, that isn't right. It will hand you something beautifully written, technically structured, and entirely credible, and it will still be wrong.

This is a different failure to the ones most teams have built review processes around. A missing section is obvious. A contradictory number is findable if someone happens to compare the two clauses that contain it. A specification that reads as coherent, confident, and internally consistent, while quietly asserting something false, is the failure mode a manual review is worst at catching, because a manual review relies on something looking wrong before anyone stops to check it. It's the same gap assessors keep finding in traceability audits: a requirement can be correctly, bidirectionally traced to its parent and still be wrong, because a trace link confirms a connection exists, not that what it connects makes sense.

Virk uses a tool comparison to explain why this matters at scale, and it holds up under weight. A more powerful drill gets more good work done in a day, and it does more damage the moment someone stops paying attention to what it's cutting into. AI is the more powerful drill. Left unchecked, he says, it becomes "a very powerful way of scaling mistakes." Which leads him to a distinction he thinks most of the industry hasn't made yet: knowing who owns AI matters, but knowing who is accountable for the quality of what it produces matters more. An industry racing to generate requirements, specifications and code faster has an uncomfortable question sitting underneath the speed gain, which is whether all that speed is simply producing poor engineering content faster.

That distinction is not a semantic one, and the contracts prove it. Read the terms behind any of the major AI providers and the position is strikingly consistent: output is supplied as-is, no warranty is given as to its accuracy or its fitness for a purpose, and responsibility for what is done with that output sits with the customer. The vendors own the model. Not one of them accepts accountability for the quality of what it writes into your specification. That liability was transferred the moment the tool was adopted, whether or not anyone in the organisation read the clause that transferred it. Which makes the question unavoidable rather than optional: the accountability already sits inside your engineering organisation. The only thing still open is whether anyone there has been told it is theirs.

The judgment that isn't a bottleneck to remove

Jeanette Ward's approach to her own teams doubles as a description of what good AI use actually requires from the human side of it. Her advice has become something of a mantra: "You've got to treat AI as if it's an individual working for you. You've got to give them clear instructions and get what you need." Most failures she sees start earlier than the output, with someone who "chuck[s] the one-liner in and then expect[s] this huge answer." Her teams are taught never to take a confident answer on trust. A tool that cites employment law fluently may be citing the wrong country's law entirely unless told otherwise, and the fluency gives no signal either way.

Nigel McMinn's example sharpens the same point from a different direction. AI candidate-matching, in his experience running a recruitment business, is "spectacularly useless" for one specific reason: it cannot read what an experienced recruiter reads instantly, the pattern of too many jobs in too few years, the way the wording of a CV signals whether a conversation is worth having. He expects the technology will eventually get there. It has not yet, and until it does, that judgement isn't a gap in the tooling to be patched. It's the actual job.

Specifications have their own version of that pattern-reading. An experienced systems engineer catches a contradiction between two requirements not because a rule fired, but because they hold both clauses in mind at once and notice they cannot both be true. That instinct doesn't disappear because AI is now involved in drafting or accelerating the document. If anything it becomes the scarcest resource in the process, because the volume of content it now has to check has grown faster than the number of people who can check it.

Sign-off can't depend on who happened to read closely

This is where the accountability question stops being abstract. If the safeguard against a convincingly wrong requirement is "an experienced engineer happened to read that particular clause carefully today," that is not a review process. It is luck, applied inconsistently across a document with hundreds or thousands of requirements, most of which nobody has the time to read that carefully even once.

WYZER Detective exists for the part of this problem that doesn't scale by hiring more careful readers. It doesn't replace the engineer who knows the domain well enough to say a requirement isn't right. It gives that engineer a way to apply the same scrutiny across an entire specification instead of the fraction of it they had time to reread. On a 288-requirement AUTOSAR specification, Detective surfaced 37 duplicate candidates, 12.8% of the document, that shared no common vocabulary and would not have matched on a keyword search, alongside a full per-requirement quality breakdown. On specifications that contained genuine contradictions, it isolated the high-risk, directly-conflicting pairs, roughly 2% of the total requirement set, out from the much larger volume of low-risk overlap that looks similar at a glance but isn't the same problem. Separating those two categories reliably is exactly the problem hybrid approaches to contradiction detection are built to solve, since language similarity alone can't tell a genuine conflict from harmless overlap. The same engine has run against specifications from a few hundred requirements up to 25,000, and across requirements expressed as text, tables, and embedded diagrams, not just prose.

None of that removes the human from the loop; it changes what the human is asked to do. Instead of scanning a thousand requirements hoping the wrong one catches their eye, the engineer responsible for sign-off gets pointed at the handful that are actually in tension, with an explanation of why, and makes the call from there. We help engineers find the needles in the haystack. The accountability stays exactly where it was. What changes is whether that accountability is backed by something that scales with the document, or by whoever happened to be paying the closest attention that week.

Accountability that scales with the work

None of the four leaders in Ennis & Co's research describe AI adoption as a comfortable curve to climb at a steady pace; everyone now has an AI strategy, and the industry is dividing around who has done the less visible work underneath it. As Virk puts it, most organisations are still on the early steps of that work. Owning the technology rather than delegating it is the foundation. Knowing where it fails is what makes that ownership safe rather than cosmetic.

The part of his argument that's easiest to skip past is this: speed without a matching increase in scrutiny doesn't produce an advantage, it produces more of whatever was already wrong, faster. A specification signed off on the strength of how confident it sounds doesn't become safer because AI wrote parts of it; its errors simply travel further before anyone catches them, since the volume being produced has outpaced the humans checking it. Accountability, in the end, is not a line in someone's job title or an objective on a slide. It is the answer to a much smaller question: when this specification goes out the door, who looked closely enough at it to be the reason it's right, and did they actually have the means to look that closely at all of it, not just the parts they had time for.

About the Authors

Peter Virk

Co-founder at Wyzer — building Sherlock to find what specifications hide

30+ years in automotive technology and digital innovation, including senior roles at Jaguar Land Rover, FORSEVEN, BlackBerry QNX, and Lotus Cars.

Patrick Bartsch

Co-founder at Wyzer — turning requirements intelligence into engineering confidence

20+ years in automotive software, cloud, and AI, including roles at Volkswagen Group, Audi, Jaguar Land Rover, and AWS, with a PhD in Electronics and Computer Science.