Auditable evidence of your AI's quality: the 7 questions the auditor will ask
An AI auditor doesn't ask about intentions: they ask for records. The seven questions they will ask, and the concrete artifact that answers each one.

There is a quick way to find out whether your AI’s quality is auditable: picture the meeting. An auditor —from the regulator, from an enterprise customer, from the ISO certification your company is chasing— sits down across from your team and starts asking questions. They are not interested in intentions or demos: they are interested in records. The difference between “we test a lot” and “here is the evidence” is exactly the difference between a bad afternoon and a passed audit.
This article is that meeting, simulated. The seven questions an AI auditor asks in some order, and the concrete artifact that answers each one. If you can produce all seven, you have auditable evidence; every gap is your to-do list. It is, at bottom, the practical test of Responsible AI: not what the organization declares, but what it can show.
Question 1: “What exactly did you test?”
The artifact: the versioned test plan. The list of cases —question, context, success criterion— organized by topic and by risk, with its version history. The first thing auditors look at: that the cases cover the system’s declared risks (if you mapped “the agent could make up interest rates”, the plan that tests for it must exist). A plan with no traceability back to the risk analysis answers what you tested but not why that was what needed testing.
Question 2: “Against what criterion did you decide it ‘passes’?”
The artifact: the criteria and thresholds written before the run. What every answer must satisfy, which behaviors are blocking, and what minimum score each dimension requires, dated earlier than the results. That chronological order matters: criteria defined after seeing the results are the QA version of drawing the target around the arrow, and an experienced auditor spots it by comparing dates.
Question 3: “Who or what evaluated the answers, and why should I trust that evaluator?”
The artifact: the judges’ calibration evidence. This is the question most teams cannot answer. If the answers were scored by an LLM-as-a-judge, the reasonable auditor asks the obvious thing: and who evaluated the judge? The auditable answer is the calibration record: which evaluators were used, in which version, with what calibration status, and what deviation against expert judgment. A score from an uncalibrated judge is an opinion with decimals; the same score with its calibration on record is a measurement.
Question 4: “Show me the results. All of them.”
The artifact: complete, immutable run records. Every run with its date, the version of the agent under evaluation, the cases executed, the answers obtained and the scores per dimension —including the runs that went badly. The temptation to file only the green runs is exactly what turns evidence into propaganda: the file’s credibility comes from showing March’s 71% next to June’s 94%, with the intermediate actions documented.
Question 5: “And when did you last test this?”
The artifact: the time series. AI testing done once is archaeological evidence, not quality evidence: the models change underneath you, the business context changes above you. What an auditor wants to see is the cadence —runs on every relevant change (prompt, model, policies) plus scheduled runs— and production monitoring with its alerts. The underlying question is whether the organization would know that the system had degraded; a time series with thresholds answers it.
Question 6: “What happened when something failed?”
The artifact: the incident log with its full cycle. Detection, diagnosis, fix, and —the link that sets mature teams apart— the permanent test case born out of the incident. A file with no incidents recorded does not reassure an auditor: it alarms them, because it means either the detection system isn’t working or nobody is logging. The credible file shows failures found, handled, and turned into regression.
Question 7: “Who signed off on go-live, and on what information?”
The artifact: the documented approval decision. Name, date, and the results report the decision was made on, with each dimension against its threshold at the moment of signing. This closes the loop with governance: AI risk management frameworks (NIST AI RMF, ISO/IEC 42001) and regulations with traceability and technical documentation obligations (such as the EU AI Act for high-risk systems) converge on the same idea: someone identifiable decided, on identifiable evidence.[1]
The mirror test
Read together, the seven questions have something in common: none of them can be answered retroactively. Auditable evidence is not manufactured the week before the audit: it accumulates on its own when the quality process generates it as a by-product. That is the yardstick for your current process. Does your testing produce these seven artifacts with no extra work, or does it produce results that someone would have to turn into a file by hand?
It is also ArtificialQA’s design criterion: every run is recorded with its date, its cases, its criteria, the scores per dimension and the calibration status of the evaluators that produced them; runs are compared over time, and the whole set is exportable as the file the seven questions ask for. The audit stops being a project and becomes a query.
Frequently asked questions
What does it mean for evidence to be “auditable”? That a third party can reconstruct what was tested, against what criterion, with which evaluator, when, with what result, and who decided what based on that information —from dated, versioned records, not from the team’s memory.
Who can ask me for this evidence? Regulators (depending on jurisdiction and industry), certification auditors (ISO/IEC 42001), enterprise customers in their due diligence processes, insurers assessing coverage, and your own board after an incident. The list grows every year; the file is the same for all of them.
Do LLM-as-a-judge results count as evidence? They do if the judge is defensible: recorded version, calibration documented against expert judgment, and explicit evaluation criteria. Without that, the auditor is right to treat them as automated opinion.
How long do records need to be retained? It depends on the jurisdiction and the industry; the EU AI Act’s documentation obligations for high-risk systems, for example, have their own timeframes.[1:1] Define retention with your legal team. Technically, what matters is that records are complete and immutable from day one.
What if my current evidence shows mediocre results? A file is better than no file: evidence of 78% with an improvement plan and an upward trend is defensible; the absence of measurement is not. The frameworks ask for risk management, not perfection.
This article is informational and does not constitute legal advice. Verify your specific obligations with your legal or compliance team.
Product Manager of ArtificialQA at QAlified, with 15+ years in software testing and automation. She works at the intersection of quality and AI: designing and evaluating approaches to test non-deterministic systems and ensure their behavior in production.



