Auditors need a controlled evidence chain, not a chat transcript
A September 17 discussion among CFOs and controllers asked what happens when external auditors inspect AI used in finance. The most useful answers were not about model internals. They were about ordinary accounting discipline: no black-box entries, formula-backed workpapers, source-to-ledger traceability, execution logs, named review, exception evidence, and a human gate before anything posts.
One controller described rejecting an AI-built payroll workpaper because it was hard-coded. The replacement needed formulas and visible work so another person could pick it up and understand it. Another practitioner summarized the durable rule: AI may process, but it cannot approve. That is a more useful design target than preserving thousands of tokens of prompt history.
Prompt and response history can help reconstruct what a tool saw and said. It rarely proves that the input population was complete, that the source was reliable, that the correct model and configuration ran, that formulas agree with the source, that exceptions were resolved, or that a qualified reviewer approved the exact artifact used in financial reporting. A screenshot is even weaker: it usually omits report parameters, hidden filters, run identity, later changes, and control totals.
The regulatory direction reinforces this control view. The FRC's generative and agentic AI guidance keeps the human auditor accountable and asks how appropriate confidence in AI outputs is obtained. PCAOB AS 1105 centers sufficient, appropriate, relevant, and reliable evidence, including the controls over electronic information produced or received by the company. FEI's 2026 AI and ICFR framework emphasizes outcome validation, human review, curated ground truth, challenger comparison, analytics, and outlier resolution.