Back to Blog
Scheduled — appears October 1, 2026 at 1:00 PM UTC

AI-Assisted Audit Sampling: What Passes Peer Review, What Doesn't

AI in audit sampling is being actively adopted; the peer-review implications are being worked out. The Big 4 pattern (Helix, Clara, GL.ai, Argus) — AI does the mechanical work, the practitioner does the judgment. The specific AI capabilities that pass peer review vs the ones that fail, and the workpaper documentation requirements when AI is in the workflow.

Quick Answer

AI in audit sampling is being actively adopted; the peer-review implications are being worked out. The Big 4 pattern (Helix, Clara, GL.ai, Argus) — AI does the mechanical work, the practitioner does the judgment. The specific AI capabilities that pass peer review vs the ones that fail, and the workpaper documentation requirements when AI is in the workflow.

AI-assisted audit sampling is being actively adopted across the cyber-attest market in 2026. Big 4 pilots have matured into production tools (EY Helix, KPMG Clara, PwC GL.ai, Deloitte Argus). Mid-market platforms (Fieldguide, MindBridge for financial audit, DataSnipper for evidence extraction) offer AI-native workflows. And every audit-firm-side platform's marketing page name-checks AI somewhere in the fold. The AICPA has not (as of 2026) published guidance specifically directed at AI in sampling for attestation engagements, and peer reviewers are working out — engagement by engagement — what the professional-standards implications are.

This is the practitioner's view of where AI-assisted sampling fits, where it clearly passes peer review, where it clearly fails, and how to document the auditor's judgment layer when AI is in the workflow. Not the vendor's marketing view. Not a speculative view of what AI will eventually do. The 2026 reality, from the practitioner seat.

AT-C 205
the attestation standard governing SOC 2 examinations — requires the practitioner to obtain sufficient appropriate evidence to support the conclusion, with sampling as one accepted evidence-gathering procedure. The standard is technology-neutral; the requirement is not.
4 tools
AI-in-audit at Big 4 scale — EY Helix, KPMG Clara, PwC GL.ai, Deloitte Argus. Each keeps the licensed practitioner squarely in the judgment layer while automating the mechanical operations (extraction, categorization, triage). AICPA 2024 peer-review guidance on SOC 2 sampling adequacy applies directly to any mid-market adoption.
30-50%
efficiency gain from bounded AI adoption (extraction + classification + triage), holding the practitioner's judgment layer intact. The gain is real; the peer-review safety is real; deviation from the boundary is where the risk lives.

What sampling means in a cyber-attest examination

Sampling in a SOC 2 II examination (or an ISO 27001 audit, or a PCI ROC) is the auditor's mechanism for evaluating operating effectiveness across a population without testing every item in the population. The AICPA's sampling guidance (AU-C 530 for financial audits, and by extension AT-C 205 for attestation engagements) sets out a small number of decision points the practitioner must make on every sample:

  • Define the population.: Every item subject to the control activity during the covered period. For a privileged access review, every privileged account subject to the review. For a change management control, every change ticket subject to the review procedure. Population definition is judgment — some items may be excluded (e.g., service accounts with a separate control regime) with documented rationale.
  • Define the sampling unit.: The specific item that will be tested. Usually one-to-one with the population (each privileged account is a sampling unit; each change ticket is a sampling unit), but sometimes aggregated (a batch of related changes tested as one unit).
  • Determine the sample size.: Based on the population size, the expected exception rate, the tolerable exception rate, and the required assurance level. AICPA guidance provides tables (e.g., for a population of 40+ with a low expected exception rate and reasonable assurance, a sample size of 25-30 is typical). The specific number is judgment supported by the reference table.
  • Select the sample.: Random, systematic, or stratified — the method must give every sampling unit in the population a defined chance of selection appropriate to the objective. Judgmental selection is permitted for specific purposes but is not a substitute for random selection when representative testing is the objective.
  • Perform the test procedure on each selected item.: The substantive test — was the control activity performed on this item, and did it operate as designed? Documented per item, with the source-of-truth evidence attached.
  • Evaluate the results.: Exception count against expected rate. Extrapolate the exception rate to the population. Judgment on whether the extrapolated rate exceeds the tolerable rate — and if so, whether the control has failed or additional testing is warranted.

What AI-assisted sampling actually can do — and can't

The category "AI in sampling" collapses several distinct capabilities that have different peer-review postures. Being specific about which capability is being used is the first step toward a defensible workpaper.

AI capability
What it does
Peer review posture
Evidence extraction
Extracts structured data from unstructured documents (e.g., pulls sign-off dates and reviewer names from access review PDFs). Output is data, not a conclusion.
Strong. Extraction is a mechanical operation; the practitioner still makes the judgment on the extracted data. Analogous to using OCR on a scanned document.
Population classification
Categorizes population items by attributes relevant to the sampling design (e.g., stratifies privileged accounts into service, human, and shared categories). Output is a labeling that supports subsequent judgment.
Strong when the classification is deterministic (rule-based) or the practitioner reviews and validates the categorization. Weaker when the practitioner uses the labels without validating the classification methodology.
Exception identification
Flags items likely to be exceptions (e.g., access reviews with missing sign-offs, changes with missing approvals). Output is a flag, not a conclusion.
Strong when the flag is treated as a triage aid — practitioner reviews every flagged item and makes the exception determination. Weak when practitioner treats the flag itself as the exception determination without independent verification.
Sample selection
Selects the specific items to test from the population. Can be random, stratified, or judgmentally-informed by AI-identified risk factors.
Strong when the selection methodology is deterministic and documented (random with disclosed seed; stratified with disclosed strata). Weakest area for AI — AI-selected samples that are neither purely random nor practitioner-judgmental raise the question of what basis the selection actually used.
Testing conclusion
AI evaluates the tested items and produces a conclusion on operating effectiveness.
Peer review risk area. The professional standards require the practitioner to draw the conclusion; delegating the conclusion to an AI model — even a well-tested one — is not consistent with the standard. Big 4 pattern deliberately keeps the practitioner in the conclusion loop.
Full-population testing (financial audit primarily; emerging in attestation)
Tests every item in the population using AI-driven analysis (MindBridge's model for GL testing). Not sampling — an alternative to sampling.
Novel; AICPA guidance is evolving. Where it works, it reduces or eliminates sample-size judgment. Where it doesn't work (attestation contexts with heterogeneous evidence types), it can create false-confidence gaps in coverage.
The peer-review-safe pattern

The pattern that most defensibly incorporates AI into sampling: use AI for the mechanical operations (extraction, classification, exception flagging), and keep the practitioner squarely in the judgment operations (sample size determination, sample selection methodology, testing conclusion, exception disposition). Document what AI did and what the practitioner did, separately. This mirrors the Big 4 pattern with EY Helix, KPMG Clara, PwC GL.ai, and Deloitte Argus — AI as an accelerator of professional judgment, not a substitute for it. Peer reviewers who see this pattern typically pass it without a finding; peer reviewers who see AI-derived conclusions with no visible practitioner judgment layer typically write one.

Where AI-assisted sampling passes peer review

  • Extraction of structured data from client-provided documents.: Access review sign-off dates from PDFs. Ticket dispositions from screenshot exports. Vendor risk ratings from scanned questionnaires. The AI is doing OCR-plus-classification; the practitioner is using the extracted data to make audit judgments. Documentation requirement: the extraction method (which model, what prompt), a sample of items verified manually against the source, and the exception count on the verification sample.
  • Population stratification for sampling design.: The AI groups the population by relevant attributes so the practitioner can decide whether to stratify the sample. Documentation requirement: the stratification methodology, the resulting strata, and the practitioner's evaluation of whether the stratification is appropriate for the sampling objective.
  • Exception triage.: The AI reviews the tested items and flags likely exceptions for practitioner review. Practitioner reviews every flagged item and makes the exception determination independently. Documentation requirement: the flagging methodology, the count of flagged items, the practitioner's independent evaluation of each flagged item, and the ultimate exception disposition.
  • Analytical procedures on tested populations.: The AI computes analytical statistics (exception rates by stratum, trend analysis against prior periods, outlier detection). Practitioner uses the analytics as input to sampling design and to overall risk assessment. Documentation requirement: the analytical procedures performed, the results, and how the results informed sampling and risk decisions.

Where AI-assisted sampling fails peer review

  • AI-derived sample selection without documented methodology.: The AI "selected" the sample items via some risk-informed process that the practitioner cannot describe in reproducible terms. Peer reviewer asks: what selection methodology was applied? If the answer is "the AI selected the highest-risk items," the follow-up is: what is the risk metric, how was it calibrated, and can you show the methodology produces a defensible sample? Without documentation, the sample is not defensible.
  • AI-generated exception dispositions.: The AI reviewed the tested items and produced a disposition (exception vs no-exception) that the practitioner accepted without independent verification. Peer reviewer asks: how did the practitioner satisfy themselves that the AI's disposition was correct? If the answer is "we trust the model," the finding writes itself.
  • AI-generated conclusions on control operating effectiveness.: The AI evaluated the sample results and produced the overall control conclusion. Practitioner signed the workpaper without adding a substantive judgment layer. Peer reviewer asks: what is the practitioner's substantive judgment on this control? The professional standards do not permit the practitioner to delegate the conclusion.
  • AI hallucination in workpaper documentation.: The AI generated workpaper narrative that references facts not supported by the underlying evidence. Practitioner signed the workpaper without verifying the narrative. This is a hard failure — the workpaper contains unsupported assertions, which is a direct violation of professional standards regardless of who or what authored the assertion.
  • Missing documentation of the AI's role.: The workpaper does not disclose that AI was used in the procedure. Peer reviewer asks: what testing procedures were performed, and by whom? If the practitioner cannot describe the AI's role in reproducible terms, the peer review finding cites both the missing documentation and the professional-standards concern about undisclosed reliance.

The critical documentation requirements when AI is involved

The gap between AI-assisted sampling that passes peer review and AI-assisted sampling that fails is almost always the workpaper documentation. The AICPA quality-control standards require the workpaper to enable an experienced practitioner not previously involved with the engagement to understand what was done and the significant conclusions reached. AI in the workflow does not change the requirement; it raises the bar on what the workpaper must say.

The AI-workpaper documentation checklist

For any procedure where AI was used, the workpaper should document:

1. The specific AI capability used. Not "we used AI." Name the tool (Fieldguide, DataSnipper, in-house model), the specific capability invoked (extraction, classification, triage), and the input/output of the capability.

2. The practitioner's independent verification. For every AI output the practitioner relied on, a documented verification step — either verifying every item, or verifying a defensible sample of items and documenting the exception rate on the verification sample.

3. The practitioner's judgment layer. Where the practitioner made a professional judgment (sample size, sample selection methodology, exception disposition, testing conclusion), the judgment and its supporting rationale — separate from any AI output.

4. The methodology reproducibility. Enough detail that another practitioner could re-run the same procedure and arrive at comparable results — model version, prompt or configuration, deterministic seeds where applicable.

5. The disclosure of AI in report communications (as required). The SOC 2 II report's description of testing procedures typically does not need to enumerate every tool used, but where AI plays a material role — full-population testing, AI-driven risk stratification — disclosure to the report user is increasingly expected.

The Big 4 pattern and what it tells the mid-market practice

EY Helix, KPMG Clara, PwC GL.ai, and Deloitte Argus are the reference implementations of AI in audit at scale. Each is a Big-4-scale investment (hundreds of millions of dollars over years) that mid-market practices cannot replicate. But the pattern they set — the specific way the AI is bounded and the practitioner's role is preserved — is directly instructive.

What the Big 4 tools do

Extract structured data from unstructured evidence at scale. Perform analytical procedures across full populations for financial-audit contexts (GL testing, journal entry analysis). Categorize and stratify populations for sampling design. Flag items for practitioner review. Compute trend analysis across prior periods. Generate draft workpaper narratives from testing results (with practitioner review and edit before signoff).

What the Big 4 tools do NOT do

They do not autonomously conclude on operating effectiveness. They do not sign workpapers. They do not select samples in ways the audit-methodology team hasn't validated. They do not produce report narrative that goes to the client without licensed practitioner review. The professional-judgment layer stays with the practitioner in every documented Big 4 methodology.

What the mid-market can adopt from this

The pattern, not the scale. A mid-market cyber-attest practice using Fieldguide or vCISO Lite for Auditors can adopt the same bounded-AI-with-practitioner-judgment shape without the Big 4 engineering investment. The tools available in the mid-market segment now support extraction, classification, and triage capabilities that were exclusive to Big 4 tooling five years ago.

The important adoption is the workflow discipline, not the tool selection. A tool that offers autonomous conclusion generation is only defensible if the practitioner treats the output as advisory input to their own judgment; a tool that never offered such a capability makes the discipline default.

What to actually adopt in 2026

The peer-review-safe pattern for a mid-market cyber-attest practice in 2026 has three specific components:

  • Extraction and classification as first-line adoption.: The lowest peer-review-risk use of AI. Deploy on every engagement for evidence extraction (IPE reports, walkthrough documentation, exception dispositions). The efficiency gain is substantial (30-50% reduction in evidence-processing time) and the workpaper posture is straightforward — extraction is a mechanical operation with documented verification.
  • Triage as second-line adoption.: Use AI to flag likely exceptions and prioritize practitioner review. Adopt only when the practitioner's review discipline is strong — every flagged item independently evaluated, exception dispositions documented per item, verification sample of unflagged items to test the triage's negative-predictive value.
  • Analytical procedures as third-line adoption.: AI-computed analytics as input to risk assessment and sampling design. Peer-review-safe when the practitioner documents how the analytics informed the specific audit judgments they support. Do not treat the analytics as evidence in themselves; treat them as risk-assessment input.

Deliberately NOT recommended for adoption in 2026: AI-driven sample selection without deterministic methodology, AI-generated exception dispositions, AI-generated testing conclusions, and AI-drafted report narrative that goes to the client without practitioner-authored edits. These are the specific capabilities where the peer-review posture is unclear and the professional-standards violations are direct.

The bottom line

AI-assisted audit sampling is a mature enough discipline in 2026 that a cyber-attest practice can adopt it without material peer review risk — provided the adoption is bounded to the capabilities the professional standards support (mechanical assist, not judgment substitution) and the workpaper documentation reflects the discipline. The Big 4 pattern is the reference; mid-market platforms make the pattern accessible; the practitioner's judgment layer remains the professional-standards non-negotiable. Practices that adopt with discipline gain 30-50% efficiency on the AI-assisted procedures and preserve their peer-review posture. Practices that adopt without discipline are one AICPA cycle away from a material finding.

AI in cyber-attest, on the peer-review-safe pattern

vCISO Lite for Auditors deploys AI on the bounded pattern that survives peer review: extraction and classification of client-provided evidence, triage of items likely to be exceptions, analytical procedures for risk assessment, and AI-assisted draft workpaper narrative with mandatory practitioner review and edit before signoff. The AI does the mechanical work; the practitioner does the judgment. Every workpaper documents the AI's role, the practitioner's independent verification, and the professional judgment separately.

If your practice is adopting AI in the sampling workflow and wants the discipline built into the tool rather than layered on by policy, visit firm.vcisolite.com to see the AI-assist boundaries.

Where this matters next

Share this article:

Ready to build your security program?

See how easy it can be.