7 Prove? Evidence Collection and Testing
You can read your way to a suspicion. You can only evidence your way to an opinion.
🗂️ Tessera Case: Full access granted
The email lands on a Monday morning, and it changes the engagement. Tessera’s CTO has signed the access schedule. You have credentials for the AWS environment with read-only audit permissions, exported log streams for the last twelve months, the complete HR leavers list, calendar slots with the head of people and the IT operations lead, and a standing invitation to walk the Perth office. For six weeks you have been reading about Tessera. From today you are looking at it. The preliminary opinion you wrote at the close of Act I, promising on paper, not yet proven in practice, now meets the only thing that can upgrade or overturn it: the evidence itself.
7.1 The act break: from suspicion to proof
Act I ended with a disciplined confession. You had read the ISMS, the risk register, the Statement of Applicability, and every policy beneath them, and formed an honest opinion: designed well, operating unconfirmed. Every gap you named at the close of Chapter 6 was a control you had read about and never seen run. Full access is the thing that closes those gaps.
Notice the posture change. In Act I you accepted documents because documents were all you had; your scepticism was sharp but pointed at the page. Now it has somewhere to land. A policy that says access is revoked within twenty-four hours of a leaver’s last day is, on the page, a promise. A log entry showing the account disabled at 16:11 that day is the promise kept; one showing it still active a week later is the promise broken. Both are evidence, and only full access could have produced them.
Here is the trap that comes with the door opening. Access feels like proof, and that feeling is the new danger. A dashboard, a log, a configuration export, a confident engineer: each can look like evidence and still fail to substantiate anything. This chapter is the craft that turns raw access into a defensible opinion, which evidence to gather, how much, how to keep it so it survives challenge, and which procedures prove a control operates rather than merely exists.
7.2 Four kinds of evidence
Every conclusion rests on evidence, and evidence is not all the same stuff. The profession sorts it into four kinds, and knowing which kind you are holding tells you how much it can carry.
Observational evidence is what you see yourself: a leaver’s badge denied at the side entrance, a privileged elevation granted, used, and revoked beside you. It is the most direct evidence that a control operates, with a hard limit: it only proves operation for the moment you watched. A control that ran perfectly under your gaze can fail the day after you leave.
Documentary evidence is the records the organisation and its systems produce: the access-revocation log, the CloudTrail export, the signed offboarding checklist, the supplier’s SOC 2 report. This is the bulk of what Act II gathers, and its reliability depends on its source and the controls over how it was made. A system-generated log nobody can alter is strong; a spreadsheet a manager typed last week is weak.
Analytical evidence is your own work: reconciling the asset register against the live estate, re-running the calculation behind a tenant-isolation check, computing a deviation rate from a sample. It is powerful precisely because you produced it on terms you set, independent of the organisation’s goodwill.
Oral evidence is what people tell you: the interview, the attestation, the corridor explanation. It is the weakest kind alone, because memory is selective and self-interest is real, but it is also where almost every other kind of evidence begins. A good interview does not conclude anything. It points you at the document, the log, or the observation that will.
There is a persuasiveness hierarchy worth memorising. Evidence from outside the organisation beats evidence from inside it; what you obtain yourself beats what the client hands you; the original record beats a copy, a copy beats a summary, and a contemporaneous record beats a reconstruction. Keep that ordering in mind and you will feel the right scepticism for each artefact you are handed.
7.3 Sufficiency and appropriateness
Two words do most of the judgement work, and they answer different questions. Sufficiency is the quantity of evidence: is there enough of it to support the conclusion? Appropriateness is the quality: is it relevant to the assertion, and is it reliable?
Appropriateness has two halves, and students trip on the second. Relevance asks whether the evidence actually tests what you are trying to prove: a clean access-revocation log is irrelevant if your assertion is about badge access. Reliability asks whether you can trust the source: was the record produced by a system whose integrity is controlled, at or near the event, by someone without motive to distort it?
The mistake is to chase one axis and ignore the other. A mountain of irrelevant evidence proves nothing, however large the pile; a single relevant item is a start but not a conclusion, because sufficiency is a property of the population, not of one match. High assurance needs both, enough of the right kind, and the craft of this chapter is deciding, control by control, how much of what kind will defend the opinion.
Three artefacts look like evidence and are not, and each sinks an audit that trusted it. A screenshot with no timestamp, source, or chain of custody is a picture someone made; it could be staged, old, or edited. A self-attestation, the manager’s confident “we do that,” is oral evidence, the weakest kind, and worthless uncorroborated. An AI summary standing in for the underlying logs is a secondary account of a primary record, and summarisation always loses the detail that decides whether a control actually ran.
Then the sufficiency trap, subtler and more common. One corroborating item does not prove a control operates across a population. A single leaver whose access was revoked on time is reassuring; it is not evidence the control works for every leaver. Sufficiency asks how many, across how much, with what confidence. A single match answers none of those questions, and treating it as if it did is the same error as trusting a screenshot: mistaking the appearance of evidence for evidence itself.
7.4 Sampling: why you cannot test everything, and how to size what you do
Chapter 4 decided where to dig. Sampling decides how much. You cannot test every login, every leaver, every firewall rule across a year, so you test a defensible subset and reason about the whole from it. Done badly, sampling is a licence to miss the thing that matters. Done well, it is the most rigorous thing an auditor does, because the rigour is in the design, not the counting.
The first rule, and the one most often broken, is that a sample proves nothing about a population you have not completely defined. To say something about “every leaver in the period,” you need the complete list of every leaver, reconciled to a source you trust, such as payroll. A sample drawn from a convenient subset, the leavers IT happens to remember, proves something about that subset and nothing else. Population completeness comes first, always.
For testing controls, the tool is attribute sampling. You are not measuring a sum of money; you are estimating a rate, the proportion of times a control deviated from its objective. The leaver’s access was either revoked in the policy window or it was not, each test a yes or a no, and the sample lets you estimate the deviation rate across the whole population.
Here is where the discipline bites. The sample size is not a number you feel; it is derived from three things. The tolerable deviation rate is how much failure you can accept and still conclude the control works: tighter tolerance, bigger sample. The expected deviation rate is what you think you will find: the worse you expect, the more you test. The assurance level is the confidence the engagement requires: certification work demands high assurance, and high assurance means a larger sample. Set those three and the size follows. Pick a round number because it sounds thorough, and you have a sample you cannot defend.
7.5 The five testing procedures
You have your population and your sample. Now you need procedures that actually extract evidence from each item. The audit toolbox has five, rising in persuasive power as you move down the list.
Inquiry is asking. You ask the IT operations lead how offboarding runs. Inquiry is where nearly every test begins, because it tells you where the evidence lives, and where nearly every test must not end, because what someone tells you is oral evidence, the weakest kind. “We disable accounts within twenty-four hours” is a lead, not a finding.
Observation is watching a control run: you attend an offboarding and observe the steps being performed. It proves the control operated at the moment you watched, and only that moment. It is excellent for physical and process controls and useless for anything that has to hold across the year.
Inspection is examining records and assets: the access-revocation log, the signed checklist, the CloudTrail export, the configuration file. Inspection is the workhorse of IS audit, and most of Act II is inspection done well, the right record, the right period, read sceptically.
Re-performance is re-running the control yourself, independently. Rather than trust the reconciliation was done, you redo it; rather than trust the access review caught the orphaned account, you pull the live IAM state and look yourself. It is powerful because it does not depend on the organisation having done anything correctly: you are checking the outcome, not the report of the outcome.
Re-computation is the numeric cousin: you recompute the calculation a control is supposed to perform, whether a deviation rate, a key-rotation interval, or a backup-retention tally. It turns an assertion into a number you can verify.
The order matters, because persuasive value rises as you go. An opinion built on inquiry alone is built on talk; one built on inspection, re-performance, and re-computation is built on what the evidence shows. State the control objective, pick the procedure that tests it at the assurance the engagement demands, and reach for the stronger procedure whenever the control is critical.
This is the kind of structured work an AI drafts fluently, and so the kind you must evaluate hardest. Ask it to draft the work package for testing Tessera’s access-revocation control and it returns a clean document in seconds: sampling plan, test steps, results table. Read it as an evaluator, because it is wrong in two instructive ways.
First, the sample size. The AI proposes “test 25 leavers.” A round, confident number with no basis. It carries no population, because the AI never asked how many leavers there were, and no tolerable deviation rate, expected rate, or assurance level, because it never fixed them. The number was chosen for how it reads, not derived from what the opinion requires. Your correction is to set the parameters the AI skipped: the complete HR leavers list as the population, a tight tolerable deviation rate, a low expected rate, and the high assurance a certification engagement demands. The derivation lands higher than twenty-five, and against a small population backing a control this critical, the defensible answer may be to test all of them. The AI picked a number. You derived one.
Second, the test steps. The AI drafts three. “Verify that access was revoked within twenty-four hours of the leaver’s last day.” Performable, if the revocation timestamps exist, and they do. “Confirm the leaver’s manager was notified within the policy window.” Not performable: Tessera produces no record of that notification, so the step collapses into pure inquiry and proves nothing. “Determine whether the leaver still holds production access.” Conflates now with then. For a departure three months ago, you cannot observe the state on their last day from a current snapshot; you depend on the historical audit trail, not a live check. The AI wrote steps that read as tests. Two of the three cannot be executed as written.
The AI produced the draft quickly and well-structured, and that is the trap. The sample size was a guess dressed as a method; the test steps were plausible sentences that do not survive contact with the evidence Tessera actually keeps. You supplied the parameters, rewrote the steps to what the trail can support, and turned a generic work package into one that will hold up under review. That is the drafter-to-evaluator shift, on the artefact that matters most this week.
7.6 Chain of custody
Evidence you gather has to survive challenge, sometimes years later, in a certification review, a board meeting, or a regulator’s inquiry. That survival depends on chain of custody: a complete, defensible record of who collected what, when, from where, by what method, and what happened to it afterwards. An extract with no provenance is just a file someone made, and “someone made a file” is not a basis for an opinion.
The mechanics are unglamorous and non-negotiable. Every extract gets a hash, typically SHA-256, recorded at capture so any later alteration is detectable. Copies are read-only. Each artefact carries its source system, the query or export that produced it, the date and time in a clear timezone, and the name of the auditor who took it. Transfers between people, the log moving from the analyst who pulled it to the lead who filed it, are themselves logged. The whole chain is sealed into the working papers, and another auditor following them should be able to trace every conclusion back to a specific artefact captured on a specific day.
I learned this the hard way. The first time full access landed, I went straight for the logs and forgot to record how I got them. Three weeks later a reviewer asked where a particular extract had come from, and I could not say. The extract was fine. Its provenance was gone. An evidence extract with no provenance is not evidence; it is an unsupported claim, which is exactly what we refuse to accept from the client. The discipline cuts as deep for us as it does for them, and it should.
7.7 Auditing AI: can you trust a log an AI produced?
This is the question earlier chapters promised to take up in full, and full access is where it lands. Tessera runs AI in two places inside the boundary you drew in Chapter 4: a support assistant that triages alerts, and an anomaly detector in the logging stack. Both were named as evidence gaps at the close of Act I. Both now produce logs, and the question is whether those logs count as evidence you can stand behind.
Take the assistant. When it auto-resolves a low-priority alert, the log entry recording the resolution is, in a real sense, AI-produced. Can you trust it? Treat it as any system log, then add a layer. A log’s reliability always depends on the controls over how it was generated, but an AI’s record of its own action has a failure mode a deterministic system’s log does not: it can confabulate, truncate, or summarise away the detail that would tell you whether the action was correct. The tidy “alert reviewed, no action required” may faithfully record a sound decision, or paper over a misjudgement; you cannot tell from the summary. Corroborate against the underlying system’s own raw audit trail, the original alert event and the downstream actions, and treat the AI’s log as one account among several, not the account of record.
Now the harder question, the one Chapter 5 raised: how do you test a control that an AI runs? The assistant triaging alerts is itself a detective control, and testing whether it operates means measuring its decisions against ground truth. Did it escalate what it should have? Auto-close what it should not have? That needs a labelled set, a sample of alerts where you know independently what the right outcome was, so you can measure the assistant’s deviation rate the way you measure any control’s. The five procedures apply: inquire how it was tuned, observe it on a live alert, inspect its decision log, then re-perform by re-judging the cases yourself against the labels. Re-performance, for an AI-run control, is exactly the drafter-to-evaluator posture this book argues for: the machine decided, and you are the evaluator who decides whether the decision holds.
The conclusion is steadier than the anxiety around AI usually allows. You can audit a control an AI runs, and use evidence an AI helped produce, on one condition: never let the AI’s account of itself be the only account. Corroborate against the raw trail, test its decisions against ground truth, and treat its summaries as leads to the primary evidence rather than the primary evidence itself. An AI’s log is admissible. It is not self-authenticating.
7.8 Interviewing for evidence
Interviews are oral evidence, the weakest kind alone, and also where almost every stronger kind begins. The skill is running them so they open doors instead of closing opinions.
Prepare before you sit down. Know the control objective you are testing and the evidence you seek. Ask open questions: not “do you revoke access within twenty-four hours,” which invites the answer you want, but “walk me through what happens when someone leaves.” Let silence do the work; people fill it, and what they fill it with is often the detail that matters.
The rule that turns an interview into an audit: never accept “we do that” without the record. Every claim about a process is a hypothesis until you produce the artefact that proves it. “Show me the last three times this ran” is the sentence that converts an attestation into a test. When the head of people says offboarding is always completed, ask for the last three checklists; when the SRE says privilege elevation is time-boxed, ask for the elevation log. Note where the claim and the record diverge: that is where the findings live.
Corroborate across interviewees. The IT operations lead and the head of people should describe the same offboarding in compatible terms. Where their accounts differ, one describes aspiration and the other practice, and the documents will tell you which. A good interview produces leads. The records produce the conclusions.
🗂️ Tessera Case: A sampling plan and a chain-of-custody log for access revocation
The artefact for this chapter is the test that closes the single most important gap Act I named: confirmation that the access-revocation control operates for every leaver. The control spans A.6.5 (responsibilities after termination or change) and A.8.3 (information access restriction), and its objective is testable: for every leaver in the period, all logical access to Tessera systems is revoked within twenty-four hours of the last working day, and the revocation is logged.
The population is every leaver in the twelve-month period, sourced from the complete HR list and reconciled against payroll to confirm completeness. There are forty-eight. The sample is derived, not chosen: a tight tolerable deviation rate, a low expected rate, and the high assurance the engagement demands. Against forty-eight leavers backing a control this critical, the defensible answer is to test all of them, because sampling a fraction and extrapolating buys little when the whole population is small enough to test in full.
The procedure runs per leaver: inspect the HR termination record for the last working day, inspect the IAM and single-sign-on audit logs for the revocation event, recompute the elapsed time, and flag any deviation over twenty-four hours. For every deviation, inquire for the cause and corroborate against the offboarding checklist.
Every artefact the test touches enters a chain-of-custody log. A representative slice:
| Artefact | Description | Source | Captured | By | Hash (prefix) |
|---|---|---|---|---|---|
| EVD-07-01 | HR leavers list, full period | HR system export | 2026-06-15 09:12 AWST | M. Borck | 3f9a… |
| EVD-07-02 | IAM revocation log, leaver A | AWS CloudTrail | 2026-06-15 09:41 AWST | M. Borck | b1c7… |
| EVD-07-03 | SSO session log, leaver A | Identity provider | 2026-06-15 09:46 AWST | M. Borck | 8e2d… |
| EVD-07-04 | Offboarding checklist, leaver A | HR document store | 2026-06-15 10:03 AWST | M. Borck | 4a55… |
Each row is one extract, hashed at capture and sealed into the working papers. When the test finds forty-five of forty-eight revocations inside the window and three outside, that finding is traceable, leaver by leaver and artefact by artefact, back to the records that produced it. That traceability is what makes the conclusion defensible rather than merely asserted. You can run the same test against the live companion site at tessera.locoensayo.org.
Act II’s locker stops being narrative and becomes working papers in earnest. Lock in three things this week.
First, a sampling plan for one Tessera control of your choosing: the population defined and reconciled for completeness, the control objective stated as a testable assertion, the tolerable and expected deviation rates and the assurance level stated explicitly, and the sample size derived from them rather than chosen. Second, a chain-of-custody record for every artefact the test touches: artefact, source, capture time, extractor, hash. Third, an interview-notes template that separates, on every line, what the interviewee claimed from what you subsequently corroborated, so that aspiration and practice never blur on the page.
These three are not exercises. By Week 12 they are the spine of your Assessment 3 working papers, and the discipline they enforce, sufficiency derived not assumed, custody recorded not implied, claims corroborated not trusted, is what makes an opinion defensible.
7.9 The question, answered
Can we substantiate our suspicions? Yes, when the evidence is sufficient in quantity and appropriate in quality, gathered under a chain of custody that survives challenge, and tested by procedures strong enough to prove operation rather than existence. Act I left you with suspicion, calibrated and named. Act II gives you the means to convert it, item by item, into proof. The access-revocation control is no longer a promise on a page. It is forty-eight tested log entries, each traceable to a hashed extract captured on a known day, and the deviation rate they reveal is a number you can defend.
That is the shift this chapter enacts, and the shift the rest of Act II runs on. The polished policy meets its operational reality, and every suspicion you earned in Act I either hardens into evidence or dissolves. The evidence has started to speak. The question now is whether the policy behind the controls was ever any good, because a control can run perfectly and still serve a policy whose intent has drifted from its operation.
Next week the question turns from can you prove it works? to does the policy itself hold up? The question is Comply?