Control testing is the set of audit procedures used to determine whether a control is designed to meet its objective (design effectiveness) and whether it actually operated as designed over a period (operating effectiveness). Published methodology, chiefly PCAOB AS 2201 and AS 2315, the AICPA Audit Guide: Audit Sampling, and the IIA Global Internal Audit Standards, describes three approaches to operating-effectiveness testing: a test of one (for automated controls under effective IT general controls), a sample sized from the tolerable deviation rate, expected deviation rate and acceptable risk, or a test of the full population. Every test depends on a complete and accurate population, and every deviation is documented and evaluated for cause.
Control testing is the application of audit procedures to a control to obtain evidence about whether the control is suitably designed and whether it operated effectively during a defined period. The method is common to external auditors of internal control over financial reporting, internal audit functions, and the independent testers who review compliance programs such as a BSA/AML program. The accepted method is set out in published standards: the Public Company Accounting Oversight Board's AS 2201 and AS 2315, the AICPA Audit Guide: Audit Sampling, and the Institute of Internal Auditors' Global Internal Audit Standards (2024).
Design effectiveness and operating effectiveness
Control testing addresses two separate questions, and a control can pass one and fail the other.
Design effectiveness asks whether the control, if performed as described by a person with the necessary authority and competence, would meet its objective and prevent or detect the error it targets. PCAOB AS 2201, paragraph .42, states this test, and paragraph .43 notes that a mix of inquiry, observation and inspection of documentation, typically performed as a walkthrough, is ordinarily sufficient to evaluate design.
Operating effectiveness asks whether the control actually operated as designed, consistently, throughout the period under review, and whether the person performing it had the necessary authority and competence (AS 2201, paragraph .44). Paragraph .45 lists the procedures used: inquiry, observation, inspection of relevant documentation, and re-performance. Paragraph .50 ranks those procedures by the evidence they ordinarily produce, from least to most: inquiry, observation, inspection, re-performance. Inquiry alone does not establish that a control operated.
The distinction applies outside financial reporting. In a compliance program, a written escalation procedure that routes high-risk alerts to a senior reviewer may be well designed; testing a sample of alerts shows whether the routing occurred in practice. The reference on training design versus operating effectiveness applies the same distinction to compliance training.
The three approaches to operating-effectiveness testing
| Approach | When it is used | Basis in published methodology |
|---|---|---|
| Test of one | A fully automated control, where the system applies the same logic to every item and IT general controls over program changes are effective. | AS 2201, Appendix B, "Benchmarking of Automated Controls" (paragraphs .B28 to .B33). |
| Sampling | A manual or partly manual control applied to a large population, where examining every item is not practical. | AS 2315 (audit sampling); AICPA Audit Guide: Audit Sampling. |
| Full-population (100 percent) testing | The population is available in electronic form and the control attribute can be tested by a data query or script against every item. | IIA Standard 14.1, Considerations for Implementation; AS 2315, paragraph .01 (sampling is defined as testing less than 100 percent). |
A test of one rests on the premise that an automated control performs consistently unless the program is changed, which is why the IT general controls over change management are tested alongside it. Full-population testing removes sampling risk for the attribute tested, although it does not remove the need to confirm that the population itself is complete. The IIA's Standard 14.1 guidance states that internal auditors should consider whether to test a complete data population or a representative sample, and notes that data analysis software facilitates testing of complete or targeted populations.
How sample sizes are chosen
PCAOB AS 2315 defines audit sampling as the application of an audit procedure to less than 100 percent of the items in a population (paragraph .01). For tests of controls, paragraph .38 identifies three factors that determine sample size:
- Tolerable deviation rate: the maximum rate of deviation from the control that the tester would accept and still conclude the control is effective. A lower tolerable rate requires a larger sample.
- Expected (likely) deviation rate: the rate of deviation the tester expects to find, based on prior results or the understanding of the control. A higher expected rate requires a larger sample.
- Allowable risk of assessing control risk too low: the risk that the sample supports a conclusion that the control is effective when it is not (the sampling risk defined in paragraph .12). A lower allowable risk requires a larger sample.
The AICPA Audit Guide: Audit Sampling translates these factors into sample-size tables for statistical and nonstatistical attribute sampling. The arithmetic behind the tables is binomial. For example, where no deviations are expected, a tolerable deviation rate of 5 percent and an allowable risk of 5 percent produce a sample of about 59 items, because the probability of drawing 59 compliant items from a population that actually deviates at 5 percent (0.95 raised to the 59th power) is just under 5 percent. AS 2315 permits nonstatistical sampling, but paragraph .38 states that a properly applied nonstatistical approach ordinarily produces a sample comparable to, or larger than, an efficiently designed statistical sample.
Selection method matters as much as size. AS 2315, paragraph .24, requires that sample items be selected so that the sample can be expected to be representative of the population, meaning every item has an opportunity to be selected; it names haphazard and random-based selection as two means of obtaining such a sample, and the AICPA guide also describes systematic selection with a random start. A sample drawn only from the most recent month, or only from items a control owner supplies, is not representative of the period.
The IIA standards require the method to be recorded. Under the Considerations for Implementation to Standard 13.6 (Work Program), when sampling is used the work program should state the sampling methodology, the population, the sample size, and whether results can be projected to the population.
Data-quality checks before testing
Every sampling or full-population conclusion depends on the population being complete and accurate. The standard preliminary procedures are:
- Completeness: reconcile the population to an independent source, such as a general ledger total, a core-system record count, or a count from the source system for the same period.
- Accuracy: trace a selection of population records back to source documents or source-system screens, and confirm that the fields used in the test (dates, amounts, risk ratings) are populated and correct.
- Period and scope: confirm that the extract covers the full test period and every business line or entity in scope.
- Provenance: record who produced the extract, from which system, with what query, and on what date, so the population can be reproduced.
IIA Standard 14.1 requires internal auditors to gather information that is relevant, reliable and sufficient, and it describes reliability as strengthened when the information is obtained directly by the internal auditor or from an independent source, is corroborated, or comes from a system with effective controls. In transaction monitoring, the New York Department of Financial Services' Part 504 rule (3 NYCRR 504.3) requires identification of all relevant data sources and validation of the integrity, accuracy and quality of data flowing into monitoring and filtering programs, which illustrates why a monitoring test begins with data lineage rather than with alerts.
Control testing in compliance examinations and independent testing
The FFIEC BSA/AML Examination Manual applies the same method to compliance programs. Its BSA/AML Independent Testing section expects independent testing to be risk-based and to include transaction testing, and its suspicious activity reporting procedures describe transaction testing of monitoring systems and reporting processes, with the sample size and selection based on identified weaknesses and the institution's overall risk profile. The program rule for banks, 31 CFR 1020.210, requires independent testing for compliance as a program component. The BSA/AML independent testing reference describes the scope of that review; control testing is the evidence-gathering technique inside it.
Internal audit functions follow the Global Internal Audit Standards, which the IIA released in January 2024 and which took effect on January 9, 2025. Standard 14.2 requires internal auditors to compare the evaluation criteria (what the control should do) with the condition (what the testing found), and treats any difference as a potential finding.
Exceptions and how they are documented
An exception, also called a deviation, is a sample item where the control did not operate as designed. The accepted handling has four steps.
- Confirm the exception. Obtain management's explanation and any evidence that the control operated in a way the test did not capture. AS 2315, paragraph .40, directs that an item the tester cannot examine, for example because documentation is missing, is ordinarily treated as a deviation.
- Evaluate frequency. Compare the sample deviation rate with the tolerable rate. If the deviation rate, adjusted for sampling risk, exceeds the tolerable rate, the sample does not support a conclusion that the control is effective. Where a methodology permits the sample to be extended, the workpaper records the rationale for the extension.
- Evaluate cause and nature. AS 2315, paragraph .42, requires consideration of the qualitative aspects of deviations, including whether they are errors or irregularities. IIA Standard 14.3 requires internal auditors to work with management to identify root causes where possible, determine potential effects, and prioritize each finding by significance.
- Document. The workpaper records the criteria, the population and its completeness check, the sample method and size, each exception with its evidence, management's response, and the conclusion. IIA Standard 14.6 requires documentation such that an informed, competent person could repeat the work and reach the same results.
A finding produced by control testing then enters an issue-management process with an owner, a corrective action and a target date, and is retested after remediation. That process is described in the reference on compliance remediation and issue management.
Primary sources
- PCAOB AS 2201, An Audit of Internal Control Over Financial Reporting: Paragraphs .42 to .45 (testing design and operating effectiveness), .50 (procedures ranked by evidence), Appendix B .B28 to .B33 (benchmarking of automated controls).
- PCAOB AS 2315, Audit Sampling: Paragraph .01 (definition), .12 (sampling risk), .24 (representative selection), .38 (sample-size factors for tests of controls), .40 and .42 (evaluation of deviations).
- AICPA Audit Guide: Audit Sampling: The AICPA's practice guide to statistical and nonstatistical sampling, including sample-size tables for tests of controls.
- The IIA, Global Internal Audit Standards (2024): Standards 13.6 (work program), 14.1 (relevant, reliable, sufficient information), 14.2 and 14.3 (findings and root cause), 14.6 (documentation); effective January 9, 2025.
- FFIEC BSA/AML Examination Manual, BSA/AML Independent Testing: Risk-based independent testing, including transaction testing of the program.
- 31 CFR 1020.210: Anti-money laundering program requirements for banks, including independent testing for compliance.
- 3 NYCRR 504.3: New York DFS transaction monitoring and filtering program requirements, including data identification and validation.