Skip to content
Reference

Control Testing Methods: Design, Operating Effectiveness and Sampling

The short version

Control testing is the set of audit procedures used to determine whether a control is designed to meet its objective (design effectiveness) and whether it actually operated as designed over a period (operating effectiveness). Published methodology, chiefly PCAOB AS 2201 and AS 2315, the AICPA Audit Guide: Audit Sampling, and the IIA Global Internal Audit Standards, describes three approaches to operating-effectiveness testing: a test of one (for automated controls under effective IT general controls), a sample sized from the tolerable deviation rate, expected deviation rate and acceptable risk, or a test of the full population. Every test depends on a complete and accurate population, and every deviation is documented and evaluated for cause.

Control testing is the application of audit procedures to a control to obtain evidence about whether the control is suitably designed and whether it operated effectively during a defined period. The method is common to external auditors of internal control over financial reporting, internal audit functions, and the independent testers who review compliance programs such as a BSA/AML program. The accepted method is set out in published standards: the Public Company Accounting Oversight Board's AS 2201 and AS 2315, the AICPA Audit Guide: Audit Sampling, and the Institute of Internal Auditors' Global Internal Audit Standards (2024).

Design effectiveness and operating effectiveness

Control testing addresses two separate questions, and a control can pass one and fail the other.

Design effectiveness asks whether the control, if performed as described by a person with the necessary authority and competence, would meet its objective and prevent or detect the error it targets. PCAOB AS 2201, paragraph .42, states this test, and paragraph .43 notes that a mix of inquiry, observation and inspection of documentation, typically performed as a walkthrough, is ordinarily sufficient to evaluate design.

Operating effectiveness asks whether the control actually operated as designed, consistently, throughout the period under review, and whether the person performing it had the necessary authority and competence (AS 2201, paragraph .44). Paragraph .45 lists the procedures used: inquiry, observation, inspection of relevant documentation, and re-performance. Paragraph .50 ranks those procedures by the evidence they ordinarily produce, from least to most: inquiry, observation, inspection, re-performance. Inquiry alone does not establish that a control operated.

The distinction applies outside financial reporting. In a compliance program, a written escalation procedure that routes high-risk alerts to a senior reviewer may be well designed; testing a sample of alerts shows whether the routing occurred in practice. The reference on training design versus operating effectiveness applies the same distinction to compliance training.

The three approaches to operating-effectiveness testing

ApproachWhen it is usedBasis in published methodology
Test of oneA fully automated control, where the system applies the same logic to every item and IT general controls over program changes are effective.AS 2201, Appendix B, "Benchmarking of Automated Controls" (paragraphs .B28 to .B33).
SamplingA manual or partly manual control applied to a large population, where examining every item is not practical.AS 2315 (audit sampling); AICPA Audit Guide: Audit Sampling.
Full-population (100 percent) testingThe population is available in electronic form and the control attribute can be tested by a data query or script against every item.IIA Standard 14.1, Considerations for Implementation; AS 2315, paragraph .01 (sampling is defined as testing less than 100 percent).

A test of one rests on the premise that an automated control performs consistently unless the program is changed, which is why the IT general controls over change management are tested alongside it. Full-population testing removes sampling risk for the attribute tested, although it does not remove the need to confirm that the population itself is complete. The IIA's Standard 14.1 guidance states that internal auditors should consider whether to test a complete data population or a representative sample, and notes that data analysis software facilitates testing of complete or targeted populations.

How sample sizes are chosen

PCAOB AS 2315 defines audit sampling as the application of an audit procedure to less than 100 percent of the items in a population (paragraph .01). For tests of controls, paragraph .38 identifies three factors that determine sample size:

The AICPA Audit Guide: Audit Sampling translates these factors into sample-size tables for statistical and nonstatistical attribute sampling. The arithmetic behind the tables is binomial. For example, where no deviations are expected, a tolerable deviation rate of 5 percent and an allowable risk of 5 percent produce a sample of about 59 items, because the probability of drawing 59 compliant items from a population that actually deviates at 5 percent (0.95 raised to the 59th power) is just under 5 percent. AS 2315 permits nonstatistical sampling, but paragraph .38 states that a properly applied nonstatistical approach ordinarily produces a sample comparable to, or larger than, an efficiently designed statistical sample.

Selection method matters as much as size. AS 2315, paragraph .24, requires that sample items be selected so that the sample can be expected to be representative of the population, meaning every item has an opportunity to be selected; it names haphazard and random-based selection as two means of obtaining such a sample, and the AICPA guide also describes systematic selection with a random start. A sample drawn only from the most recent month, or only from items a control owner supplies, is not representative of the period.

The IIA standards require the method to be recorded. Under the Considerations for Implementation to Standard 13.6 (Work Program), when sampling is used the work program should state the sampling methodology, the population, the sample size, and whether results can be projected to the population.

Data-quality checks before testing

Every sampling or full-population conclusion depends on the population being complete and accurate. The standard preliminary procedures are:

IIA Standard 14.1 requires internal auditors to gather information that is relevant, reliable and sufficient, and it describes reliability as strengthened when the information is obtained directly by the internal auditor or from an independent source, is corroborated, or comes from a system with effective controls. In transaction monitoring, the New York Department of Financial Services' Part 504 rule (3 NYCRR 504.3) requires identification of all relevant data sources and validation of the integrity, accuracy and quality of data flowing into monitoring and filtering programs, which illustrates why a monitoring test begins with data lineage rather than with alerts.

Control testing in compliance examinations and independent testing

The FFIEC BSA/AML Examination Manual applies the same method to compliance programs. Its BSA/AML Independent Testing section expects independent testing to be risk-based and to include transaction testing, and its suspicious activity reporting procedures describe transaction testing of monitoring systems and reporting processes, with the sample size and selection based on identified weaknesses and the institution's overall risk profile. The program rule for banks, 31 CFR 1020.210, requires independent testing for compliance as a program component. The BSA/AML independent testing reference describes the scope of that review; control testing is the evidence-gathering technique inside it.

Internal audit functions follow the Global Internal Audit Standards, which the IIA released in January 2024 and which took effect on January 9, 2025. Standard 14.2 requires internal auditors to compare the evaluation criteria (what the control should do) with the condition (what the testing found), and treats any difference as a potential finding.

Exceptions and how they are documented

An exception, also called a deviation, is a sample item where the control did not operate as designed. The accepted handling has four steps.

  1. Confirm the exception. Obtain management's explanation and any evidence that the control operated in a way the test did not capture. AS 2315, paragraph .40, directs that an item the tester cannot examine, for example because documentation is missing, is ordinarily treated as a deviation.
  2. Evaluate frequency. Compare the sample deviation rate with the tolerable rate. If the deviation rate, adjusted for sampling risk, exceeds the tolerable rate, the sample does not support a conclusion that the control is effective. Where a methodology permits the sample to be extended, the workpaper records the rationale for the extension.
  3. Evaluate cause and nature. AS 2315, paragraph .42, requires consideration of the qualitative aspects of deviations, including whether they are errors or irregularities. IIA Standard 14.3 requires internal auditors to work with management to identify root causes where possible, determine potential effects, and prioritize each finding by significance.
  4. Document. The workpaper records the criteria, the population and its completeness check, the sample method and size, each exception with its evidence, management's response, and the conclusion. IIA Standard 14.6 requires documentation such that an informed, competent person could repeat the work and reach the same results.

A finding produced by control testing then enters an issue-management process with an owner, a corrective action and a target date, and is retested after remediation. That process is described in the reference on compliance remediation and issue management.

Primary sources

Common questions

What is the difference between a test of design and a test of operating effectiveness?
A test of design determines whether a control, performed as described, would meet its objective; a walkthrough is usually sufficient (PCAOB AS 2201, paragraphs .42 and .43). A test of operating effectiveness determines whether the control actually operated as designed throughout the period, using inspection and re-performance on a sample or the full population (AS 2201, paragraphs .44 and .45).
How is a sample size for a test of controls determined?
Under PCAOB AS 2315, paragraph .38, the tester considers the tolerable deviation rate, the expected deviation rate, and the allowable risk of assessing control risk too low. The AICPA Audit Guide: Audit Sampling converts those factors into sample-size tables. With zero expected deviations, a 5 percent tolerable rate and 5 percent risk produce a sample of about 59 items.
When is a test of one acceptable?
A test of one is used for fully automated controls, on the premise that a system applies the same logic to every transaction unless the program changes. It depends on effective IT general controls over program changes, which is why benchmarking of automated controls under AS 2201, Appendix B, is paired with testing of those general controls.
Does full-population testing replace sampling?
Where the population is available electronically and the control attribute can be tested by query, full-population testing removes sampling risk for that attribute. It does not remove the need to confirm the population is complete and accurate, and it does not replace judgment-based review of items that require evaluation, such as the quality of an investigation narrative.
How are control-testing exceptions documented?
Each exception is confirmed with management, compared against the tolerable rate, evaluated for cause and nature, and recorded in the workpaper with the supporting evidence and management's response. IIA Standards 14.3 and 14.6 require evaluation of significance and root cause and documentation from which an informed person could repeat the work and reach the same result.
About this library

This reference library is maintained by Rupture Labs. We perform BSA/AML independent testing and automated control testing against every record, with each requirement cited to its rule. Talk to a practitioner.