A training-program audit produces two scores, not one: design effectiveness grades the program as written, and operating effectiveness grades the program as run. Examiners weight what actually happened over what the binder promised, so a well-designed program that was never executed on schedule fails the same as a program with no design at all. The regulatory floor is thin: 31 CFR 1020.210(b)(4) requires role-tailored, current, recurring training with tracked completion, and the FFIEC manual asks for personnel coverage and documentation of training dates, with testing materials required only if testing is used. Everything past that, objectives, citation trails, sign-off gates, is best practice layered on top of a minimum standard. Evidence-absent findings are recorded separately from evidence-of-failure findings, and a failing cohort is read as a signal about the content before it is treated as a disciplinary matter for the class.
A training-program audit produces two separate assessments: design effectiveness, which grades the program as written, and operating effectiveness, which grades the program as run. The first tests the curriculum, its scoping, and its regulatory grounding. The second tests whether the curriculum reached the people it was written for, on the schedule it set, with a recorded consequence where it did not.
The two assessments are graded separately and are not substitutes for one another. A program graded on design alone carries no evidence about its operation, and a program graded on operation alone carries no evidence about whether its content is fit for the obligations it teaches.
Design effectiveness and operating effectiveness
An auditor reviewing a training program asks two distinct questions and treats an answer to one as no answer to the other. The first is whether the program is designed to work: whether the curriculum covers the right roles, on the right cadence, with the right regulatory grounding. The second is whether the program operated: whether the people in scope completed it, passed it, and were escalated when they did not.
Compliance practitioners call these design effectiveness and operating effectiveness. Design effectiveness (DE) grades the program on paper: the scoping matrix, the stated objectives, the materials, the sign-off gates, the citation trail back to the rule each module teaches. Operating effectiveness (OE) grades the program in the wild: completion records, quiz results and retakes, escalation paths that fired when someone missed a deadline, and whether the refresh cycle happened on the calendar it promised.
The two scores can diverge. A binder can describe a complete training program: role-tailored objectives, a citation trail to every regulation, a two-gate sign-off process, a documented escalation ladder. None of that evidences whether the people who needed the training received it, understood it, and were held to it when they did not. A program can score well on design and fail on operation, and an undocumented program that reaches its population on schedule can perform better in examination than a documented program that was not delivered.
Examination practice weights operation over description: what the records show outweighs what the program document states should occur. A design that was not executed does not mitigate the operational gap. An examiner who finds a well-scoped training matrix alongside a completion log with six-month gaps writes the finding against the gaps.
What design effectiveness measures
Design effectiveness assesses the program as specified, before any record of delivery is considered. It asks whether the plan, if followed exactly, would produce a defensible training program. Four elements carry most of the weight.
- The scoping matrix. Who is in scope, and why. A defensible matrix ties each role to the BSA/AML obligations and red flags that role actually touches, rather than handing everyone the same module regardless of exposure.
- Stated objectives. What each module is supposed to teach, written down before the content is built, not reverse-engineered from whatever got produced.
- The materials themselves. Content current to the institution's own policies and the regulations behind them, not a generic overview of money laundering that could belong to any institution.
- Sign-off gates and a citation trail. Evidence that someone with subject-matter authority reviewed the content for correctness, that a separate check reviewed it for production readiness, and that every regulatory claim in the material traces to a named source.
Design effectiveness is assessable from documents alone. An auditor can sit with the scoping matrix, the objectives, the materials, and the sign-off record, and form a defensible view of the program's design without interviewing a single employee. That property is also the score's limit: it establishes whether the program was built to work, and carries no evidence about whether it operated.
What operating effectiveness measures
Operating effectiveness is the same program observed after it left the design phase and met real people on real deadlines. It is not assessable from a policy binder. It requires pulling records and testing them against what the design promised.
- Completion records. Who was assigned the training, who finished it, and when, matched against the scoping matrix so gaps are visible rather than assumed away.
- Quiz results and retakes. Where testing is used, individual pass rates, and how the program handled a failed first attempt: a real retake and a remedial session, or a passing score on attempt two with no explanation on file.
- Escalation paths that actually fired. A program can document a ladder from reminder to manager notice to a formal out-of-compliance flag. Operating effectiveness asks whether that ladder was ever climbed, or whether every late completion got quietly absorbed with no record of escalation.
- Refresh actually happening. Annual training that ran on the calendar it promised, and targeted training that followed a new product launch or a regulatory change within a reasonable window, not at the next scheduled cycle months later.
Design work is typically performed once, by a party accountable for the program. Operating discipline has to hold every cycle, across every role, with each exception recorded. Operating effectiveness is accordingly the score that requires sustained maintenance, and it is the one an examiner samples first.
The regulatory floor, and where it stops
What the rule requires and what a well-run program does beyond the rule are distinct categories, and the sections below separate them.
The binding requirement sits in 31 CFR 1020.210(b)(4): covered institutions must provide training that is role-tailored, current, and recurring, and completion has to be tracked. That is the statutory floor for the training pillar. On its own it does not require learning objectives written to a particular standard, a quiz, or a formal sign-off process.
The FFIEC BSA/AML Examination Manual's exam procedures, in the BSA/AML Training section, ask examiners to confirm that personnel whose duties require BSA knowledge are covered, and to review documentation of training-session dates and of the materials used for training and testing, if testing is used. The conditional clause governs the requirement: the manual does not mandate a test, and asks for documentation of the testing materials only where testing occurs. A program with no quiz is not deficient on that point, while a program that administers a quiz and cannot produce the materials is.
Board training sits in a different category again. The manual describes board and senior-management training as a supervisory expectation, tied to the board's responsibility for an adequate program, and examiners cite its absence routinely. It is not, however, a hard mandate under 1020.210 the way role-tailored staff training is. Grading it as a statutory failure overstates the requirement. Grading it as a recurring, well-documented finding pattern is accurate.
A further training requirement sits outside the training pillar: institutions that participate in FinCEN's Section 314(a) information-sharing program must provide ongoing training on how to handle that information, under 31 CFR 1010.520. A program scoped to the five pillars can omit this obligation because the obligation is not filed under the training pillar.
These requirements are a floor rather than a ceiling. A program built only to the floor, with no objectives, no citation trail, and no sign-off gates, satisfies the rule and provides limited operational assurance. The additional practices exist because the statutory minimum was not written as a design standard.
Evidence-absent findings and evidence-of-failure findings
An auditor who requests a completion log and receives nothing has established an absence of evidence rather than a failed program. The two results are distinct, and recording them as the same finding misstates the condition in either direction.
If a program cannot produce the record, that is its own finding, worth citing on its own terms: the evidence trail is deficient, whatever the underlying training actually did. If a program produces the record and it shows a 40 percent completion rate, that is a different, worse finding: the program ran and it did not reach the people it was supposed to reach. Writing both up as "training deficiencies" without distinguishing them erases information the institution needs to fix the actual problem, which might be recordkeeping, or might be the program itself, or might be both.
The distinction runs in both directions. A compliance officer defending a program cannot rely on missing records as evidence that the training occurred as intended; an absent record supports no conclusion about compliance or about failure. The defensible position states what the available records establish and identifies what they do not.
A two-column self-check
The questions below are the ones an examination puts, sorted onto the score each one tests.
| Design effectiveness | Operating effectiveness |
|---|---|
| Does the scoping matrix name every role and tie it to a specific obligation? | Is a completion log available, matched against that matrix gap by gap? |
| Are learning objectives written down before the content, not after? | Do quiz results exist for every module that claims to test comprehension? |
| Does every regulatory claim in the material trace to a named source? | When someone fails a first attempt, is there a real retake on record, not just a passing second score? |
| Is there a documented sign-off from someone with subject-matter authority? | When someone misses a deadline, did the escalation ladder actually move, with a record at each step? |
| Does the program name what triggers a refresh: a new product, a new segment, a rule change? | Did the last refresh happen on the calendar the design promised, or later? |
The two columns are not graded against each other. A program can score well on the left column and fail on the right, and the right column is where an examination typically begins. Where the columns disagree, the operating-effectiveness column governs the finding.
A further principle governs how a poor result is read. Where a training cohort fails at scale, the more defensible diagnosis concerns the material rather than the cohort. A first-attempt pass rate that comes back low across a whole population is a signal about the content, the delivery, or the timing rather than about the individuals who sat the module. The individual pass bar and the cohort pass bar answer different questions: one asks whether a person understood the material, the other asks whether the material was teachable. A program that investigates the module before scheduling remedial sessions is performing the operating-effectiveness analysis; a program that begins with remedial sessions has assigned the result to the wrong half of the score.