Audit of AI and Automated Decisions
When a model influences a payment, an assessment, a credit decision or an eligibility outcome, somebody has to be able to explain why it produced that answer. This is assurance work, and it extends naturally from what an audit practice already does: testing that a control operates as management claims. The subject happens to be a model rather than a reconciliation.
1
Scope
Which decisions matter
2
Document
Training and lineage
3
Test
Bias and accuracy
4
Explain
Can it be justified
5
Report
Findings and controls
How We Audit
01
Scoping the Decisions
We start from the decisions rather than the technology. A model producing a suggestion a person reviews carries different risk from one that approves a payment automatically, and the audit depth should follow that difference.
- Inventory of models influencing material decisions
- Degree of automation, from advisory to fully automated
- Consequence of an incorrect output per decision
- Regulatory exposure attaching to each decision type
02
Model Documentation and Lineage
We assess whether the organisation can account for how the model was built. Training data provenance, versioning and the record of what changed and when are frequently absent, which alone is a reportable weakness.
- Training data provenance and permitted use
- Model versioning and change history
- Documentation of features and their justification
- Records of validation performed before deployment
03
Bias and Performance Testing
We test performance across subgroups rather than only in aggregate. A model that is accurate overall while performing markedly worse for one group is a legal and reputational exposure, not merely a technical shortcoming.
- Accuracy measured across relevant subgroups
- Disparate impact testing where decisions affect people
- Performance drift since deployment
- Behaviour tested against edge and adversarial cases
04
Explainability and Human Override
Somebody affected by an automated decision is entitled to an explanation, and POPIA gives specific rights where a decision is made solely by automated means. We test whether an explanation can genuinely be produced and whether override actually works.
- Whether individual decisions can be explained
- POPIA rights around automated decision making
- Human override tested rather than assumed
- Escalation route for a contested decision
05
Control Testing and Reporting
Finally we test the controls around the model: who can change it, whether changes are approved, whether outputs are monitored and whether anyone would notice degradation.
- Access and change control over models
- Approval required before a model reaches production
- Ongoing monitoring of output quality
- Findings reported in audit language with root cause
What You Receive
- Inventory of models by decision type and automation level
- Assessment of model documentation and lineage
- Bias and performance testing results across subgroups
- Explainability and human override findings
- Control testing over model access, change and monitoring
- Findings reported with root cause and remediation actions
Indicative Timeline
An audit of a single model normally runs three to five weeks. Documentation quality is the largest variable: where nothing was recorded during development, reconstructing lineage takes considerably longer than testing does.
- Scoping and model inventory: one week
- Documentation and lineage review: one week
- Bias and performance testing: one to two weeks
- Control testing and reporting: one week
What We Test
Assurance covers the model, the data behind it, and the controls around both.
Training Data
Provenance, permitted use, representativeness and whether personal information was lawfully processed.
Model Accuracy
Performance against a held out set and against real outcomes since deployment.
Bias
Whether outcomes differ materially across groups in ways that cannot be justified.
Explainability
Whether a specific decision can be explained to the person it affected.
Human Oversight
Whether override exists, is used, and is not merely a rubber stamp.
Change Control
Who can alter a model, what approval is required, and whether changes are logged.
Frequently Asked Questions
Why would an accounting firm audit AI?
Because it is assurance work. Testing whether a control operates as management asserts, gathering evidence and reporting in audit language is exactly what an audit practice does. The novelty is the subject matter, not the discipline.
Can you audit a model we bought rather than built?
Yes, though vendor transparency limits depth. Where a supplier will not disclose training data or performance characteristics, that constraint is itself a finding and should inform how much reliance you place on the output.
What does POPIA say about automated decisions?
It gives data subjects rights where a decision affecting them is based solely on automated processing, including the ability to contest it and require human intervention. If you make such decisions, being able to explain and override them is a legal requirement rather than good practice.
How do you test for bias?
By measuring performance across relevant subgroups rather than in aggregate. A model that is ninety percent accurate overall may be seventy percent accurate for one group, and only disaggregated testing reveals that.
What if the model performs well but cannot be explained?
That is a finding, and its severity depends on the decision. For an internal efficiency tool it may be acceptable. For a decision affecting a person, an unexplainable model is a serious exposure regardless of accuracy.
Can you audit AI we deployed years ago?
Yes, and drift makes it worthwhile. Models degrade as the world moves away from their training data, and a model unmonitored for two years is usually performing worse than anyone realises.
Related Services
This sits inside our AI Services practice. Related work: AI Governance and Policy for the framework assurance tests against, and Application and Data Integrity Controls for automated processing controls more broadly.
