Audit of AI and Automated Decisions

When a model influences a payment, an assessment, a credit decision or an eligibility outcome, somebody has to be able to explain why it produced that answer. This is assurance work, and it extends naturally from what an audit practice already does: testing that a control operates as management claims. The subject happens to be a model rather than a reconciliation.

1

Scope

Which decisions matter

2

Document

Training and lineage

3

Test

Bias and accuracy

4

Explain

Can it be justified

5

Report

Findings and controls

How We Audit

01

Scoping the Decisions

We start from the decisions rather than the technology. A model producing a suggestion a person reviews carries different risk from one that approves a payment automatically, and the audit depth should follow that difference.

02

Model Documentation and Lineage

We assess whether the organisation can account for how the model was built. Training data provenance, versioning and the record of what changed and when are frequently absent, which alone is a reportable weakness.

03

Bias and Performance Testing

We test performance across subgroups rather than only in aggregate. A model that is accurate overall while performing markedly worse for one group is a legal and reputational exposure, not merely a technical shortcoming.

04

Explainability and Human Override

Somebody affected by an automated decision is entitled to an explanation, and POPIA gives specific rights where a decision is made solely by automated means. We test whether an explanation can genuinely be produced and whether override actually works.

05

Control Testing and Reporting

Finally we test the controls around the model: who can change it, whether changes are approved, whether outputs are monitored and whether anyone would notice degradation.

What You Receive

Indicative Timeline

An audit of a single model normally runs three to five weeks. Documentation quality is the largest variable: where nothing was recorded during development, reconstructing lineage takes considerably longer than testing does.

What We Test

Assurance covers the model, the data behind it, and the controls around both.

Training Data

Provenance, permitted use, representativeness and whether personal information was lawfully processed.

Model Accuracy

Performance against a held out set and against real outcomes since deployment.

Bias

Whether outcomes differ materially across groups in ways that cannot be justified.

Explainability

Whether a specific decision can be explained to the person it affected.

Human Oversight

Whether override exists, is used, and is not merely a rubber stamp.

Change Control

Who can alter a model, what approval is required, and whether changes are logged.

Frequently Asked Questions

Because it is assurance work. Testing whether a control operates as management asserts, gathering evidence and reporting in audit language is exactly what an audit practice does. The novelty is the subject matter, not the discipline.

Yes, though vendor transparency limits depth. Where a supplier will not disclose training data or performance characteristics, that constraint is itself a finding and should inform how much reliance you place on the output.

It gives data subjects rights where a decision affecting them is based solely on automated processing, including the ability to contest it and require human intervention. If you make such decisions, being able to explain and override them is a legal requirement rather than good practice.

By measuring performance across relevant subgroups rather than in aggregate. A model that is ninety percent accurate overall may be seventy percent accurate for one group, and only disaggregated testing reveals that.

That is a finding, and its severity depends on the decision. For an internal efficiency tool it may be acceptable. For a decision affecting a person, an unexplainable model is a serious exposure regardless of accuracy.

Yes, and drift makes it worthwhile. Models degrade as the world moves away from their training data, and a model unmonitored for two years is usually performing worse than anyone realises.

Related Services

This sits inside our AI Services practice. Related work: AI Governance and Policy for the framework assurance tests against, and Application and Data Integrity Controls for automated processing controls more broadly.

Discuss AI assurance