Guide

How to monitor AI systems for drift

Every AI assessment describes the system on the day it was signed. Drift is the distance that opens afterwards. What to watch, how to set thresholds a person will actually act on and why vendor model changes are the drift source most organisations miss.

How do you monitor an AI system for drift?

Treat the assessment as a description of the system on the day it was signed, and monitor for the ways the system departs from that description. Watch the inputs, because data drifting from what the system was assessed on is the earliest signal. Watch the outputs, meaning quality, error patterns and behaviour against the standards the assessment assumed. Watch usage, because a system used for purposes or populations it was not assessed for has drifted even if the model has not changed. And watch the vendor, because for organisations that consume AI rather than build it, the model changing underneath you is the largest drift source and the least visible. Set thresholds that name a person and an action, since a drift signal nobody owns is a dashboard, and route what monitoring finds back into the governance record so the assessment gets revisited rather than quietly outdated.

An AI risk assessment is a photograph. On the day it was signed it described the system, and every day afterwards the system and the world move while the photograph does not. Drift is the name for that widening distance, and the reason it deserves its own practice is that the expensive AI failures are rarely systems nobody assessed. They are systems that kept running after their behaviour left the assessed envelope, with the time between departure and detection deciding the size of the damage.

Four places drift comes from

The inputs. The data flowing into a system moves away from what the system was assessed and, where relevant, trained on. Customer profiles shift, upstream systems change formats, the world the data describes changes. Input drift is usually the earliest signal, arriving before output quality visibly degrades.

The outputs. Quality, error patterns and behaviour move against the standards the assessment assumed. Output drift is the kind people mean when they say a system “got worse”, and by the time humans notice it unaided, it has usually been happening for a while.

The usage. The system itself may be unchanged while the organisation’s use of it drifts towards new purposes, new populations and new stakes. A system assessed for internal drafting that quietly becomes the engine of customer-facing decisions has drifted as surely as if the model had changed, and no amount of model monitoring will catch it.

The vendor. For organisations that consume AI rather than build it, which is most Australian organisations, the largest drift source is the model changing underneath you. Vendors update models, adjust defaults and switch on features, sometimes announced, often not. The system you assessed is then simply not the system you run, and nothing in your own telemetry announces the moment it happened.

Thresholds are a governance decision

Raw drift metrics are easy to produce and easy to ignore. The work that makes monitoring real is deciding, per system, what movement matters and who does what when it moves. The form that works is a sentence with an owner in it, along the lines of “if this signal passes this point, this person reviews, retests or escalates”. Setting the threshold is a governance decision in itself, because it encodes how far from its assessed behaviour your organisation is willing to let a system travel before anyone looks, and for higher-risk systems that tolerance should be tighter and written down.

Scale the effort to the risk. Monitoring everything equally is a way of monitoring nothing well. The systems whose drift can hurt people or the organisation earn continuous signals and tight thresholds, while a low-stakes internal tool may earn a usage check at review time, and the risk assessment is exactly the instrument that tells you which is which.

Close the loop into the record

A drift signal that leads nowhere is trivia. The point of noticing that a system has left its assessed envelope is that the governance responds, so the assessment is revisited, a control is adjusted, a guardrail is tightened or the finding matures into an incident and is handled as one. That only happens reliably if what monitoring finds lands in the governance record, against the system it concerns, in front of the person who owns it. Monitoring that lives in an engineering dashboard while governance lives in documents produces two accounts of the same system that stop matching, which is drift of a second and more embarrassing kind.

Where Aicura fits

Aicura’s monitoring lands the signals where the governance already lives. Drift and quality signals, dynamic risk scoring, thresholds and alerting, vendor model change signals and telemetry ingestion all attach to the same register entries the assessments hang off, so a moving picture is visible to the owner accountable for the system, and what monitoring observes becomes part of the record alongside the assessments and the enforcement events. Monitoring is part of the Enterprise tier.

The indicators are operational signal, not judgement. Aicura provides guidance, not legal advice, and does not certify, audit or issue a compliance verdict, and what a moving indicator means for a system stays with the owner the alert reaches.

Common questions

Is drift only a problem for models we train ourselves? No, and believing so is how consuming organisations get caught. A vendor updating the model inside a product you use is drift from your side of the relationship, because the system you assessed is no longer the system you run. Organisations that only consume AI arguably need change monitoring more, since they neither control nor necessarily hear about the changes.

How is drift different from an incident? Drift is the system moving away from its assessed behaviour. An incident is harm, or a near miss. Left unwatched, the first matures into the second, and the time between them is where the damage compounds. Monitoring exists to act in that window.

What thresholds should we set? Ones that name a person and an action. The form that works is “if this signal moves past this point, this owner does this”, even if the action is only a review. Thresholds without owners produce dashboards that everyone can see and nobody answers for.

Does monitoring replace periodic review? No, they answer different questions. Periodic review asks whether the assessment still describes the system on a schedule. Monitoring asks it continuously and tells you when the answer turns to no between reviews. A system with both gets reviewed when it matters rather than when the calendar says.

A note on this page

This guide describes the practice of monitoring AI systems for drift in governance terms. The right signals and thresholds depend on your systems and your risk, and the statistical detail of drift detection is its own field. It is general information, not advice.


Related guides

Monitoring in Aicura

Aicura's monitoring lands drift and quality signals, dynamic risk scoring, thresholds and vendor model change signals against the same register entries the assessments hang off, so when the picture moves the accountable owner hears about it. Monitoring is part of the Enterprise tier.