Move from evidence collection to observed performance

A supplier evidence pack tells a buyer what controls are designed to do. An acceptance test checks what actually happens in the product, configuration and user journey the buyer intends to use.

Article 50 distinguishes between provider and deployer obligations. Depending on the system, these can include informing people that they are interacting with AI, machine-readable marking of generated or manipulated content, notices for emotion recognition or biometric categorisation, and disclosure of deepfakes or certain public-interest text.

The test below is a procurement and governance control, not a universal legal classification. Confirm the organisation's role, system scope, deployment context and applicable exceptions before setting pass criteria.

Gate 1: freeze the scope and responsibility boundary

Record the product, model, version, configuration, integration, channels, languages, user groups and content types being tested. Identify which features involve direct AI interaction or generate or manipulate text, audio, images or video.

Set out the provider, deployer and upstream responsibilities for every control. Include who renders a notice, embeds a mark, applies a visible label, retains logs, responds to a failure and approves a material change.

Fail the gate if the tested version cannot be identified or if a required control falls into an undocumented gap between buyer and supplier.

Gate 2: test direct-interaction notices

Run the first interaction as a new user and verify that any required AI notice is clear and distinguishable from the start. Repeat the journey for returning sessions, embedded widgets, mobile layouts, voice channels and supported accessibility modes.

Capture the notice, timestamp, environment and expected result. The Commission says the obviousness exception should be interpreted restrictively, so a supplier relying on it should provide a documented assessment for the actual audience and context.

Fail the gate if a user must open help content, inspect terms or complete a meaningful part of the interaction before learning that the interface is AI-driven.

Gate 3: test generated-content marking and detectability

Generate representative text, image, audio and video outputs for every supported route. Check whether in-scope outputs carry the supplier's documented machine-readable marking and whether an appropriate detector can identify it after ordinary delivery and storage.

Repeat agreed transformations such as resizing, compression, export, format conversion or platform upload. Record where the mark persists, degrades or disappears, together with known technical limits and any relevant exception.

Fail the gate when a claimed control is absent from an in-scope output, cannot be tested with the documented method or differs materially from the supplied specification without an approved explanation.

Gate 4: test visible disclosure at first exposure

For deepfake content, test that a natural person receives a clear and distinguishable disclosure by first exposure at the latest. Do not treat an invisible machine-readable mark as a substitute for the deployer's visible or audible disclosure duty.

For AI-generated or manipulated public-interest text, test the publishing workflow, label rule and any claimed human-review or editorial-control route. The Commission describes substantive review by suitably knowledgeable people; spelling or grammar checks alone are not enough.

Fail the gate if the label disappears in a supported channel, depends on specialist tools to perceive or can be bypassed through an ordinary publishing workflow.

Gate 5: prove failures can be contained and corrected

Deliberately exercise an agreed negative case, such as a missing localisation, unsupported output path or disclosure-rendering failure. Verify that monitoring, user reporting or quality checks can detect the problem and route it to a named owner.

Ask the supplier to show the incident record, containment option, correction process and evidence used to confirm remediation. Agree when the buyer must be notified and whether the feature can be disabled or rolled back.

Fail the gate if neither party can identify the affected users and outputs, preserve evidence or stop continued exposure while the issue is investigated.

Gate 6: make acceptance conditional and repeatable

Record each test case with an identifier, applicable obligation or control, precondition, steps, expected result, observed result, evidence link, status, owner and retest date. Approve exceptions only with a reason, risk owner, compensating control and expiry date.

Connect acceptance to supplier change control. A model, user interface, content format, integration, localisation or detection-method change should trigger impact assessment and proportionate retesting before production use.

For the documentary inputs, start with our transparency evidence-pack guide. AI Act Ready helps procurement teams combine supplier evidence with observed acceptance results that remain useful after go-live.