01
What documenting an AI system means
Documenting an AI system means recording in a verifiable way what the system does, who provides it and who uses it, which data feeds it, and what information is communicated to the people involved. Regulation (EU) 2024/1689 defines an AI system as an automated system that, for explicit or implicit objectives, infers from input how to generate output such as predictions, content, recommendations or decisions (citation-2). It also defines provider and deployer as the parties that develop or use the system under their own authority (citation-2). Solid documentation distinguishes these roles, the data used, and the transparency controls required.
02
Data categories to track
The regulation distinguishes several categories of data that should be tracked separately: training data, used to adapt the system's parameters; validation data, used to evaluate the trained system and avoid under- or over-fitting; testing data, used to confirm expected performance before market placement; input data, meaning data supplied to or acquired by the system; and biometric data, obtained through technical processing of a person's physical, physiological or behavioural characteristics (citation-1). Tracking these categories separately makes it possible to reconstruct how the system was built and verified over time.
03
The documentation flow, step by step
Steps
- Record the system's definition: objectives, output produced and environments it may affect, according to the regulation's definition of an AI system (citation-2).
- Identify and document the provider role and the deployer role for each system, as defined by the regulation (citation-2).
- Classify the data used into training, validation, testing, input and, where present, biometric data (citation-1).
- Ensure staff involved in operating and using the system receive AI literacy appropriate to their tasks (citation-4).
- Prepare disclosures for people exposed to emotion recognition, biometric categorisation or artificially generated content, provided no later than the first interaction (citation-0, citation-3).
04
A hypothetical case
Hypothetical example: a European company uses a generative AI system to produce customer support texts. The compliance team records that the company is a deployer of the system, while the external provider marks outputs as artificially generated in a machine-readable format (citation-0). Texts published to inform the public carry a notice of artificial generation, unless subject to human editorial review with assigned editorial responsibility (citation-3). The documentation also lists the input data used by operators (citation-1).
05
Transparency controls to verify
- Verify that synthetic audio, image, video or text output is marked and detectable as artificially generated or manipulated (citation-0).
- Check that people exposed to emotion recognition or biometric categorisation are informed about how the system works (citation-0).
- Verify that deep fakes are disclosed to the people concerned, except where exceptions provided by law apply (citation-0, citation-3).
- Check that disclosures are provided clearly and distinguishably no later than the first interaction or exposure (citation-3).
- Verify that texts generated to inform the public state the artificial generation, unless subject to human editorial review (citation-3).
06
Operational choices for supporting software
A company can choose to configure, together with its software provider, a register that links each AI system to its provider, deployer, data categories and disclosures given to the people involved. This register can keep evidence of markings on synthetic outputs and of notices provided before the first interaction. Suggested operational choices include: assigning review roles to confirm disclosures and periodically checking recorded data categories, based on the regulation's definitions (citation-2, citation-1, citation-0).
07
Frequent clarifications on documentation
- Who is considered a provider and who a deployer? The provider develops or has developed the system and places it on the market; the deployer uses it under its own authority (citation-2).
- Which data categories should be distinguished in the documentation? Training, validation, testing and input data, and biometric data where applicable (citation-1).
- When must disclosures be given to exposed people? No later than the first interaction or exposure, in a clear and accessible way (citation-3).
08
The next useful step
The next useful step is to start an inventory of AI systems active in the company, assigning to each the role of provider or deployer, the data categories, and a person responsible for disclosures, before extending documentation to output marking controls (citation-2, citation-1, citation-0).
FAQ
Frequently asked questions
What distinguishes a provider from a deployer under the regulation?
The provider develops or has developed the system and places it on the market or puts it into service under its own name; the deployer uses it under its own authority.
Which types of data must be tracked separately in documentation?
Training data, validation data, testing data, input data and, where present, biometric data, each serving a distinct function in the system's lifecycle.
When must the artificial nature of content be communicated?
Information must be provided to the people concerned clearly and distinguishably no later than the first interaction or exposure, except where exceptions provided by law apply.
✓