# EU AI Act Article 10: Data and data governance

Source: https://aiexponent.com/eu-ai-act/article-10 · Content verified 2026-10-04

Article 10 of the EU AI Act requires high-risk AI providers to govern training, validation, and testing data: quality criteria, examination for bias and protected attributes, and representativeness for the intended purpose. Datasets must be appropriate to the geographical, contextual, behavioural or functional setting in which the system is intended to be used (Art. 10(4)).

- Status: Applies 2 Dec 2027
- Who: Providers of high-risk AI systems that train, validate or test on data.
- From when: 2 Dec 2027 (Annex III), 2 Aug 2028 (Annex I) (Art. 113(c)(i), as amended)
- Maximum fine: €15M or 3% (Art. 99(4), point (a), through the provider obligations in Art. 16)

## What Article 10 says

> **10(1), as amended by Regulation (EU) 2026/1744** 1. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1) whenever such data sets are used.

> **10(2)** 2. Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system. Those practices shall concern in particular:

> **10(2)(f)** (f) examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations;

> **10(2)(g)** (g) appropriate measures to detect, prevent and mitigate possible biases identified according to point (f);

> **10(3)** 3. Training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used. Those characteristics of the data sets may be met at the level of individual data sets or at the level of a combination thereof.

> **10(4)** 4. Data sets shall take into account, to the extent required by the intended purpose, the characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting within which the high-risk AI system is intended to be used.

Selected paragraphs, quoted exactly from Regulation (EU) 2024/1689: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (checked 4 Oct 2026).

## What changed

Regulation (EU) 2026/1744, in force since 27 Jul 2026. Removed and context lines quote Regulation (EU) 2024/1689 as adopted; added lines quote the amending Regulation.

### Article 10(1), (5) and (6)

```diff
- 1. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2 to 5 whenever such data sets are used.
+ 1. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1) whenever such data sets are used.
- 5. To the extent that it is strictly necessary for the purpose of ensuring bias detection and correction in relation to the high-risk AI systems in accordance with paragraph (2), points (f) and (g) of this Article, the providers of such systems may exceptionally process special categories of personal data, subject to appropriate safeguards for the fundamental rights and freedoms of natural persons. In addition to the provisions set out in Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680, all the following conditions must be met in order for such processing to occur:
- (a) the bias detection and correction cannot be effectively fulfilled by processing other data, including synthetic or anonymised data;
- (b) the special categories of personal data are subject to technical limitations on the re-use of the personal data, and state-of-the-art security and privacy-preserving measures, including pseudonymisation;
- (c) the special categories of personal data are subject to measures to ensure that the personal data processed are secured, protected, subject to suitable safeguards, including strict controls and documentation of the access, to avoid misuse and ensure that only authorised persons have access to those personal data with appropriate confidentiality obligations;
- (d) the special categories of personal data are not to be transmitted, transferred or otherwise accessed by other parties;
- (e) the special categories of personal data are deleted once the bias has been corrected or the personal data has reached the end of its retention period, whichever comes first;
- (f) the records of processing activities pursuant to Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680 include the reasons why the processing of special categories of personal data was strictly necessary to detect and correct biases, and why that objective could not be achieved by processing other data.
- 6. For the development of high-risk AI systems not using techniques involving the training of AI models, paragraphs 2 to 5 apply only to the testing data sets.
+ 6. For the development of high-risk AI systems not using techniques involving the training of AI models, paragraphs 2, 3 and 4 of this Article and Article 4a(1) shall apply only to the testing data sets.
```

Point (6) of the same Article 1 inserts Article 4a, "Processing of special categories of personal data for bias detection and correction".

Source: Regulation (EU) 2026/1744, Article 1, point (9), https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng (verified 4 Oct 2026).

### Article 113, third paragraph, point (c)

```diff
  It shall apply from 2 August 2026.
- (c) Article 6(1) and the corresponding obligations in this Regulation shall apply from 2 August 2027.
+ (c) Chapter III, Sections 1, 2, and 3, with the exception of Article 6(5), shall apply from:
+ (i) 2 December 2027 as regards AI systems classified as high-risk pursuant to Article 6(2) and Annex III; and
+ (ii) 2 August 2028 as regards AI systems classified as high-risk pursuant to Article 6(1) and Annex I;
```

The context line is the second paragraph of Article 113, which set the date for Annex III systems before the change.

Source: Regulation (EU) 2026/1744, Article 1, point (40), https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng (verified 4 Oct 2026).

## What you must produce

Records of how your training, validation and testing data sets are governed.

- 10(2)(a) to (c): Data governance and management practices: design choices, data collection and origin, data preparation
- 10(2)(f) and (g): An examination for possible biases, and the measures to detect, prevent and mitigate them
- 10(2)(h): The data gaps or shortcomings found, and how they are addressed
- 10(3): Evidence that the data sets are relevant, sufficiently representative and, to the best extent possible, free of errors and complete

A summary to help you plan. The quoted text above is the law.

## The tool: License Compliance Checker

Coverage: Partial.

License Compliance Checker partially addresses Article 10. It surfaces training-data licence and provenance risk via the dataset risk registry (top-50 known datasets flagged with critical / high / medium tiers). Full Article 10 coverage (dataset lineage across runs, bias examination, statistical-property characterisation) is the scope of TraceForge, which is not scheduled and is built on request. LCC covers the licence and provenance part today.

```bash
pip install license-compliance-checker
```

Writes: OSS + model licence report (JSON).

Tool docs and tool FAQ: https://aiexponent.com/docs/license-compliance-checker.md

For coding agents:

Install license-compliance-checker from PyPI into a virtual environment (Python 3.11 or newer); the command is `lcc`. In the project root run `lcc scan .` for a licence summary, `lcc scan . --format json --output scan-report.json` for a machine-readable report, and `lcc sbom generate scan-report.json --format cyclonedx --output sbom.json` for a CycloneDX SBOM (version 2.0.2 or newer). Exit code 0 means no violations. Docs: https://aiexponent.com/docs/license-compliance-checker.md

## Questions about Article 10

### When does Article 10 apply?

It applies from 2 Dec 2027 for high-risk systems listed in Annex III (Art. 113(c)(i), as amended), and from 2 Aug 2028 for high-risk systems covered by Annex I (Art. 113(c)(ii), as amended). Before Regulation (EU) 2026/1744, the dates were 2 Aug 2026 and 2 Aug 2027.

### What is the maximum fine for breaching Article 10?

Up to €15 million or 3% of total worldwide annual turnover for the preceding financial year, whichever is higher (Art. 99(4), point (a), through the provider obligations in Art. 16). For SMEs, including start-ups, the fine is capped at whichever of the two is lower (Art. 99(6)). Since 27 Jul 2026, the same lower cap applies to small mid-cap enterprises (Art. 99(6a)).

### Did the Digital Omnibus change Article 10?

Yes. Regulation (EU) 2026/1744, in force since 27 Jul 2026, makes these changes. Article 10: paragraphs 1 and 6 replaced, paragraph 5 deleted. High-risk dates move to 2 Dec 2027 (Annex III) and 2 Aug 2028 (Annex I). The section "What changed" quotes the old and new text.

### Is there an AiExponent tool for Article 10?

Partly. License Compliance Checker covers part of Article 10. It writes a OSS + model licence report (JSON).

---

Not legal advice. Not a notified body. The tools produce evidence, not conformity assessment.
All docs as Markdown: https://aiexponent.com/llms.txt · Guide for coding agents: https://aiexponent.com/agents.md
