Writing
What Grok 4.5's New EU Disclosure Adds to the Training-Data Record
xAI has published its first EU training-content summary, covering Grok 4.5. It places the model in the largest available data-size bands for text, images, audio and video, and reports using all six categories of training-data sources in the European Commission's template.
That makes Grok 4.5 newly comparable with the summaries already collected in my AI training-data transparency dataset. It also provides a useful test of the template itself: the source questions reveal meaningful differences among providers, while the size bands often group very different models together.
Bottom line: Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to publish a summary of their training content using the AI Office's standardized template. In the 11-entry dataset snapshot used for the comparisons below, six models report more than 10 trillion text tokens and five report between 1 billion and 10 trillion. None uses the smallest band. Because the top band has no upper bound, Grok 4.5, GPT-5.5, Gemini 3 Pro and a three-billion-parameter open model receive the same size classification. The filing tells us which kinds of sources xAI says it used, but not how Grok's corpus compares in scale with the other models in that band.
Broad size bands limit comparison
The size question is likely to interest copyright holders, researchers and regulators, but it provides limited differentiation. Providers select one of three bands, each covering a wide range.
| Text training-data size band | Models filing it |
|---|---|
| More than 10 trillion tokens, with no upper bound | 6 |
| 1 billion to 10 trillion tokens | 5 |
| Less than 1 billion tokens | 0 |
The top band contains Google's Gemini 3 Pro family, OpenAI's GPT-5.5, xAI's Grok 4.5 and Meta's Muse Spark. It also contains Swiss AI's Apertus-70B and Hugging Face's SmolLM3-3B, two open models, one with three billion parameters. Whatever separates their training corpora, the filed size answer is the same.
The middle band is also broad: “1 billion to 10 trillion tokens” spans four orders of magnitude. Microsoft's Phi-4 family, SpeakLeash's Bielik and Bria all fall within it. Bria provides the only precise figure in this snapshot, writing “up to 19.2 billion tokens” and “479 million images” into the size fields. That text adds useful precision beyond the standard band.
The same limitation applies to modality-specific comparisons. Gemini 3 Pro, GPT-5.5, Grok 4.5 and Muse Spark all report the maximum bands for text, images, audio and video. The summaries do not provide enough detail to rank their modality-specific datasets by scale.
Source categories provide more differentiation
The data-source questions give a clearer view of how providers say their models were built:
| Provider and model | Public | Licensed | Private third-party | Crawled | User data | Synthetic |
|---|---|---|---|---|---|---|
| Meta, Muse Spark | Yes | Yes | Yes | Yes | Yes | Yes |
| OpenAI, GPT-5.5 | Yes | Yes | Yes | Yes | Yes | Yes |
| xAI, Grok 4.5 | Yes | Yes | Yes | Yes | Yes | Yes |
| SpeakLeash, Bielik v3 11B | Yes | Yes | No | Yes | No | Yes |
| Hugging Face, SmolLM3-3B | Yes | No | No | No | No | Yes |
| Swiss AI, Apertus-70B | Yes | No | No | No | No | No |
| Bria, Bria 3.2 | No | Yes | No | No | No | No |
Meta, OpenAI and xAI answer Yes to every category: public data, licensed data, private third-party data, their own crawlers, user data and synthetic data. Several open and licensed models report narrower source mixes. Swiss AI's Apertus uses public datasets and says it did not operate its own crawler. Bria reports a different strategy again, saying that Bria 3.2 was trained exclusively on commercially licensed images.
There are limits to the comparison. Google's summary answers only three of the six source questions, so a blank means “not stated,” not “No.” Microsoft answers some nominally binary fields with prose rather than Yes or No. The template makes comparison possible, but providers do not always complete it in comparable ways.
xAI also states that it has not signed the GPAI Code of Practice and that it honors opt-out signals directly instead. It is the first provider in this dataset to publish the template without signing the Code. Signing the Code and publishing a training-content summary are separate actions with different timelines; neither should be used as a proxy for the other.
A second search found six more summaries
After adding Grok 4.5, I ran another provider-by-provider search. It identified six credible public summaries not yet represented in the 11-entry snapshot above:
- Adobe Firefly Image Model 5
- Domyn Large
- FastwebMIIA
- Thinking Machines' Inkling
- Microsoft MAI-Image-2
- Meta Muse Image
These documents are being validated and normalized before they enter the comparable dataset, so I have not mixed their values into the counts above. That separation matters: discovery is not the same as ingestion, especially when checkboxes, modality definitions and publication formats vary.
I could not find a standardized Article 53 summary for Moonshot AI's Kimi K2.5 on its official site, API documentation or official model page. Its technical report describes approximately 15 trillion mixed visual and text tokens, but that is not the standardized EU disclosure. Kimi K2.5 was released after the obligation took effect; whether the requirement applies also depends on whether Moonshot placed the model on the EU market, so the absence alone should not be labeled non-compliance.
Reaction so far
I found no public reaction from the European Commission, the AI Office or an identified independent expert to the Grok 4.5 filing specifically. That should not be interpreted as either approval or criticism.
The wider debate about the template provides context. The Commission calls it a “common minimal baseline”, designed to improve transparency while protecting trade secrets rather than provide a technically complete inventory. A coalition of 38 creator and rightsholder organizations has expressed dissatisfaction with the implementation package, arguing that it does not provide enough information for rightsholders to exercise their rights. The European Parliament has also called more generally for stronger transparency, remuneration and the ability to prevent protected content from being used.
Those positions concern the framework as a whole, not Grok 4.5. The new filing illustrates both sides: it creates a comparable record of source categories that did not previously exist, but its broad size bands neither reveal how much data xAI used relative to other providers nor identify individual copyrighted works.
Caveats. Every figure is self-reported and unaudited. The 11-entry comparison is a snapshot, not a survey of the market, and three entries are Microsoft Phi variants. Missing source rows mean the summary is silent, not that the answer is No. Data cutoffs range from March 2024 to June 2026. Models placed on the EU market before August 2, 2025 have until August 2, 2027 to publish, so a missing summary is not necessarily evidence of non-compliance.
Data: ai_training_metrics, Transparency Report API; initial snapshot of 11 model entries from nine providers, with filings published from June 2025 through July 2026. Follow-up discovery performed July 23, 2026.