The practices of major AI manufacturers when collecting training data are becoming increasingly questionable. Heise online has documented in a recent analysis that the industry is apparently shifting boundaries deliberately – and in doing so, is not shying away from legal gray zones.
The essentials
- Systematic escalation: AI makers employ increasingly aggressive data collection methods without regard for copyright or data protection
- Heise online documents this as a trend, not isolated incidents
- Compliance gaps: Many of these practices operate in legal gray areas that existing laws do not adequately regulate
- Business model pressure: Competition for larger and better training datasets drives the escalation forward
Data competition without rules
The pressure to train ever larger and better models is driving AI providers to resort to increasingly questionable means. While early generations of language models relied on licensed or publicly available data, current systems now also use sources whose legal permissibility is disputed. Heise online observes a continuous shifting of boundaries – what was considered unacceptable yesterday is treated as standard today.
Copyright and data protection under strain
Particularly problematic is the handling of copyrighted works. Authors, photographers, and artists see their work processed in training data without consent – a conflict that has already led to multiple lawsuits. At the same time, data protection problems emerge: personal data from millions of people ends up in AI systems without those affected being informed or asked for consent.
Existing regulation lags behind this development. The EU AI Act does address transparency requirements for training data, but concrete enforcement mechanisms are still missing.
What this means for decision-makers
For German companies deploying or developing AI systems, the question of data provenance becomes a critical compliance issue. Anyone using an AI model whose training data is based on questionable sources potentially carries liability risks – for instance, if copyrights were violated or data protection breaches occurred. This applies especially in regulated sectors such as financial services or healthcare.
The recommendation: ask questions. What data did the AI provider use? Are licenses in place? How was data protection handled? Anyone unable or unwilling to answer these questions signals elevated risk. The era in which AI systems could be accepted as black boxes is drawing to a close.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




