What Do the European Data Protection Board’s Web Scraping Guidelines Mean for AI Training Datasets?

On July 7, 2026, the European Data Protection Board published draft guidelines on web scraping for generative AI, aimed at giving practical guidance on how EU data protection law applies when personal information is pulled from the open internet to train models. The guidelines cover both organizations that do the scraping themselves, directly or through a contractor, and organizations that license or reuse datasets someone else already scraped. For any company buying a generative AI system trained on internet-sourced data, the guidelines, as drafted, do not shift the data protection responsibility away from the buyer. The EDPB’s public consultation on the draft remains open until 30 October 2026, so the text can still change before it is finalized.

Related Content