Your legal team chose legitimate interest as the lawful basis for processing training data. Your DPO documented a balancing test. Your AI team scraped publicly available profiles to fine-tune a recommendation engine. Then a supervisory authority sends a preliminary finding: your legitimate interest assessment doesn't hold up, and you never had valid consent. Now you're defending a processing operation built on a legal basis that was never defensible.
This pattern repeats across organizations deploying AI systems under the GDPR. The regulation offers six lawful bases for processing personal information under Article 6(1), but teams consistently misapply them when the use case involves AI training. The mistakes aren't about ignorance, they're about misunderstanding how AI processing maps to the GDPR's framework and what supervisory authorities will accept when they audit your AI operations.
Why These Mistakes Keep Happening
AI training creates a disconnect between how your team thinks about data use and how the GDPR structures lawful processing. Your data scientists see training data as an input to model development. The GDPR sees it as personal information processing that requires a lawful basis before you collect the first record. That gap produces predictable errors: teams pick a basis that sounds plausible without testing whether it can withstand regulatory scrutiny for this specific processing purpose.
The Garante Per La Protezione Dei Dati Personali has signaled that when you're using publicly sourced data to train AI, consent or legitimate interest are the only plausible lawful bases. That guidance narrows your options and raises the stakes for getting the choice right.
Mistake 1: Treating "Publicly Available" as a Legal Basis
Why it happens: Your team scraped data from social media profiles, public directories, or open forums. The information was accessible to anyone. It feels like you're not taking anything private, so surely you don't need consent.
The consequence: Public availability is not a lawful basis under Article 6(1). You still need to identify which of the six bases applies. Supervisory authorities have rejected the argument that scraping public data exempts you from the consent or legitimate interest analysis. If you can't articulate a valid basis, the entire training dataset becomes unlawfully processed personal information.
The fix: Document which basis you're relying on before you collect the data. If you choose legitimate interest, draft a balancing test that explains why your AI development purpose doesn't override the rights of the individuals whose data you're processing. If you choose consent, design a mechanism to obtain it before you scrape, or acknowledge that retrospective consent for already-collected public data is not feasible and reconsider your approach.
Mistake 2: Assuming Legitimate Interest Always Wins the Balancing Test
Why it happens: Your AI project serves a business purpose, improving customer recommendations, detecting fraud, personalizing content. That purpose feels legitimate, so you document it in a legitimate interest assessment and move forward.
The consequence: Article 6(1)(f) requires that your legitimate interest not be "overridden" by the interests or fundamental rights and freedoms of the data subject. Supervisory authorities scrutinize whether individuals could reasonably expect their publicly posted information to be used for AI training, whether the processing creates new risks (like profiling or automated decision-making), and whether you've offered meaningful safeguards. A business benefit alone doesn't win the balancing test.
The fix: Your legitimate interest assessment must address the data subject's perspective. Can someone who posted their information on a forum reasonably expect it to train your AI model? Does your processing create risks that weren't present in the original context? Document the safeguards you've implemented, anonymization where feasible, limitations on how the model will be deployed, transparency about the training process. If the balancing test reveals that individual rights likely override your interest, don't force the basis to fit. Reconsider the processing or switch to consent.
Mistake 3: Collecting Consent After You've Already Processed the Data
Why it happens: You've built a training dataset from publicly available sources. Later, you realize you need a stronger legal basis and decide to obtain consent. You send notifications to individuals asking them to approve the use of their data for AI training.
The consequence: Consent under the GDPR must be obtained prior to processing. Retrospective consent doesn't cure unlawful processing that already occurred. If you've already trained a model on personal information without a valid basis, asking for consent now doesn't make the initial processing lawful, it just creates a new processing operation going forward.
The fix: If you're relying on consent, design your data collection to obtain it before you process the information. For publicly sourced data, this often means you cannot use consent as a practical basis, you'd need to contact individuals before scraping their data, which defeats the efficiency of automated collection. In that case, legitimate interest may be your only viable option, and you must ensure your balancing test is defensible.
Mistake 4: Confusing "Necessary to Perform a Contract" with Business Operations
Why it happens: Your AI system improves a service you provide to users under a contract. Training the model feels necessary to deliver that service, so you cite Article 6(1)(b) as your basis.
The consequence: "Necessary to perform a contract" applies only when the processing is objectively required to fulfill your contractual obligations to the individual. Training an AI model to improve recommendations or personalize content is not necessary to perform the core contract, it's a value-add or business optimization. Supervisory authorities have consistently rejected expansive interpretations of contractual necessity that encompass any processing that benefits the service.
The fix: Reserve Article 6(1)(b) for processing that's genuinely required to deliver what you promised in the contract. If your AI training enhances the service but isn't essential to its basic function, use legitimate interest or consent instead. Be precise about what the contract actually requires versus what you'd like to offer.
Mistake 5: Failing to Document the Legal Basis Before Processing Begins
Why it happens: Your AI team starts collecting training data because the project timeline is tight. The legal basis analysis gets deferred to the compliance review, which happens after the model is already in development.
The consequence: Article 6(1) requires that you identify your lawful basis before you begin processing. If you're audited and cannot show that you documented a valid basis at the time of collection, the processing is unlawful, even if you could have justified it under legitimate interest or consent. The absence of documentation is itself a compliance failure.
The fix: Make legal basis identification a gate in your AI project workflow. Before your data science team ingests the first record, your DPO or legal team must document which Article 6(1) basis applies, why it applies, and what safeguards you're implementing. For legitimate interest, complete the balancing test. For consent, design the collection mechanism. Treat this as a hard prerequisite, not a post-hoc rationalization.
Prevention Checklist
- Identify the lawful basis under Article 6(1) before collecting any training data
- If using legitimate interest, complete a balancing test that addresses data subject expectations and processing risks
- If using consent, design a mechanism to obtain it prior to processing (not retrospectively)
- Confirm that "necessary to perform a contract" applies only to objectively required processing, not business enhancements
- Document your legal basis analysis before processing begins, not during a compliance review
- Review whether publicly sourced data realistically supports consent as a basis, or whether legitimate interest is your only viable option
- Implement safeguards (anonymization, purpose limitation, transparency) that support your chosen basis
- Train your AI and data science teams to flag legal basis questions before starting data collection




