In recent years, large language models (LLMs) have taken on a central role in the development of generative artificial intelligence. These models, based on advanced neural architectures such as transformers, are capable of understanding, generating, and manipulating natural language with an unprecedented level of sophistication. Their versatility makes them powerful tools in a wide range of applications: from virtual assistance to automatic content generation, from language translation to the synthesis of complex documents.
LLMs operate on the basis of an intensive training process, which requires the processing of enormous amounts of textual data, often collected from publicly accessible sources on the web. It is precisely at this stage that some of the most sensitive and controversial issues regarding personal data protection arise. As highlighted in the document published by the Data Security Council of India (DSCI), the training of LLMs may involve the use of personal data that is not always collected in accordance with the fundamental principles of privacy regulations, such as the European Union’s GDPR.
In particular, the use of “publicly available” data does not automatically exclude the need to comply with information obligations, purpose limitations, minimization principles, and the rights of data subjects. Furthermore, the generalist and adaptable nature of LLMs makes it difficult to define the purposes of processing ex ante, hindering the application of traditional legal concepts such as data controller and data processor. Added to this are issues related to transparency, data retention, and the ability of users to effectively exercise their rights.
This bulletin therefore aims to analyze the strains between the technical functioning of LLMs and the legal principles of personal data protection, with the goal of highlighting critical areas and proposing ideas for a more harmonious regulatory and technical evolution.