How can a legal AI tool provide useful support while ensuring that the information behind its answers is reliable, relevant and handled responsibly? This is one of the key questions addressed by EuroLegalBot’s D2.1 – Methodology for Dataset Creation.
The deliverable establishes the methodological foundations for building the dataset used to train EuroLegalBot. Drawing on existing research and best practices in the legal domain, it defines procedures for sourcing, selecting, filtering and organising legal information, with particular attention to data relevance and accuracy.
At the same time, the methodology places legal and ethical safeguards at the centre of dataset development. Data protection, GDPR compliance, anonymisation, fairness and the potential risks of bias in legal datasets are considered as essential elements of the process.
D2.1 therefore provides the framework on which the subsequent creation of the EuroLegalBot knowledge base and the development of a trustworthy AI-supported tool for European judicial cooperation can be built.