Escrever uma avaliação PREreview

LEGRA: A Pipeline for Building Graph-Based Representations of Polish Court Rulings for Legal Retrieval-Augmented Generation

de Szymon Dobrowolski e Waldemar Bauer

Publicado: 24 de novembro de 2025
Servidor: Preprints.org
DOI: 10.20944/preprints202511.1742.v1

Efficient access to similar legal cases is a crucial requirement for lawyers, judges, and researchers. Traditional text-based search systems often fail to capture both the semantic similarity and the relational context of legal documents \cite{article}. To address this challenge, we present LEGRA, a novel graph-based dataset of Polish court rulings designed for Retrieval-Augmented Generation (RAG) and legal research support \cite{https://doi.org/10.48550/arxiv.2005.11401}. LEGRA is automatically constructed through an end-to-end pipeline: rulings are collected from public sources, converted and cleaned, chunked into passages, and enriched with TF-IDF vectors and embedding representations. The data is stored in a Neo4j graph database where documents, chunks, embeddings, judges, courts, and cited laws are modeled as nodes connected through explicit relations. This structure enables hybrid retrieval that combines semantic similarity with structural queries, allowing legal professionals to quickly identify not only textually related cases but also those linked through judges, locations, or legal references. We discuss the construction pipeline, the graph schema, and potential applications for legal practitioners. LEGRA demonstrates how graph-based datasets can open new directions for AI-powered legal research.

Você pode escrever uma avaliação PREreview de LEGRA: A Pipeline for Building Graph-Based Representations of Polish Court Rulings for Legal Retrieval-Augmented Generation. Uma avaliação PREreview é uma avaliação de um preprint e pode variar de algumas frases a um parecer extenso, semelhante a um parecer de revisão por pares realizado por periódicos.

Antes de começar

Vamos pedir que você faça login com seu ORCID iD. Se você não tiver um iD, pode criar um.