Citation Recognition in Legal Documents
Named Entity Recognition with Transformer-Based Deep Learning Methods

- Autor:innen
- Mirio EggmannJasmin FitzDario Glasl
- Station
- Bachelor of Science in Informatik, OSTNote 6.0
- Publikation
- Paper (DOI)
Abstract
Online platforms are popular among legal professionals for daily searches of information needed for drafting opinions, preparing court cases etc. These platforms must handle large volumes of documents ensuring timely availability of information. A key challenge is identifying and linking citations in the documents to their sources within the limited time available for nightly re-indexing. Traditional methods, such as regular expressions, are fast but lack flexibility and do not support context-dependent citations. The latest state-of-the-art generative language models lack efficiency for the task, processing only 36 documents per hour. We propose transformer-based encoder-only models to recognize citations. We demonstrate that the discriminative BERT-based models can process 3,000 documents in about 32 minutes, exceeding the performance requirements and ensuring scalability and efficiency of large-scale legal platforms.
Task Description
This bachelor thesis develops a system for recognizing citations in legal documents. It is carried out in collaboration with Swisslex, an online research platform for legal documents. In order to link those documents in their web application, citations and their specific parts need to be identified in the text. The current citation recognition system is based on regular expression matching. This approach is not robust and is hard to maintain.
Approach
The task is divided into two stages: recognizing citations, as well as their parts. First, text chunks are processed using a fine-tuned Google BERT multilingual base model (uncased). This identifies citations and classifies them into the categories CASELAW, LAW and LITERATURE. The recognized citations are then passed to a second fine-tuned model, based on DistilBERT (uncased). This identifies the specific parts of a citation, such as COURT, DATE, ARTICLE, etc. Sparse data poses a challenge for recognizing parts of a citation. Few-shot prompting with an LLM yields good results, but experiments show that it is prohibitively slow in practice. Computational requirements of such large models hampers their adoption. Therefore, DistilBERT is fine-tuned via the supervision of Llama 3.3 70B through knowledge distillation. The 1'044 times smaller DistilBERT achieves similar performance compared to Llama. The final .NET based solution offers bulk processing as well as a web interface for user interaction. FastAPI is used to serve the fine-tuned models. Results are stored in an MS SQL database.
Results
The fine-tuned Google BERT achieves an F1 test score of 95.7% for recognizing citations. The proposed solution offers many benefits apart from the at par performance with the regex-based system. The solution is more robust and can handle deviations such as typing errors and new citation formats more easily. Moreover, the cherry on top is the improved performance in recognizing parts of a citation with more granular labels. The fine-tuned DistilBERT recognizes parts with an F1 test score of 98.9%. Further improvement is possible via increasing diversity in the training samples.
Downloads








