Abstract
Legal argumentation analysis represents a critical yet underexplored domain in computational linguistics, particularly for cross-lingual contexts where high-quality parallel corpora with argumentative annotations remain scarce. This paper introduces the Hong Kong Legal Judgments Argumentative Corpus (HKLJ-Arg), a novel bilingual Chinese-English parallel corpus comprising 557 professionally translated paragraph pairs extracted from 15 legal judgments spanning 2012-2025. Our corpus uniquely leverages Hong Kong's bilingual legal framework to provide high-fidelity parallel translations with comprehensive argumentative structure annotations, including claims, premises, examples, and their logical relationships. The annotation methodology employs a prompt-engineered approach using large language models, validated through expert evaluation achieving Cohen's kappa of 0.79 for component identification and 0.72 for relationship classification. We further present the Argumentation Preservation Assessment Framework (APAF), a systematic methodology for evaluating how effectively machine translation systems preserve argumentative structures across linguistic boundaries. Statistical analysis reveals diverse argumentative patterns across constitutional, criminal, and civil law domains, with systematic preservation of logical structures across languages. This resource addresses a critical gap in legal natural language processing by providing the first large-scale bilingual corpus specifically designed for cross-lingual legal argumentation mining, translation quality assessment, and comparative legal reasoning analysis.
| Original language | English |
|---|---|
| Journal | DATA INTELLIGENCE |
| Publication status | Published - 2025 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 16 Peace, Justice and Strong Institutions
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver