Legal Systems and Artificial Intelligence (CBR project)


Project team

  • Project leader: Simon Deakin (CBR), Mihoko Sumida (Hitotsubashi University, Tokyo)
  • Co-Investigators: Jennifer Cobbe, Jon Crowcroft, Jat Singh (Computer Laboratory, University of Cambridge); Felix Steffek (Faculty of Law, University of Cambridge); Christopher Markou, Linda Shuku, Helena Xie (CBR); Yuishi Washida, Kazuhiko Yamamoto, Keisuke Takeshita, Mikiharu Noma, Wataru Uehara (Hitotsubashi University); Nanami Furue (Tokyo University of Science); Motoyuki Matsunaga (Institute for International Socio-Economic Studies, Tokyo); Takashi Araki, Chikako Kanki (Tokyo University)
  • Researchers: Bhumika Billa, Vanessa Cheok, Anca Cojocaru, Narine Lalafaryan, Chris Pang, Joana Ribeiro De Faria, Holli Sargeant, Lucy Thomas

Project status

Completed

Project dates

2020-2023

Funding

ESRC and Japanese Science and Technology Agency (2020-2023); Keynes Fund (2024)


Overview

Background

The aim of this project is to assess the implications of the introduction of Artificial Intelligence (AI) into legal systems in Japan and the United Kingdom. The main part of the project was jointly funded by the UK’s Economic and Social Research Council, part of UKRI, and the Japanese Society and Technology Agency (JST), and involves collaboration between the University of Cambridge (the CBR, Computer Laboratory and Faculty of Law) and Hitotsubashi University, Tokyo (the Graduate Schools of Law and Business Administration).

The main project was completed in December 2023, but outputs will continue to appear, and members of the 2 project teams (Cambridge and Tokyo) are continuing their collaboration. In December 2023 Simon Deakin and Linda Shuku received funding from the Keynes Fund to support their ongoing research using machine learning and natural language processing to analyse historical poor law and workmen’s compensation cases. The use of machine learning (ML) and natural language processing (NLP) to replicate aspects of legal decision making is well advanced. A number of ‘Legal Tech’ applications have been developed by law firms and commercial suppliers and are being used, among other things, to model litigation risk. Data analytics are informing decisions on legally consequential matters including probation, predictive policing and credit evaluation.

The next step will be to use ML to replicate core functions of legal systems, including adjudication. At the same time there are already signs of push-back against the use of ML in the legal sphere. Critics point to the biases in current algorithmic decision-making processes which systematically disadvantage the poor and minority groups. Concerns over the constitutionality of automating judicial processes prompted, for example, the passage Art. 33 of French Law 2019-222, which bars the use of personally identifiable data of judges and other court officials with a view to “evaluating, analysing, comparing or predicting their professional performance, real or supposed”.

Aims and objectives

In this context there is an urgent need for informed debate over the uses of AI in the legal sphere. The project will advance this debate by exploring stakeholders’ perceptions of the acceptability of AI-related technologies in the legal domain identifying and addressing legal and ethical risks associated with algorithmic decision making understanding the potential of, and limits to, the computational techniques underlying law related AI.

Methods

The project has been organised through 3 work packages which will deploy, respectively, the methods of Horizon Scanning (WP1), and machine learning, deep learning, natural language processing, and computational linguistics (WPs 2 and 3).

WP1: Constructing Future Scenarios for the Uses of AI in Law: A Horizon-Scanning Approach

Project leaders: Washida, Sumida, Deakin

The Horizon Scanning Method was developed principally by the Stanford Research Institute in the late 1960s. The method avoids the assumption that the future will tend to deviate from a linear extension of current circumstancesm, and attempts instead to develop more realistic predictions of the future by focusing on the collection and analysis of information that does not lie on the path of this linear extension. In implementing the Horizon Scanning approach we will firstly produce a database containing a range of information sources on the uses of AI in law, drawn from press reports and commentary and secondary academic literatures. The database will be used as the basis for discussion at a series of workshops. We will invite experts, researchers, corporate professionals and users across a broad range of fields of activity and different age ranges to take part in the workshops. Emergent scenarios will describe different possible combinations of advantages and risks stemming from the use of AI.

WP2: Computation of Complex Knowledge Systems: Law and Accounting

Project leaders: Deakin, Markou, Crowcroft, Singh, Cobbe, Shuku, Noma

This WP considered whether the juridical reasoning underpinning employment status decisions could be statistically represented using historical data from decided cases. We used machine learning (ML) and natural language processing (NLP) to analyse legal decisions for latent or hidden variables that can help inform and refine the model. We then explored how far the same techniques could be applied to the digitisation of knowledge systems used in accounting.

WP3: Predicting the Outcome of Dispute Resolution: Feasibility, Factors and Ethical Implications

Project leaders: Steffek, Xie, Yamamoto

This WP dealt with the prediction of dispute outcomes and aimed to advance understanding of the use of artificial intelligence in case outcome predictions. Analysis was carried out on a large dataset of 42 English court cases. The dataset was used to test different ML approaches to predicting dispute outcomes. A parallel study using Japanese court data was also undertaken. In addition, this WP developed ethical guidelines for regulating Artificial Intelligence in dispute resolution. The development of the guidelines was supported by through input of stakeholders in the UK Ministry of Justice, the OECD Department on Access to Justice, leading representatives of the UK judiciary, and LawTech firms.

Progress

The project began in January 2020 and a planning meeting and workshop was held in Cambridge in early March, with the participation of the Japanese team. Shortly afterwards lockdowns were initiated in both Cambridge and Tokyo and work on the project was formally paused for a 3-month period. Research was resumed in the summer of 2020.

In WP1, the collection of abstracts for use in the Horizon Scanning Method began in August 2020. A horizon scanning workshop, originally planned to take place in Cambridge in December 2020, was postponed because of COVID-19. The workshop was rescheduled to take place once COVID-related restrictions on travel had come to an end, and was successfully completed in March 2023, with the joint participation of the Cambridge and Hitotsubashi teams from WP1. The workshop took as its theme the impact of artificial intelligence (‘AI’) on the future of work. Just over 30 experts ranging from human resources (HR) professionals and lawyers to trade unionists and academics took part. The participants were divided into 3 break-out groups, each of which brainstormed future scenarios based on a dataset of summaries of around 100 media opinions disseminated constructed by the Cambridge project team before the workshop. The dataset included news, blog posts, and op-eds published in English across the world in the last 4 years and were curated by google search for recent writing on AI and work. Initial results from the deliberations were published in the form of a blog written by Bhumika Billa and Simon Deakin, and further analysis of the workshop findings, applying the horizon scanning methodology, was undertaken by the Japanese team in the course of 2023 and 2024.

In WP2 progress was made in developing the conceptual framework for the work, and has resulted in a series of publications including an edited collection, Is Law Computable? Critical Reflections on Law and Artificial Intelligence, which was published by Hart/Bloomsbury in November 2020, and papers published in the Journal of Cross-Disciplinary Research in Computational Law and the Northern Ireland Legal Quarterly. In addition, substantial progress has been made on constructing a dataset of historical employment cases which is being used to test hypotheses concerning the long-run dynamics of legal change and the coevolution of law with social and economic development. Simon Deakin, Linda Shuku and Vanessa Cheok completed a dataset of poor law cases in the spring of 2023, and a paper based on analysis of the data will appear in the Journal of Law and Society in 2024. Funding from the Keynes Fund is supporting this work on a continuing basis.

In WP3, work was carried out on the dataset of English cases and the possibility of creating similar datasets of Japanese cases was explored with relevant stakeholders. Progress was also made in developing the ML and NLP methods to be used in analysis of the judicial data. A dataset of English court cases, the Cambridge Law Corpus, was completed by the summer of 2023, and published shortly afterwards.

Both WP2 and WP3 organised multiple meetings between the British and Japanese sides, via Zoom, to coordinate progress and ensure continuing cooperation notwithstanding the impossibility of meeting in person during the COVID-19 emergency, before in-person meetings were resumed in Cambridge in March 2023. The final conference of the project, at which several working papers were presented, took place in Tokyo in December 2023.

Findings

The horizon scanning workshop held in Cambridge in March 2023 discussed scenarios in 3 key areas of potential impact of AI in the workplace: HR performance assessments; protections for freelance workers in the gig economy; and the resolution of workplace disputes. There was general agreement on the risks arising from AI, which included over-reliance on biased and error-prone systems, which had to be weighed in the balance against potential cost improvements. Increases in freelance work driven by AI would empower employers at the cost of worker autonomy if new rights were not enacted, and made effective, with respect to surveillance and the maintenance of a living wage. Automation of dispute resolution in and beyond the workplace, advanced as a means of improving access to justice, could also lead to a loss of worker voice through information asymmetries and an absence of representation. Our work on the use of ML for case prediction suggests that while there is huge potential for this form of law-related AI, the techniques involved are still at an early stage. So far, published studies have been confined to demonstrating correlations between different parts of the same judgment text. Since judges write their opinions knowing the outcome, there is a high risk of cross-contamination between the test and training data used in these studies. Our experience of building large scale corpora of legal cases suggests that some of the challenges involved in curating datasets and annotating cases for analysis can be overcome using large language models such as GPT-4, which can be used to facilitate automated annotation. Our historical research suggests that NLP techniques such as sentiment analysis can be used to identify trends in judicial decision making and to address the issue of how shifts in legal language are linked to wider changes in the political, economic and technological context of the law.

GDPR notice

The Cambridge Law Corpus (CLC) is a dataset of more than 250,000 court cases for legal AI research. The CLC has been developed as part of the Legal Systems and Artificial Intelligence research project funded by UKRI (UK Research and Innovation) and JST (Japan Science and Technology Agency). Most cases in the CLC are from the 21st century, but the corpus includes cases as old as the 16th century. The CLC only contains decisions of UK courts and tribunals (together referred to as courts in this notice) that have been made available by the relevant courts for publication. All decisions in the CLC have already been published before either by the courts themselves or by other information providers.

In the UK, court cases are not anonymised. Parties, judges, barristers and other persons involved in court proceedings should, therefore, expect to be named in judgments because courts uphold the principle of open justice, promote the rule of law and ensure public confidence in the legal system. However, courts will anonymise a party if the non-disclosure is necessary to secure the proper administration of justice and to protect the interests of that party. The CLC contains the texts of decisions as they were made available by the courts themselves. As a result, the CLC contains the names of persons involved in court decisions and other personal data as reported by the official court decision.

The CLC is the basis of research experiments conducted as part of this project. They include the identification of the sentences in judgments that contain the outcome of the case, the determination of topics that the courts deal with and the prediction of case outcomes.

Ethical approval has been granted by the Research Ethics Committee of the Centre of Business Research at the University of Cambridge.

GDPR requirements: research exemptions

Compliance with the Data Protection Act 2018 (DPA) and UK General Data Protection Regulation (GDPR) is the basis of the legality of the CLC and its use for research. The personal data in this corpus was not collected directly from data subjects and was only undertaken for research purposes in the public interest. Both these circumstances offer exemptions from obligations in the GDPR.

Given the practically impossible and disproportionate task of informing all individuals mentioned in this corpus and that these cases are publicly available and being processed exclusively for research purposes, the CLC is exempt from notification requirements. Further, research in the public interest is privileged as regards secondary processing and processing of sensitive categories of data restrictions. In particular, this aids the protection of the integrity of research datasets.

Safeguards

We apply safeguards in compliance with legal and ethical requirements and ensure that:

  • Appropriate technical and organisational safeguards exist to protect personal data.
  • Processing will not result in measures being taken in respect of individuals and no automated decision-making takes place.
  • There is no likelihood of substantial damage or distress to individuals from the processing.
  • Users who access the corpus must agree to comply with the DPA and the GDPR in addition to any local jurisdiction.
  • Any individual may request the removal of a case or certain information which will be immediately removed.
  • The corpus will not pose any risks to people’s rights, freedoms or legitimate interests.
  • Access to the corpus will be restricted to researchers based at a university or other research institution whose Faculty Dean (or equivalent authority) confirms, inter alia, that ethical clearance for the research envisaged has been received.
  • Researchers using the corpus must agree to not undertake research that identifies natural persons, legal persons or similar entities. They must also guarantee they will remove any data if requested to do so.

Further information

Further information on the legality and ethical administration of the CLC can be found in the related research paper “The Cambridge Law Corpus: a dataset for legal AI research”.

Further information on the terms and conditions that researchers applying for access to the CLC have to comply with are available on the Department of Computer Science and Technology’s CLC project page. The project page also contains information and a link for those applying for the removal of a case from the CLC.

Output

Journal articles

Billa, B. (2023) “Law as code: exploring information, communication and power in legal systems.” Journal of Cross-Disciplinary Research in Computational Law, 2(1)

Cobbe, J., Veale, M. and Singh, J. (2023) “Understanding accountability in algorithmic supply chains.” In: ACM Conference on Fairness, Accountability, and Transparency (DOI: 10.1145/3593013.3594073)

Deakin, S. and Markou, C. (2022) “Evolutionary interpretation: law and machine learning.” Journal of Cross-Disciplinary Research in Computational Law, 1(2)

Deakin, S. and Markou, C. (2021) “Evolutionary law and economics: theory and method.” Northern Ireland Legal Quarterly, 72(4): 682-712 (DOI: 10.53386/nilq.v72i4.939)

Deakin, S. and Shuku, L. (2025) “Exploring computational approaches to law: the evolution of judicial language in the Anglo-Welsh poor law, 1691-1834.” Journal of Law and Society, 52(1): 3-33 (DOI: 10.1111/jols.12521)

Datasets

Deakin, S., Shuku, L. and Cheok, V. (2024) English Poor Law cases, 1690-1815. [data collection]. UK Data Service. SN: 856924, DOI: 10.5255/UKDA-SN-856924.

Östling, A. et al (2024) The Cambridge law corpus: a corpus of court decisions for legal and AI research.

Other publications

Billa, B. (2024) “Folúkẹ́ Adébísí, Decolonisation and legal knowledge: reflections on power and possibility [Book review].” Modern Law Review, 87: 1603-1606 (DOI: 10.1111/1468-2230.12895)

Top