AI-Assisted Judges in Pakistan Closed 6.3% More Cases, Study Finds
A randomized trial of 1,559 Pakistani judges found targeted training, not mere access to a GPT-4 assistant, drove the resolution gains.
What happened
A team of economists built JudgeGPT, a GPT-4-based chatbot that retrieves from 128,292 Pakistani judicial opinions and 943 statutes, to help trial judges with legal research and drafting. They then randomized 1,559 trial judges across 118 courts — roughly half the country’s trial judges and about 80% of its district courts — into three groups: JudgeGPT plus targeted training, JudgeGPT plus a generic technology seminar, and a no-AI control.
The targeted training consisted of six 90-minute sessions built with Pakistan’s Federal Judicial Academy. Judges who received it logged in and used JudgeGPT roughly four times more than judges given only the generic seminar: after 40 weeks, trained judges averaged 56 to 60 logins and more than 200 prompts, against 10 to 20 logins and under 50 prompts for the comparison group.
Districts with trained judges resolved about 1,848 additional cases a year, a 6.3% increase over baseline, a figure independently reported by the Pakistani outlet DAWN.COM rather than simply relayed from one source. On a quality check, blind pairwise comparisons rated rulings from trained judges as better in 59% of comparisons against 42% for control-group rulings; appeal rates fell slightly, and the paper reports no increase in gender- or religion-related bias in judicial language, according to the-decoder.com. The researchers estimate a return of about $38.50 in judicial cost savings per dollar spent running the tool, with a conservative floor of “at least $10 per dollar.”
“We do find an increase in cases resolved, and we don’t find any corresponding decrease in decision quality,” said study author Sultan Mehmood.
The paper, “Courts of Tomorrow,” is authored by Mehmood (New Economic School, Moscow), Christoph Goessmann, and Elliott Ash (ETH Zurich), and circulates as an NBER working paper and CEPR Discussion Paper No. 21783, dated July 14, 2026. It has not been reported as peer-reviewed or published in a journal.
What this means (and what it does not)
The gap between the two AI-access groups is the central finding: giving judges the tool was not enough on its own to change how much they used it, and the resolution gains are attributed to the trained group specifically. Only about a quarter of participating judges had used a large language model before the experiment, in a system with fewer than 2 judges per 100,000 residents — against 22 in the EU and 30 in England and Wales — and 2.26 million cases pending at the end of 2024, 82% of them in trial courts, per the-decoder.com.
The result does not show that AI can substitute for judicial judgment: the researchers themselves caution that the findings “don’t support replacing judges with AI” and stress that the gains depend on structured training, not tool access alone. It also does not establish that these results would hold in another judiciary, or that they were produced under conditions free of interest: the study’s own authors have a professional stake in a headline-grade positive result ahead of peer review; Pakistan’s judiciary and Federal Judicial Academy, which co-designed the training, have a political interest in showing their AI modernization program working against the backlog; and OpenAI, whose GPT-4 model underlies JudgeGPT, benefits from a favorable case study for generative AI generally, though the sources reviewed do not list OpenAI as a study author or funder. As MIT economist David Autor, who was not involved in the study, put it: “It’s not easy to do large-scale field experiments in civil service.”
What we still do not know
The paper has not passed peer review, so the 6.3% resolution figure and the quality findings are not yet independently confirmed. The measured window runs to roughly 40 weeks, and no source reviewed reports whether the effect holds, grows, or fades over a longer horizon. Nothing in the sources breaks the gain down by case type — criminal versus civil, simple versus complex — or by judge experience and seniority. The quality assessment relied on two Pakistani lawyers and a GPT-5-mini comparison rating pairwise rulings, but how those raters were selected, whether they were blinded to treatment status under a documented protocol, and who chose the case pairs are not described. No source discloses the study’s funding source. The authors reportedly suggest their GPT-4-based results may be a floor rather than a ceiling given newer models, but this is their own claim about an untested scenario, not a measured result. And whether Pakistan’s court structure, backlog severity, or training design would transfer to other legal systems is untested and unclaimed by the authors themselves.
Sources
Every source cited in this article, gathered in one place.
- https://elliottash.com/papers/Mehmood-Goessmann-Ash-Courts-of-Tomorrow-Evidence-Nationwide-Rollout-Generative-AI.pdf — elliottash.com
- https://the-decoder.com/an-ai-system-helped-pakistani-judges-clear-massive-backlogs-at-38-50-return-per-dollar-invested/ — the-decoder.com
- https://spectrum.ieee.org/judgegpt-experiment — spectrum.ieee.org
- https://www.dawn.com/news/2016213 — dawn.com
- https://conference.nber.org/conf_papers/f247308.pdf — conference.nber.org
Also available in Portugues (BR)