Legal AI is desperate for a credible benchmark built by experienced lawyers on real cases.
More and more lawyers are finding that agents do not perform well enough on legal work, especially on complex matters. Part of why that is hard to fix is that we cannot say precisely where they fall short.
Meanwhile, partners and GCs are not short on products. They are short on anyone they can trust to tell them which one is better. And the most widely used open benchmark in legal AI today was built by a player in the market.
So here is what we are doing about it:
Legal Tech Frontier, our community, is collaborating with the Agents’ Last Exam team at the Center for Responsible, Decentralized Intelligence (RDI) at University of California, Berkeley on an independent, open-source benchmark for legal agents. The project is “Agents’ Last Exam for Legal”, part of ALE:
- RDI Page ➡️ https://rdi.berkeley.edu/blog/agents-last-exam/
- ALE Page ➡️ https://agents-last-exam.org/
ALE has covered 55 occupations and 1,500+ tasks, each one sourced from a real project a working professional actually completed.
When OpenAI announced GPT-5.6, ALE was the first benchmark it led with.
We are thrilled to have invited Wei Chen, founder of The Atticus Project and a leading expert in legal benchmarking, to join us.
Open call: Engineering / Research Lead, ALE for Legal
We are now looking for a full-time Engineering / Research Lead for ALE for Legal.
The lead will be co-first author on the legal domain benchmark paper, and named as “Domain Lead” in ALE’s general submission to Nature family.
Estimated time commitment: 30–40 hours per week
Qualifications & expectations
- Strong expertise in AI Agent evaluation, experimentation, and testing
- Deep passion for building domain-specific Legal AI benchmarks
- Master’s or Ph.D. in AI, Computer Science, or a closely related field (US/North America-based preferred)
Core responsibilities (Domain Lead)
- Design & Build: Lead the architecture of long-horizon tasks and evaluation environments tailored for the legal domain.
- Coordinate & Curate: Collaborate with domain experts and attorneys to design, review, and ingest high-quality benchmark tasks and workflows.
- Drive End-to-End Execution: Actively drive ongoing benchmark iterations, model evaluations, and co-authoring the benchmark paper.
Please note this is a volunteer research role / non-salaried academic collaboration, ideal for researchers, PhDs, or engineers looking to lead a high-impact open-source benchmark with top-tier academic visibility.
If that sounds like a fit, please send your resume and background to legal@agents-last-exam.org, and CC helenfan@legaltechfrontier.com, with the email subject line: “ALE for Legal - Engineering Lead (Legal Tech Frontier)”.
While we may not be able to reply to every email individually, please feel free to send a quick DM to Helen Fan on LinkedIn if your inquiry is time-sensitive.
We will also open a separate call for experienced lawyers as core contributors as the next step. We want a number of people who have spent years doing legal work and can tell in one look whether a task is real. So stay tuned!
Let’s do something big to help reshape the industry!