Brodies’ AI pilot: what law firms should measure
Brodies chose Legora after a three-month, 180-person pilot. Our proposed scorecard helps law firms test quality, review time and readiness before buying.
Brodies announced a partnership with Legora on 25 September after a three-month pilot involving 180 colleagues across legal and business support functions. The firm says its evaluation considered performance, security, governance and practical use. Brodies’ announcement describes positive feedback, including ease of use, but does not publish a comparative dataset of time savings, error rates or financial returns.
That distinction matters to a firm planning its own purchase. An announcement can establish that an organisation tested a product and decided to proceed. It cannot tell another buyer what result to expect on different work, with different staff and controls.
Start with the decision the pilot must support
Our proposed approach begins with one sentence: “At the end of this pilot, we will decide whether to use this tool for these tasks, with these reviewers and these limits.” Make the tasks specific enough that a result can change the decision.
For example, a fictional pilot might examine the preparation of a first summary from a defined set of documents. Agree in advance what an acceptable summary must contain, which omissions would make it unusable and who is qualified to judge it. Use material approved for the exercise.
Write down the existing process too. Without a comparison, a convincing demonstration may only show that the tool can perform the task, not that the new process is worth adopting.
Use four measures that survive scrutiny
The following scorecard is our recommendation, not Brodies’ disclosed methodology.
- Total task time. Include selecting and preparing the material, writing instructions, generating output, reviewing it and correcting it. Report training and administration separately so they do not disappear from the buying decision.
- Quality after review. Classify errors by consequence. A missing qualification that changes the meaning should not count the same as a formatting correction. Keep the reviewer’s reasons alongside the score.
- Completion without rescue. Record whether the intended user can finish the task using normal support. Distinguish a successful independent attempt from one completed by a specialist standing beside them.
- Retrievable evidence. Ask a second authorised reviewer to identify the source documents, generated version, corrections and person who approved the final output. Record any missing link.
A hypothetical task that takes ten minutes to generate and forty minutes to repair has taken at least fifty minutes. The comparison belongs against the full time and quality of the existing process. Generation speed alone would answer the wrong question.
Keep difficult cases in the result
Include incomplete inputs, conflicting documents and routine exceptions in the agreed sample. Keep the result for each kind of task visible. An average can conceal a useful tool for one activity and a poor fit for another.
Where practical, have reviewers score outputs without being told which process produced them. Record the limits of the comparison, including differences in experience or task difficulty. A small internal pilot can guide a purchase without pretending to be a scientific trial.
Make the exit decision explicit: proceed for the tested tasks, narrow the use, repeat a defined part of the exercise, or stop. Avoid an open-ended pilot that produces encouraging anecdotes but never resolves the purchase.
Test the handoff into client work
Before expanding use, follow one approved output into the place where a colleague or client receives it. Can the recipient distinguish a draft from an authorised answer? Can a later reviewer establish which source version supported it? Those are formal communication questions about authority and the record of an exchange.
The practical lesson we draw from the announcement is to demand decision-ready evidence from a pilot. For a smaller firm, the appropriate scale may be much smaller than Brodies’ reported programme. The useful output is a justified decision with clear limits, not a participation target.
For the supplier side of the same decision, see our walkthrough for testing data control using one fictional client file.
This record analyses Brodies’ own announcement, accessed on 29 September 2026. We have not inspected its pilot data or independently evaluated Legora. The scorecard and hypothetical time example are original editorial proposals, not measured results or statements about Brodies’ internal method.