The Problem
You use one AI tool for everything because it is familiar. You do not know whether another tool your firm has approved would suit a particular task, because you have never compared them on your own work.
The Solution
Compare the tools on the tasks you do, using material where you already know the right answer, and keep a record.
This guide does not say which tool is best at anything. This site has not published a comparison, so it has no result to show you. No model version is named here either, because vendors replace them. Test the tools your firm has approved, and test them again when a vendor changes the model.
Before You Start
- Compare only the tools your firm has approved. Ask which tools are approved, and on which plan, before you open one.
- Keep client material out of the comparison. Use a precedent, a published agreement or a published judgment. You do not need a client's document to find out how a tool behaves.
- Check before you use any tool with client material. What you may put into a tool depends on the plan your firm has, its agreement with the vendor and your firm's policy. Claude Projects for Client Matters sets out the checks for one tool.
What to Compare
Each heading below is a kind of task. Under it are the questions to answer for each tool, and a prompt to test with.
Legal Research
- Does the tool link to its sources?
- Does each link open the judgment or the legislation itself?
- Does the source say what the answer says it does?
- Does every case it cites exist?
Test prompt:
Recent England and Wales cases on directors' duties in takeover
situations. Give the neutral citation for each and link to the judgment.
Run it on a point of law where you already know the leading authorities.
In R (Ayinde) v London Borough of Haringey and Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin), the Divisional Court said that freely available generative AI tools "are not capable of conducting reliable legal research" (paragraph 6). Whichever tool does best in your comparison, you still check its research against authoritative sources before you use it: see The Citation Verification Rule.
Long Document Review
- Can the tool take the whole document? Each vendor publishes its own limits. Those for Claude are in Claude Context Window Optimisation.
- Does it find a provision you know is near the end of the document?
- Does it give clause numbers, and are they right?
- Does it find a conflict between clauses that you know is there? Does it report one that is not there?
Test prompt:
[Upload a precedent or a published agreement]
Identify any internal contradictions or inconsistent provisions.
Give the clause number for each.
First Drafts
- How much do you have to change before the draft is usable?
- Does it follow your instructions on length, tone and format?
- Does it apply the law of England and Wales?
Test prompt:
Draft a confidentiality clause for a software services agreement
governed by the law of England and Wales.
Client Communication
- Is the tone right for the reader?
- Does it state anything as certain that is not?
- Would you send it after editing, or would you rewrite it?
Test prompt:
Draft an email to a client explaining that a deal has fallen through.
Empathetic but not apologetic. Professional.
Use a placeholder for every name.
Case and Document Summaries
- Does it get the facts, the issue and the decision right?
- Does it leave out something that matters?
- Does it add something that is not in the document?
Test prompt:
[Upload a published judgment you have read]
Summarise: Facts, Issue, Decision, Reasoning (5 bullets)
Tools Built for Lawyers
Your firm may have a tool built for legal work. Test it in the same way as the others. What a vendor says about its own product is not a test result: ask the vendor what evidence it has published, and run your own comparison.
How to Run a Comparison
- Choose one task you do often
- Choose the material: no client information, and an answer you already know
- Write one prompt and use the same wording in each tool
- Decide what a good answer is before you look at any output
- Run the prompt in each tool and save each output
- Mark each output against what you decided in step 4
- Record the result, with the date and the model name as the tool shows it
- Run it again when a vendor changes the model
One run on one document is a single observation. Repeat the comparison with different material before you draw a conclusion. Tracer Bullets explains how to start with a simple case and add complexity.
Keeping the Record
Task: [e.g. summarise a judgment]
Material: [e.g. published judgment, with its neutral citation]
Prompt: [the exact wording]
Date: [date]
Tool and model: [as shown in the tool]
Good answer is: [what you decided before you ran it]
Result: [what was right, what was wrong, what was missing]
Time to correct: [how long it took you to make the output usable]
Pro Tips
Know your firm's approved tools:
- Your firm may restrict which tools you can use, and on which plan
- Check before you use any tool with client information
- Ask who approves a new tool before you try one
Use as few tools on a matter as the work needs:
- Each tool you use on a matter is another vendor that receives information about it
- Moving a draft from one tool to another sends the same material to both
Use projects and memory with care:
- Claude Projects for Client Matters covers projects
- ChatGPT Memory Management covers what ChatGPT carries from one conversation to the next
Common Mistakes
❌ Choosing a tool on reputation: what other people say about a tool is not a test on your work
❌ Comparing with different prompts: you cannot tell whether the tool or the prompt made the difference
❌ Using a client's document for the comparison: a precedent or a published document does the job
❌ Relying on an old result: the tool may have changed since you recorded it
✅ Same prompt, same material, a known answer and a dated record
Quick Reference
To compare tools:
- Approved tools only, and no client material
- One task, one prompt, material with a known answer
- Decide what a good answer is before you run the prompt
- Run it, mark it and record it, with the date and the model
- Repeat when the model changes
Whichever tool you choose: check its output before you rely on it.