The Problem
You've crafted the perfect prompt for analysing complex M&A agreements. You feed it a 200-page SPA with schedules. The AI output is… wrong. But you don't know where it went wrong, or why, because you tested the most complex version first.
This is like debugging a rocket ship by launching it and hoping it works. In software engineering, they use "tracer bullets"—simple versions that validate the core logic before adding complexity.
Legal AI needs the same discipline.
What Are Tracer Bullets?
A tracer bullet is the simplest possible version of your task that can still demonstrate success or failure.
Instead of: "Analyse this 200-page share purchase agreement and identify all material adverse change provisions, assess enforceability, compare to market standard, and flag risks"
Start with: "Identify the material adverse change clause in this 2-page simplified agreement"
If the AI can't handle the simple version, it definitely can't handle the complex version. But if it can handle the simple version, you now have a baseline to build from.
Why This Matters for Legal AI
1. Isolate Failure Points
When a complex prompt fails, you're dealing with multiple variables:
- Token limit issues
- Ambiguous instructions
- Missing context
- Edge cases in the document
- Hallucination tendencies
- Format confusion
With tracer bullets, you test one variable at a time.
2. Build Confidence Incrementally
You know the AI can handle simple cases because you tested them. That gives you a baseline when you come to check a complex output.
3. Create Reference Points
Your tracer bullet becomes your regression test. Every time you modify the prompt, run it against the simple case first. If it breaks the tracer bullet, don't test the complex version.
The Method
Step 1: Identify the Core Task
Strip away all modifiers, qualifiers, and edge cases. What's the absolute simplest version?
Complex task: "Review this shareholders agreement for anti-dilution protections, assess whether they're weighted-average or full-ratchet, calculate economic impact under three scenarios, and compare to market terms"
Core task: "Find the anti-dilution clause"
Step 2: Create a Minimal Test Document
Don't use real contracts initially. Create a toy example with:
- Only the relevant clauses
- Clear, unambiguous language
- Obvious right answers
SHAREHOLDERS AGREEMENT (Simplified)
Clause 5.3 Anti-Dilution Protection
In the event of a down round, the conversion price shall be
adjusted using the weighted-average method.
This is your tracer bullet target—deliberately simple, deliberately obvious.
Step 3: Test and Validate
Run your prompt against the minimal document. The AI should easily succeed.
If it fails on this simple version:
- Your prompt has fundamental issues
- The AI doesn't understand the core task
- Your instructions are ambiguous
Fix these before adding complexity.
If it succeeds:
- Document what success looks like
- Save this as your regression test
- Proceed to Step 4
Step 4: Add Complexity Incrementally
Now add one complexity at a time:
Complexity Layer 1: Add ambiguity
Clause 5.3 Price Adjustments
Upon certain equity events, pricing mechanisms may be adjusted
as determined by the Board in accordance with customary methods.
Does the AI still identify this as an anti-dilution clause? If yes, proceed. If no, refine prompt.
Complexity Layer 2: Add length and noise
- Include multiple clauses
- Add irrelevant clauses
- Increase document to 10 pages
Complexity Layer 3: Add edge cases
- Multiple anti-dilution clauses
- Conflicting provisions
- Cross-references to schedules
Complexity Layer 4: Use real documents
Start with a precedent or a published agreement. Before you use a client's document, check that you may. What you may put into an AI tool depends on the plan your firm has, its agreement with the vendor and your firm's policy: Claude Projects for Client Matters sets out the checks for one tool.
Each layer tests a different failure mode. By isolating them, you know exactly what breaks and can fix it systematically.
Worked Example: Contract Review Prompt
This example is illustrative. It shows how the method runs, not the result of a published test.
Without Tracer Bullets
Attempt 1: "Review this 150-page facility agreement and summarise all representations and warranties"
Result: AI produces plausible-sounding summary, but misses critical representations in clause 12 and hallucinates one that doesn't exist.
Problem: You don't know if the issue is:
- The prompt structure
- Document length/token limits
- Specific language in clause 12
- Something else entirely
Outcome: Frustrated debugging with no clear path forward.
With Tracer Bullets
Tracer 1: "Review this 1-page agreement with two representations (solvency, due incorporation) and list them"
Result: Perfect. AI correctly identifies both.
Tracer 2: "Review this 5-page agreement with eight representations (including one buried in a sub-clause) and list them"
Result: AI misses the buried one.
Diagnosis: The prompt doesn't handle nested sub-clauses well.
Fix: Add explicit instruction: "Review all clauses and sub-clauses, including those nested within other provisions."
Tracer 2 Retest: Now finds all eight.
Tracer 3: "Review this 25-page agreement with realistic complexity"
Result: Success.
Tracer 4: Original 150-page document
Result: Success, with confidence that the prompt logic is sound.
Practice Area Applications
Litigation: Disclosure Review
Tracer 1: Single email, obvious relevance determination
Tracer 2: Email chain with partial relevance
Tracer 3: 50 emails with varying relevance
Tracer 4: 10,000 email corpus
Corporate: Due Diligence
Tracer 1: Single contract, identify key terms
Tracer 2: Three contracts, compare key terms
Tracer 3: 20 contracts with inconsistencies
Tracer 4: Full data room
Regulatory: Compliance Review
Tracer 1: Policy document, check single regulation
Tracer 2: Policy document, check five regulations
Tracer 3: Multi-document policy suite
Tracer 4: Enterprise-wide compliance review
Advanced Technique: Negative Tracer Bullets
Test what the AI shouldn't do:
Positive Tracer: Document contains indemnity clause → AI finds it ✓
Negative Tracer: Document doesn't contain indemnity clause → AI says "no indemnity clause found" ✓
This catches over-eager AI that hallucinates findings to be helpful.
Common Mistakes
Mistake 1: Skipping Straight to Complexity
"But my real use case is complex, so I need to test the complex version."
Reality: The complex version will fail in unclear ways. You'll waste time debugging symptoms instead of root causes.
Mistake 2: Making Tracer Bullets Too Simple
A tracer bullet should test the core task, not a different task entirely.
Too simple: "Does this document contain the word 'indemnity'?"
Right level: "Identify the indemnity clause and state who indemnifies whom"
Mistake 3: Not Documenting Successful Tracers
Your tracer bullets become your regression test suite. Save them.
Create a folder:
/prompt-tests
/contract-review
tracer-1-simple.pdf
tracer-2-nested-clauses.pdf
tracer-3-full-length.pdf
Every time you modify the prompt, run all tracers first.
Integration with Testing Checklist
Tracer bullets and the Testing Checklist work together:
- Tracer Bullets validate prompt logic during development
- Testing Checklist validates outputs before client delivery
Think of it as:
- Tracer Bullets = Unit tests for prompts
- Testing Checklist = Quality assurance for outputs
Measuring Success
You know tracer bullets are working when:
- You can locate a failure: When a prompt fails, you can say which layer of complexity broke it
- You can explain a failure: A failure on a complex document is one you can trace to a layer you have tested
- Clearer requirements: The process of creating tracers clarifies what you actually need
- Team alignment: Other lawyers can understand prompt behaviour by reading tracer examples
Getting Started
- Take your most complex prompt
- Ask: "What's the simplest version of this task?"
- Create a 1-paragraph test document with obvious right answer
- Test it
- If it fails, fix the prompt before adding complexity
- If it succeeds, add one layer of complexity and repeat
The discipline of starting simple forces clearer thinking about what you're actually asking the AI to do—and that clarity pays dividends throughout the entire workflow.
Building a library of tracer bullet tests for your firm? The consulting page describes what is offered.