Skip to content
Back to Methodology
prompt engineeringbeginner

Tracer Bullets: Test Simple Before Complex

Start with the simplest possible version of your legal task, validate it works, then progressively add complexity. Adapted from software engineering test-driven development.

beginner level

The Problem

You've crafted the perfect prompt for analysing complex M&A agreements. You feed it a 200-page SPA with schedules. The AI output is… wrong. But you don't know where it went wrong, or why, because you tested the most complex version first.

This is like debugging a rocket ship by launching it and hoping it works. In software engineering, they use "tracer bullets"—simple versions that validate the core logic before adding complexity.

Legal AI needs the same discipline.

What Are Tracer Bullets?

A tracer bullet is the simplest possible version of your task that can still demonstrate success or failure.

Instead of: "Analyse this 200-page share purchase agreement and identify all material adverse change provisions, assess enforceability, compare to market standard, and flag risks"

Start with: "Identify the material adverse change clause in this 2-page simplified agreement"

If the AI can't handle the simple version, it definitely can't handle the complex version. But if it can handle the simple version, you now have a baseline to build from.

Why This Matters for Legal AI

1. Isolate Failure Points

When a complex prompt fails, you're dealing with multiple variables:

  • Token limit issues
  • Ambiguous instructions
  • Missing context
  • Edge cases in the document
  • Hallucination tendencies
  • Format confusion

With tracer bullets, you test one variable at a time.

2. Build Confidence Incrementally

You know the AI can handle simple cases because you tested them. That gives you a baseline when you come to check a complex output.

3. Create Reference Points

Your tracer bullet becomes your regression test. Every time you modify the prompt, run it against the simple case first. If it breaks the tracer bullet, don't test the complex version.

The Method

Step 1: Identify the Core Task

Strip away all modifiers, qualifiers, and edge cases. What's the absolute simplest version?

Complex task: "Review this shareholders agreement for anti-dilution protections, assess whether they're weighted-average or full-ratchet, calculate economic impact under three scenarios, and compare to market terms"

Core task: "Find the anti-dilution clause"

Step 2: Create a Minimal Test Document

Don't use real contracts initially. Create a toy example with:

  • Only the relevant clauses
  • Clear, unambiguous language
  • Obvious right answers
SHAREHOLDERS AGREEMENT (Simplified)

Clause 5.3 Anti-Dilution Protection

In the event of a down round, the conversion price shall be
adjusted using the weighted-average method.

This is your tracer bullet target—deliberately simple, deliberately obvious.

Step 3: Test and Validate

Run your prompt against the minimal document. The AI should easily succeed.

If it fails on this simple version:

  • Your prompt has fundamental issues
  • The AI doesn't understand the core task
  • Your instructions are ambiguous

Fix these before adding complexity.

If it succeeds:

  • Document what success looks like
  • Save this as your regression test
  • Proceed to Step 4

Step 4: Add Complexity Incrementally

Now add one complexity at a time:

Complexity Layer 1: Add ambiguity

Clause 5.3 Price Adjustments

Upon certain equity events, pricing mechanisms may be adjusted
as determined by the Board in accordance with customary methods.

Does the AI still identify this as an anti-dilution clause? If yes, proceed. If no, refine prompt.

Complexity Layer 2: Add length and noise

  • Include multiple clauses
  • Add irrelevant clauses
  • Increase document to 10 pages

Complexity Layer 3: Add edge cases

  • Multiple anti-dilution clauses
  • Conflicting provisions
  • Cross-references to schedules

Complexity Layer 4: Use real documents

Start with a precedent or a published agreement. Before you use a client's document, check that you may. What you may put into an AI tool depends on the plan your firm has, its agreement with the vendor and your firm's policy: Claude Projects for Client Matters sets out the checks for one tool.

Each layer tests a different failure mode. By isolating them, you know exactly what breaks and can fix it systematically.

Worked Example: Contract Review Prompt

This example is illustrative. It shows how the method runs, not the result of a published test.

Without Tracer Bullets

Attempt 1: "Review this 150-page facility agreement and summarise all representations and warranties"

Result: AI produces plausible-sounding summary, but misses critical representations in clause 12 and hallucinates one that doesn't exist.

Problem: You don't know if the issue is:

  • The prompt structure
  • Document length/token limits
  • Specific language in clause 12
  • Something else entirely

Outcome: Frustrated debugging with no clear path forward.

With Tracer Bullets

Tracer 1: "Review this 1-page agreement with two representations (solvency, due incorporation) and list them"

Result: Perfect. AI correctly identifies both.

Tracer 2: "Review this 5-page agreement with eight representations (including one buried in a sub-clause) and list them"

Result: AI misses the buried one.

Diagnosis: The prompt doesn't handle nested sub-clauses well.

Fix: Add explicit instruction: "Review all clauses and sub-clauses, including those nested within other provisions."

Tracer 2 Retest: Now finds all eight.

Tracer 3: "Review this 25-page agreement with realistic complexity"

Result: Success.

Tracer 4: Original 150-page document

Result: Success, with confidence that the prompt logic is sound.

Practice Area Applications

Litigation: Disclosure Review

Tracer 1: Single email, obvious relevance determination

Tracer 2: Email chain with partial relevance

Tracer 3: 50 emails with varying relevance

Tracer 4: 10,000 email corpus

Corporate: Due Diligence

Tracer 1: Single contract, identify key terms

Tracer 2: Three contracts, compare key terms

Tracer 3: 20 contracts with inconsistencies

Tracer 4: Full data room

Regulatory: Compliance Review

Tracer 1: Policy document, check single regulation

Tracer 2: Policy document, check five regulations

Tracer 3: Multi-document policy suite

Tracer 4: Enterprise-wide compliance review

Advanced Technique: Negative Tracer Bullets

Test what the AI shouldn't do:

Positive Tracer: Document contains indemnity clause → AI finds it ✓

Negative Tracer: Document doesn't contain indemnity clause → AI says "no indemnity clause found" ✓

This catches over-eager AI that hallucinates findings to be helpful.

Common Mistakes

Mistake 1: Skipping Straight to Complexity

"But my real use case is complex, so I need to test the complex version."

Reality: The complex version will fail in unclear ways. You'll waste time debugging symptoms instead of root causes.

Mistake 2: Making Tracer Bullets Too Simple

A tracer bullet should test the core task, not a different task entirely.

Too simple: "Does this document contain the word 'indemnity'?"

Right level: "Identify the indemnity clause and state who indemnifies whom"

Mistake 3: Not Documenting Successful Tracers

Your tracer bullets become your regression test suite. Save them.

Create a folder:

/prompt-tests
  /contract-review
    tracer-1-simple.pdf
    tracer-2-nested-clauses.pdf
    tracer-3-full-length.pdf

Every time you modify the prompt, run all tracers first.

Integration with Testing Checklist

Tracer bullets and the Testing Checklist work together:

  1. Tracer Bullets validate prompt logic during development
  2. Testing Checklist validates outputs before client delivery

Think of it as:

  • Tracer Bullets = Unit tests for prompts
  • Testing Checklist = Quality assurance for outputs

Measuring Success

You know tracer bullets are working when:

  1. You can locate a failure: When a prompt fails, you can say which layer of complexity broke it
  2. You can explain a failure: A failure on a complex document is one you can trace to a layer you have tested
  3. Clearer requirements: The process of creating tracers clarifies what you actually need
  4. Team alignment: Other lawyers can understand prompt behaviour by reading tracer examples

Getting Started

  1. Take your most complex prompt
  2. Ask: "What's the simplest version of this task?"
  3. Create a 1-paragraph test document with obvious right answer
  4. Test it
  5. If it fails, fix the prompt before adding complexity
  6. If it succeeds, add one layer of complexity and repeat

The discipline of starting simple forces clearer thinking about what you're actually asking the AI to do—and that clarity pays dividends throughout the entire workflow.


Building a library of tracer bullet tests for your firm? The consulting page describes what is offered.