Skip to content
Back to Methodology
quality assuranceadvanced

The Staircase of Complexity: Systematic Improvement

Build your prompt library systematically by starting with simple use cases and progressively adding complexity. Create a repeatable process for continuous improvement based on real failure data.

advanced level

The Problem

A firm can approach legal AI improvement reactively: something breaks, they fix it, they move on. Months later, they're still finding new failure modes, still patching individual issues, still lacking systematic quality improvement.

This is firefighting, not engineering.

The Staircase Principle

Improvement should be deliberate, sequential, and measurable.

Like climbing a staircase:

  1. You know which step you're on (current capability level)
  2. You know which step is next (next complexity increment)
  3. You test each step before climbing higher (validation gates)
  4. You can always return to previous steps (regression testing)

The result is a systematic progression from simple use cases to complex production workflows.

The Five Phases

The durations and the percentages in the success criteria below are suggested starting points. They are not measured thresholds, and no results have been published for them. Set your own for the risk of the work.

Phase 3 onwards uses real documents. What you may put into an AI tool depends on the plan your firm has, its agreement with the vendor and your firm's policy: Claude Projects for Client Matters sets out the checks for one tool. Removing names does not by itself make a document safe to upload, because the facts in it can still identify a client.

Phase 1: Foundation (Weeks 1-2)

Goal: Establish basic functionality with simple use cases

Activities:

  • Define the core task in simplest terms
  • Create tracer bullet test cases
  • Build initial prompt library
  • Document baseline performance

Success criteria:

  • 90%+ success rate on tracer bullets
  • Clear documentation of what works
  • Identified failure modes catalogued

Example - Contract Review:

Simple use cases:
✓ Identify indemnity clause in standard agreement
✓ Extract key dates from two-page NDA
✓ Summarise three-paragraph termination provision

Don't move to Phase 2 until these simple cases work reliably.

Phase 2: Controlled Complexity (Weeks 3-4)

Goal: Add one complexity variable at a time

Activities:

  • Introduce ambiguity in test cases
  • Test with longer documents
  • Add edge cases and exceptions
  • Measure degradation patterns

Success criteria:

  • Maintain 85%+ success rate with added complexity
  • Understand which complexity factors cause failures
  • Documented mitigation strategies

Example - Contract Review (Continued):

Added complexity:
✓ Identify indemnity in non-standard language
✓ Extract key dates from 20-page services agreement
✓ Summarise termination provision referencing other sections

Each addition isolates a single complexity variable (ambiguity, length, cross-references).

Phase 3: Real-World Scenarios (Weeks 5-8)

Goal: Test with actual documents and client situations

Activities:

  • Use real contracts, with everything that identifies a client removed
  • Test variations from different counterparties
  • Include industry-specific language
  • Compare against a solicitor's own analysis of the same document

Success criteria:

  • 80%+ success rate on real documents
  • Failure modes understood and documented
  • Clear guidance on when human review is essential

Example - Contract Review (Continued):

Real scenarios:
✓ Review supplier agreements from 5 different industries
✓ Compare standard forms from 3 major law firms
✓ Analyse heavily negotiated final agreements

Phase 4: Production Deployment (Weeks 9-12)

Goal: Live use with validation safeguards

Activities:

  • Implement all five validation layers
  • Track all outputs and corrections
  • Collect user feedback
  • Measure time savings vs. quality impact

Success criteria:

  • Quality equals or exceeds human baseline
  • Documented time savings
  • User adoption and satisfaction
  • Established feedback loop for improvement

Phase 5: Continuous Improvement (Ongoing)

Goal: Systematic refinement based on data

Activities:

  • Weekly review of failures and corrections
  • Monthly prompt library updates
  • Quarterly capability assessments
  • Regression testing against all previous phases

Success criteria:

  • Failure rate trending downward
  • Expanding use case coverage
  • Increasing user confidence
  • Measurable ROI

What to Measure

Track these metrics at every phase:

Quality Metrics

Accuracy:

  • % of outputs requiring zero corrections
  • % requiring minor corrections
  • % requiring major revisions
  • % that are unusable

Consistency:

  • Variance across different users
  • Variance across document types
  • Variance across time periods

Completeness:

  • % of relevant provisions identified
  • % of false positives
  • % of false negatives

Efficiency Metrics

Time savings:

  • Time to first draft (AI vs. human)
  • Time to final output (including validation)
  • Net time saved per task

Scalability:

  • Number of documents processed per hour
  • Maximum document complexity handled
  • Team members successfully using AI

Risk Metrics

Failure modes:

  • Types of errors by category
  • Severity distribution
  • Near-misses (caught in validation)
  • Actual errors (reached client)

Building Your Failure Database

The most important asset for systematic improvement is a comprehensive record of failures.

What to document:

For every failure (error caught in validation or by client):

  1. Context:

    • Document type and complexity
    • Prompt used
    • AI model and settings
    • User and validation path
  2. Failure details:

    • What went wrong?
    • What should it have been?
    • How was it discovered?
    • What layer caught it? (if caught in validation)
  3. Root cause:

    • Prompt ambiguity?
    • Context window issue?
    • Edge case not anticipated?
    • AI knowledge limitation?
  4. Remediation:

    • How was it fixed?
    • Prompt changes made?
    • New validation added?
    • User training needed?
  5. Prevention:

    • How to prevent recurrence?
    • New test case added?
    • Documentation updated?
    • Process changed?

Example failure log:

This entry is made up to show the format. It is not a record of a real matter.

Failure #47
Date: [date]
Context: 50-page supplier agreement, [tool and model version], standard review prompt
Failure: Missed critical change-of-control provision in clause 14.3
Should have been: Flagged as material provision requiring notice
Discovered by: Client questioned why not mentioned in summary
Root cause: Provision was in the "General" clause, prompt weighted earlier/titled clauses more heavily
Remediation: Updated prompt to explicitly review all clauses including "General" and the schedules
Prevention: Added test case with critical provisions in general clauses
Layer that should have caught: Layer 3 (adversarial) should flag missing material provisions
Regression test added: Test-case-47-general-clauses.pdf

Systematic Prompt Improvement

Use failure data to improve prompts systematically:

Monthly Prompt Review Process

Step 1: Analyse failure patterns

  • Group failures by root cause
  • Identify most common issues
  • Calculate impact (severity × frequency)

Step 2: Prioritise improvements

  • Address high-impact failures first
  • Batch similar issues
  • Quick wins before complex rewrites

Step 3: Update prompts

  • Make targeted changes
  • Version control (prompt v1.0 → v1.1)
  • Document what changed and why

Step 4: Regression test

  • Run new prompt against all historical test cases
  • Ensure old cases still pass
  • Verify new cases now pass

Step 5: Document and deploy

  • Update prompt library
  • Train users on changes
  • Monitor for new issues

Example prompt evolution:

The success rates below are made-up figures that show how to record progress. They are not measured results.

v1.0 (Week 1):

Review this contract and identify key provisions.
  • Success rate: 60%
  • Issues: Vague, inconsistent results

v1.1 (Week 2):

Review this contract and identify:
- Indemnities
- Termination rights
- Limitation of liability
- Dispute resolution

For each, provide clause number and brief summary.
  • Success rate: 75%
  • Issues: Missed cross-references, summary too brief

v1.2 (Week 4):

Review this contract systematically:

1. Read the entire document first
2. Identify all clauses containing:
   - Indemnities
   - Termination rights
   - Limitation of liability
   - Dispute resolution
3. For each provision, provide:
   - Clause number
   - Full text of key language
   - Summary of rights/obligations
   - Any cross-references to other clauses
4. Note any unusual or non-standard provisions
  • Success rate: 85%
  • Issues: Struggles with heavily negotiated/marked-up docs

v1.3 (Week 8):

[Previous instructions]

Additional instructions:
- If the language is ambiguous, provide both interpretations
- If provisions conflict, flag the conflict
- Review the "General" clauses and the schedules - material provisions sometimes hide there
- For cross-references, include the referenced text for context
  • Success rate: 90%

Each version addresses specific failures from the database. Progress is measured and documented.

Worked Example: Building a Disclosure Review Process

This example is illustrative. The figures show what a record of progress looks like; they are not results from a real project.

Month 1 - Foundation:

  • Simple emails, obvious relevance (100 test cases)
  • Success rate: 92%
  • Issues: Struggles with partially relevant emails

Month 2 - Controlled Complexity:

  • Email threads (preserving context)
  • Success rate: 88%
  • Issues: Context from earlier emails lost, technical jargon misunderstood

Month 3 - Real Documents:

  • Actual client emails, with identifying details removed
  • Success rate: 83%
  • Issues: Industry-specific acronyms, implicit references

Month 4 - Production:

  • Live review with validation
  • Success rate: 87% (after prompt improvements)
  • Checked by a senior solicitor before the disclosure list is prepared

Month 6 - Continuous Improvement:

  • 47 failure cases documented
  • 8 prompt versions deployed
  • Success rate: 93%

Regression Testing

Critical rule: When you improve a prompt, verify it didn't break old use cases.

Build a regression test suite:

  1. Tracer bullets from Phase 1 (always must pass)
  2. Edge cases from Phase 2 (document degradation tolerance)
  3. Real documents from Phase 3 (validate production quality)
  4. Historical failures (never repeat past mistakes)

Run regression tests:

  • Before every prompt update
  • Weekly on production prompts
  • After the vendor changes the model behind your tool, or you move to a different tool

Automation:

# Pseudocode for regression testing

def regression_test_suite(prompt_version):
    results = []

    for test_case in test_library:
        output = run_prompt(prompt_version, test_case.input)
        passed = validate_output(output, test_case.expected)

        results.append({
            'test': test_case.name,
            'passed': passed,
            'prompt_version': prompt_version
        })

    # Alert if any test that previously passed now fails
    regressions = find_regressions(results, previous_results)

    return results, regressions

How This Fits with the Rest of the Methodology

Systematic improvement orchestrates everything:

This methodology is the process that ties the others together.

Common Mistakes

Skipping to Production Too Early

"We tested it on a few documents, seems good, let's deploy!"

Problem: You haven't discovered the failure modes yet. They'll appear in production, undermining confidence.

Solution: Systematic progression through all five phases. Don't skip.

Not Documenting Failures

"We fixed that issue, we're good now."

Problem: Without documentation, you'll repeat the same mistakes, can't identify patterns, can't measure improvement.

Solution: Mandatory failure logging. It's not bureaucracy—it's the data that drives improvement.

Changing Too Many Variables at Once

"We updated the prompt, changed models, AND modified the validation process."

Problem: When success rate changes, you don't know what caused it.

Solution: Change one variable at a time. A/B test when possible.

No Regression Testing

"The prompt works great for our new use case!"

Problem: You broke three old use cases that were working fine.

Solution: Build regression test suite from day one. Run it before every change.

Getting Started

Week 1 action items:

  1. Define your core use case in simplest terms
  2. Create 10 tracer bullet test cases with known-good answers
  3. Build your first prompt (version 1.0)
  4. Test and document results
  5. Set up failure tracking (even a spreadsheet works)

Week 2:

  1. Add complexity to 3 test cases
  2. Update prompt based on Week 1 failures
  3. Regression test - verify original cases still pass
  4. Document what you learned

Week 3-4:

Continue staircase progression: simple → complex → real → production

Long-Term Vision

After 6 months of systematic improvement:

  • Library of prompts with recorded test results for common tasks
  • Comprehensive test suite covering known failure modes
  • Data-driven understanding of AI capabilities and limitations
  • Documented improvement trajectory showing ROI
  • Team confidence in AI-assisted workflows

This is what mature AI adoption looks like: not ad-hoc experimentation, but systematic, measurable, continuously improving professional processes.


Need help building systematic improvement processes for your firm? The consulting page describes what is offered.