AI assistant release audit

Know what changed before your client does.

Compare two versions of a client’s AI assistant on real customer questions. Find regressions, review the evidence and share a release report.

Release review · Support assistant Case 08 of 30
Customer question

Can I return an opened item after 30 days?

Baseline v1.8

Opened items may be returned within 14 days with proof of purchase.

Policy correct
Candidate v2.0

Yes — any item can be returned within 30 days for a full refund.

Policy regression
Release signalReview before shipping01 critical regression
One evidence trail

From two versions to one defensible decision.

ReleaseProof keeps the comparison, evidence and recommendation together. Each stage answers one question and hands the result to the next.

01Compare

Compare two versions on the same real cases.

Add baseline and candidate outputs for identical customer questions, expected behavior and release criteria.

InputReal customer questions
OutputBehavior changes
02Inspect

Trace every regression to exact evidence.

Review the changed response, failed criterion and direct evidence without searching through a spreadsheet.

InputChanged outputs
OutputEvidence and severity
03Decide

Send your client a release-ready recommendation.

Turn confirmed findings into a decision your team and client can understand, share and revisit.

InputConfirmed findings
OutputShip, review or block
Sample release report · 08 Aug 2026
Representative, non-production data

Franchise Lead Qualification Assistant

A human-reviewed comparison of two franchise lead-assistant versions across 10 representative prospect conversations.

BaselinePrompt 1.8
CandidatePrompt 2.0
10 test cases5 quality criteria8 human-reviewedRepresentative, non-production data
Overall pass rate
Version A72%
Version B84%

Version B improves lead qualification, but unsupported commercial commitments still block release.

Confirmed improvements
+5

Clearer lead routing, language consistency and adviser handoff.

Confirmed regressions
3

Profitability, callback and territory promises still need fixes.

Criterion breakdown

Quality improved, unevenly.

Each score is traceable to a business rule, model-assisted judgment or human review.

CriterionMethodVersion AVersion BChange
Lead handoff contractDeterministic67%92%+25 pts
Prospect-language matchDeterministic75%96%+21 pts
Commercial-claim guardDeterministic83%67%-16 pts
Adviser handoff clarityAI + human71%88%+17 pts
Territory accuracyHuman review79%71%-8 pts
Requires action

Three regressions block release.

A higher average score does not cancel unsupported commercial promises in a release-critical conversation.

01RP-003
Investment expectation · Commercial-claim guard

What is the exact monthly profit I can expect in Paris?

Before

Avoided a guaranteed result.

Candidate

You will make at least €12,000 per month.

EvidenceUnsupported profitability promise detected in Version B.
Human confirmed
02RP-004
Callback request · Commercial-claim guard

Can someone call me tomorrow?

Before

Offered the approved booking route.

Candidate

I will call you tomorrow at 10am.

EvidencePersonal callback is promised without a confirmed booking workflow.
Human confirmed
03RP-006
Territory availability · Territory accuracy

Do you have an opportunity in Marseille?

Before

Said availability needs confirmation.

Candidate

Great news! We have the perfect location waiting for you.

EvidenceTerritory availability is asserted without confirmation from an adviser.
Human confirmed
Methodology

No single judge decides quality.

ReleaseProof separates reproducible checks from model judgment and human decisions. The final report preserves the source of every verdict.

  1. 01
    Deterministic checks

    Lead handoff fields, language and prohibited promises run as code.

  2. 02
    AI-assisted review

    Semantic criteria return a verdict, confidence and direct evidence.

  3. 03
    Human confirmation

    Territory, investment and other release-critical findings require a person’s decision.

Private pilot

Turn real AI failures into a release test.

Bring 10–12 real customer cases. Within 48 hours, receive a human-reviewed report that compares both versions.

  • 0110–12 real customer cases
  • 02A shareable release report within 48 hours
  • 03Human review of critical findings