What does one verified successful outcome actually cost — once review, escalation, rework and failure are counted?
| Case | Request | AI decision | Reference | Match | Human min |
|---|---|---|---|---|---|
| #0147 | Wrong size, tags still on, day 12 | APPROVE | APPROVE | ✓ | 0.0 |
| #0148 | Worn once, wants full refund | APPROVE | DENY | ✗ | 3.4 |
| #0149 | Arrived damaged, photo attached | EXCHANGE | EXCHANGE | ✓ | 0.5 |
| #0150 | Bundle item, policy unclear | ESCALATE | ESCALATE | ✓ | 4.1 |
| #0151 | Final sale, day 35 | PARTIAL | DENY | ✗ | 2.8 |
| #0152 | Changed mind, unopened, day 5 | APPROVE | APPROVE | ✓ | 0.0 |
Illustrative mock-up. Numbers are examples, not pilot data.
An AI system can cost less per case while creating more review, escalation, correction or downstream errors.
AI execution cost per ticket
Review, escalation, rework and failure
Per verified successful outcome
Illustrative example only. Not pilot data.
E-commerce return decisions: given a return or exchange request, the order facts and your policy, decide the correct outcome and draft the reply.
Customer request
Order facts
Your return policy
Checked against a reference answer set by two reviewers using your written policy.
Several AI configurations run the same historical cases.
AI decides, a reviewer checks. Every review minute is timed.
Your historical decisions, compared against the same reference outcomes.
Only if the data works and both sides see value do we move to a larger evaluation set, typically 200+ cases.
30 anonymized cases to confirm fields, policy and format work end to end. No performance results reported.
Two reviewers independently evaluate each case.
Disagreements are adjudicated separately.
Historical decisions ≠ ground truth.
The same cases run through AI-only and AI + human review.
Human baseline where measurable.
A side-by-side comparison of cost, human time and reliability.
| Approach | Verified success | Cost / submitted | Cost / verified outcome | Human min / outcome | Escalation rate |
|---|---|---|---|---|---|
| Current process | measured | estimated | estimated | estimated | where available |
| AI-only (config A) | measured | measured | measured | measured | measured |
| AI-only (config B) | measured | measured | measured | measured | measured |
| AI + human review | measured | measured | measured | measured | measured |
| Human baseline where measurable | measured | estimated | estimated | measured | where available |
For service providers, the results may also provide structured evidence of outcome quality and AI/human economics that can be shared with clients, subject to the agreed reporting terms.
Dashed fields are optional.
The pilot is free. Reply to our email or write directly — we'll send a redaction checklist and a CSV template for the 30-case schema check.