AI Support Economics Pilot · AI × human work

AI is cheap.But isthe outcome?

What does one verified successful outcome actually cost — once review, escalation, rework and failure are counted?

Historical cases onlyNo integrationNo live customer actionsFree pilot
replay · e-com return decisions · config AILLUSTRATIVE
CaseRequestAI decisionReferenceMatchHuman min
#0147Wrong size, tags still on, day 12APPROVEAPPROVE✓0.0
#0148Worn once, wants full refundAPPROVEDENY✗3.4
#0149Arrived damaged, photo attachedEXCHANGEEXCHANGE✓0.5
#0150Bundle item, policy unclearESCALATEESCALATE✓4.1
#0151Final sale, day 35PARTIALDENY✗2.8
#0152Changed mind, unopened, day 5APPROVEAPPROVE✓0.0

Per verified outcome

Execution$0.18
Human review$0.42
Escalation$0.31
Rework$0.27
Failure$0.58
True cost$1.76

Illustrative mock-up. Numbers are examples, not pilot data.

THE PROBLEM

The price per ticket is not the price per outcome.

An AI system can cost less per case while creating more review, escalation, correction or downstream errors.

Execution+Human review+Escalation+Rework+Failure=True outcome cost
$0.18

Looks cheap

AI execution cost per ticket

+$1.58

Hidden work

Review, escalation, rework and failure

$1.76

True cost

Per verified successful outcome

AI execution
$0.18
+ Human review
$0.42
+ Escalation
$0.31
+ Rework
$0.27
+ Failure cost
$0.58
True cost
$1.76

Illustrative example only. Not pilot data.

THE WORK UNIT

One task.
Five possible outcomes.

E-commerce return decisions: given a return or exchange request, the order facts and your policy, decide the correct outcome and draft the reply.

Inputs

Customer request
Order facts
Your return policy

→

Decision

✓ APPROVE× DENY◐ PARTIAL↔ EXCHANGE↑ ESCALATE
→

Verified outcome

Checked against a reference answer set by two reviewers using your written policy.

Decision onlyNo refunds issued, no orders changed. Execution stays with you.
Low integrationNo access to your helpdesk, store or payment systems.
Escalation is costed, not freeA correct ESCALATE counts as correct — the human time it triggers still counts.
THE COMPARISON

Same work.
Different economics.

AI ONLY

Fast, inexpensive execution

Several AI configurations run the same historical cases.

How often is it actually right?
AI + HUMAN

More human minutes

AI decides, a reviewer checks. Every review minute is timed.

Does reliability offset the cost?
CURRENT PROCESS

Your existing operation

Your historical decisions, compared against the same reference outcomes.

What is the true baseline?
THE PILOT

We begin with just 30 cases.

Only if the data works and both sides see value do we move to a larger evaluation set, typically 200+ cases.

01

Schema check

30 anonymized cases to confirm fields, policy and format work end to end. No performance results reported.

02

Verified reference outcomes

Two reviewers independently evaluate each case.
Disagreements are adjudicated separately.
Historical decisions ≠ ground truth.

03

Evaluation

The same cases run through AI-only and AI + human review.
Human baseline where measurable.

04

Results

A side-by-side comparison of cost, human time and reliability.

WHAT YOU RECEIVE

Every number labelled by where it comes from.

ApproachVerified successCost / submittedCost / verified outcomeHuman min / outcomeEscalation rate
Current processmeasuredestimatedestimatedestimatedwhere available
AI-only (config A)measuredmeasuredmeasuredmeasuredmeasured
AI-only (config B)measuredmeasuredmeasuredmeasuredmeasured
AI + human reviewmeasuredmeasuredmeasuredmeasuredmeasured
Human baseline where measurablemeasuredestimatedestimatedmeasuredwhere available
measured observed or timed during the pilotestimated from your inputs: handling time, labour cost, rework, cost of a wrong decision

For service providers, the results may also provide structured evidence of outcome quality and AI/human economics that can be shared with clients, subject to the agreed reporting terms.

DATA & PRIVACY

Anonymized history.
You decide what leaves your systems.

What we need

30 anonymized casesHistorical return / exchange requests
Your policyThe policy that applied to those cases
A few cost inputsHandling time · labour cost · wrong-decision cost
View required CSV fields
case_idrequest_textrequest_dateorder_datedelivery_dateproduct_categoryorder_valueagent_decision item_conditionagent_replyfollow_uphandling_minutes

Dashed fields are optional.

How we handle it

  • Remove before sharing — names, emails, phones, addresses, payment details and identifying order numbers.
  • Historical only — no customers contacted, no actions on real orders.
  • Pilot use only — deleted on request when the pilot ends.
  • Your client's consent — if cases belong to a brand you serve, confirm you may share anonymized cases.
  • Mutual NDA — available before any data is shared.
One optional question, asked separately. After the evaluation, we may ask whether a small set of the same anonymized cases can be evaluated confidentially by other independent providers under NDA, to compare how different providers price and perform the same work. It never happens without your explicit written consent.

Measure the
outcome.

The pilot is free. Reply to our email or write directly — we'll send a redaction checklist and a CSV template for the 30-case schema check.