How we measured −70%.
Same dossiers. Manual route versus Compl.AI, logged in the officers' own KPI journals, reconciled line by line against production data.
What was counted, what was not, where it breaks, and the protocol to run the same test on your cases.
manual versus Compl.AI
across the pilot
scope · 875 → 278 is the paired dossiers. the 43 hours is the whole pilot. two scopes, not one number stated twice.
Same dossier, both routes.
The officers kept paired journals. The same case done the manual way, and done with Compl.AI.
A dossier entered the set on the officer's own sign-off. The bank chose, not us.
Real cases off a live production queue. Not a benchmark set.
Every journal line was reconciled against production data. A logged time had to match a screening that actually ran.
We publish no dossier count and no per-officer figure. Eight officers in one bank are identifiable.
The measurement, stated plainly.
Why an average, not a pooled total.
Each officer is one observation. Averaging officers stops the heaviest files deciding the headline.
One case in the pilot carried 1,993 articles. Pool everything and cases like that one write the number.
Pooled, the same dossiers ran 875 minutes to 278. Slightly less than −70%. We say so rather than pick.
Bank-approved deterministic rules can fully automate small, low-risk checks.
Keep them in the pilot and measure straight-through volume, exceptions and officer time.
What we will not claim.
No accuracy number. A figure from our own test set is not evidence to your model-risk function.
So we hand over the method instead, and let your officers produce the number.
Conclusions were compared too, not just clocks.
Speed alone proves nothing. Each pair was also checked for the conclusion it reached.
Where the two routes diverged, the case was root-caused. Every one of them, during the pilot, not afterwards in a report.
Each was fixed and put into production straight away. That loop, not the time saving, is what the officers said they valued.
210 improvements implemented · 4.5 weeks · reported issue to production in minutes to hours
We publish no count of pairs and no agreement fraction. Eight officers in one bank are identifiable and that is their business, not our marketing.
Run it on your own cases.
Ours ran 4.5 weeks in live production. Keep week one in: the learning curve is part of the cost.
Your real cases, your mix. Include the quick low-risk checks, so the spread shows instead of hiding.
Same case, both routes, same officer. Alternate which route goes first, so knowing the answer cancels out.
When it starts, when it stops, whether waiting on a run counts. Decide once, in writing, for both routes.
Both, on every dossier. A time saving without the decision attached to it proves nothing at all.
Weekly, line by line. Journals are memory. Reconciliation is what makes them evidence.
Where the two routes disagree, in writing, before anyone quotes a result from the exercise.
What would prove us wrong.
A test you cannot fail is not a test. Any of these sinks the claim on your book.
One more, whatever the clock says. If a decision cannot be replayed months later, the time saved is worthless.
The trace is what makes that checkable: what the file keeps.
Bring your own cases.
We will run this protocol with you and report what it says, including if it says less than −70%.
Questions on the method: contact@complai.ch.
related · case study · ai & model governance · security and residency