Proof, not promises.
What you get
Every solution ships with a scorecard.
Criteria are agreed per task type before the work starts, then scored against held-out examples from your own data. Nothing ships until its row clears, which sometimes means a row does not.
| Task type | Metric | Threshold | Result | Status |
|---|---|---|---|---|
| Document extraction | Field-level accuracy | ≥ 95% | 97.2% | Cleared |
| Line-item classification | Macro F1 | ≥ 0.90 | 0.94 | Cleared |
| Demand forecast | MAPE | ≤ 12% | 14.1% | Held. Held. Re-scoped to weekly buckets, re-scored at 9.8%, shipped two weeks later. |
| Generated summary | Groundedness | 100% traced | 100% | Cleared |
Evaluation does not end at go-live.
A solution clears a bar before it ships, then keeps being measured against that same bar. The criteria are written once and used twice.
Illustrative figures, shown to explain the format. Published accuracy is reported as ranges across deployments, never as a single figure from a single account.
What makes it different
Two differences, and neither is a better score.
Every vendor will show you a groundedness number, and almost all of them mean the same thing by it: did the answer match the text a retriever handed over. What differs here is what that number resolves to, and who is left running it.
Most evaluation: A chunk of text a retriever judged relevant
Here: A column in your warehouse
Most evaluation: The model was consistent with what it read
Here: The figure carries the field it was drawn from, so it can be followed home
Most evaluation: An AI engineering team you hire, on a tool you license
Here: The team delivering the solution, against criteria agreed before the work starts
Most evaluation: A license, and evaluations somebody has to keep current
Here: The solution and the evidence it works, with the method reused by the next one
Evaluation platforms do their job well. The question is who operates them, and what happens to the evidence when nobody does.
What we publish.
Method and ranges rather than single figures or named accounts. Published this way, the work ships on a cadence instead of waiting on a disclosure decision each time.
How we evaluate before go-live
The harness itself: acceptance criteria per task type, the regression suite, and what good enough to ship actually means.
Model selection benchmarks
How leading models perform on real workload types, extraction, classification, forecasting and code generation, refreshed each quarter.
Accuracy by task type
Published ranges and the method behind them, across deployments rather than from any single account.
Groundedness measurement
How we test that an answer is supported by its source, what it is scored against, and what happens when it is not.
Human acceptance rates
How often output is accepted unchanged versus corrected, and how that curve moves over a deployment's life.
Got questions? We're here to answer them for you
Have more questions?
Contact our support team to get what you need.

