How we test AI before you trust it
Every AI demo works; your Friday-afternoon paperwork is the real exam. How we test AI on your material, your questions and rival models, and share the scores.

Every AI demo you have ever seen worked. That is selection, not evidence: a demonstration runs on material chosen to behave. Your business runs on the other kind — the scanned form with a coffee ring, the recording where two people talk at once, the question asked at five to five on a Friday. Before you trust AI with any of that, someone should have tested it on exactly that. This post is how we do it.
Our starting point is the law we wrote down last time: assume no AI is perfect. A rule like that is only as good as the checking behind it. So we test three ways, and the results go to you.
First: your material, every way it actually arrives
Business information is messy in predictable ways. The same fact reaches you as typed notes, as a Word form, as a page of handwriting, as a recording of a conversation. So when we test document reading, we push the same source material through each form it takes in your business and score every one against the same fixed questions.
In a recent pilot the score sheet read: typed notes, 14 out of 14. The Word form, 14 out of 14. The handwritten scan, 13 out of 14 — and the miss was the system declaring it could not find the answer, which is the failure we want. Recordings sit the same exam after a transcription step. Two details from that exercise are worth knowing. When a document is already digital text, no AI is involved in reading it at all: the system extracts the words exactly, for pennies. And when a scan is beyond confident reading, the system refuses it and says why, so a person handles it. Guessing is not on the menu.
Second: the questions people will actually ask
For AI assistants, we keep a suite of golden questions per project: real questions, asked against the real running system, with the answers checked at the level that matters. Did it pick the right information? Did it respect who was asking? The same pipeline question from a regional manager and a director must produce two different, correctly scoped answers, and both must be right. Did it refuse what it should refuse? We include trick questions on purpose: attempts to pull the assistant off-topic, to extract credentials, to smuggle instructions inside data. Our demonstration sales assistant currently passes a suite of exactly this shape: fifteen questions, plus four deliberate attempts to talk it away from its job. When the model or the prompt changes, the suite runs again.
Third: the models against each other
The AI market comes in tiers. Every provider sells a large, expensive model and smaller, cheaper siblings (Anthropic's Fable, Opus and Sonnet, for instance), and the price gaps between tiers are wide. Which tier a task needs is an empirical question, so we treat it as one. For judgment work we use a blind protocol: one written brief, each model attempts it without knowing it is a contest, and the results are compared side by side. For task work, the golden-question suite re-runs under the candidate model and the scorecards sit next to each other, with the cost per answer alongside. Before we moved a client to a different model, that is the evidence we would want on the table, and it is shared: you see what a cheaper tier gets right and wrong before you choose it. What we will not do is publish a benchmark we have not run; where a comparison has not happened yet, we say so.
Testing never finishes
Models change under their makers' hands, and your data grows odd corners over time. So the suites live with the system: they run before a change ships, and the scorecards are kept. Sometimes the most valuable result is the one that says stop — a step where accuracy will not reach the bar, and the honest recommendation is to keep a person in it. We have given that advice before. Clients tend to hear it as good news, because it means every other step comes with scores.
None of this is exotic. It is the same discipline a good engineer applies to any new material: measure it before you build with it, and keep measuring after. AI is a new material. Treat it like one.
Want your own material put through this? A discovery call is where it starts: book one. Bring your worst-looking document; we mean that.