A prompt in your API is a feature. We test it like one.
Two public experiments, and two rules we now build in everywhere, show why a prompt needs tests, a log and a loop, like any other feature.

More and more of the briefs we see now include a line like "and the API should use AI for this bit". Summarise the notes. Rank the options. Draft the email. It is a reasonable ask, and the tempting answer is to paste a prompt behind an endpoint, watch it work once, and ship. We have spent the last month running two public experiments that show, in scoreboard form, why we don't.
A prompt behind an API is a feature
It needs the same engineering as a database query or a calculation in custom logic: rules about what it may and may not do, tests against real material, measurement of what it costs and how often it is right, and a way to change it when the evidence says so. Skip any of those and you have not built a feature. You have built a guess with a URL.
We wrote down our rules for what an AI in business software should and should not be able to do, and how we test one before a client depends on it. This post is about what happens after that: the running, the measuring, and the changing of prompts once the numbers come in.
What the football taught us in one week
On Friday we gave nine AI models from three makers an identical brief: predict ten Premier League scorelines from the same league table and form guide. On Tuesday we marked them. The models called the right winner 40% of the time, little better than chance, and they were wrong together: all nine backed Chelsea to beat Hull, when the table in front of them showed Hull unbeaten and unscored-against and Chelsea leaking goals. A human reading the same two lines called it a draw and took the point.
The cost column was the other finding. The whole nine-model round cost ten and a half cents. The most expensive point cost 74 times the cheapest, and the three cheapest models scored exactly what Claude Opus 5 and Gemini Pro scored. One round proves nothing, and we have said so. But the shape of the lesson is the one clients need: the machine can be confidently, unanimously wrong on evidence it was handed, and paying more for it bought nothing measurable that week.
Two rules we now build in, before a client ever sees the feature
The same loop runs on real work, without a scoreboard, and two of its lessons have hardened into standing rules.
A faithful pipeline propagates an upstream error. Speech-to-text hears one word as another; a handwriting model reads a seven as a one; the extraction step downstream believes it and writes it into the record with complete confidence. Neither model has failed by its own lights. So nothing an AI captures reaches a record until a person has confirmed the transcript it was extracted from. The AI does not get less useful. It gets a gate.
An AI that ranks or recommends must not be able to invent. When the job is to choose between options, the hard rules run first as plain arithmetic, and the AI ranks only what survives; nothing the rules exclude can come back. Bellwether, our demonstration CRM, shows the other half of the same principle. Its assistant can look anything up and can propose a change, but it cannot save one: a proposed discount becomes a confirmation for a person to approve, and until they do, the quote is untouched. That refusal is a feature we built, not a limitation we apologise for.
Where we are confident, and where we are not
Some of this is settled. Transcription of clear speech, reading a typed document, and pulling answers out of a transcript against known questions all work well enough that we will put them in front of a client with a review step and a log. We say so in the specification.
Prediction against a noisy world is another matter, and football is the honest extreme of it. Anything where the model must weigh a table against a reputation is a place to test before you trust, because week one showed nine models reading the reputation. And the price argument cuts both ways. A call that costs a fraction of a cent is rarely the constraint, so the money is in getting it right, not in running it. But cheap is not the same as consequence-free. A wrong answer extracted into a compliance record costs a great deal more than the token did.
How we do it
On our platform a prompt is content, not code. Its structure and its permissions live in the project definition with everything else, but the words, the model and the settings can be changed by the client's own administrator at runtime, and tried live against real input before saving. Every call writes a log line: which model answered, how many tokens, how long it took, and who asked. Never the text itself.
That gives us the loop. We run a fixed test set. We read the log for cost per model and cost per correct answer. We change the prompt, or the model, and run the set again. The football league is that loop in public: after one round we have rewritten the prompt, added midweek fixtures to the evidence pack, and will run both versions side by side on Friday so the difference is measured rather than argued about. We do not know the result yet. We do not even know the predictions until Friday. That is the point.
The experiments read as fun, and they are. They are also the test bench, and the scoreboard is the only reason to believe us when we say a prompt is engineered rather than pasted.