When not to use a model
Half the AI features I have been asked to build should have been a regular expression, a lookup table, or a better form. Here is how to tell before you spend the sprint.
Contents
A client once asked me to add a language model to extract invoice numbers from emails. The invoice numbers were all in the format INV- followed by six digits. We shipped a regular expression that afternoon, it has a 100% hit rate, it costs nothing, and it has never once been down.
That project is the reason I now start every one of these conversations the same way: what would this look like if we were not allowed to use a model at all?
The test
If you can write down the rule, write down the rule.
Models are for problems where the rule exists but nobody can state it — is this photo blurry, is this sentence rude, does this paragraph answer that question. They are not for problems where the rule is perfectly stateable and somebody simply has not stated it yet.
Three questions, in order:
- Could a careful junior write the rules in an afternoon? Then it is code. Code is testable, debuggable, free, instant, and does not change its mind on a Tuesday.
- Is the answer already in a database somewhere? Then it is a query. I have watched a team ask a model to classify a product’s category when the category was a column in the products table.
- Would a better form make the question disappear? A dropdown with six options beats a model classifying free text into six categories every single time, and it is less work.
Only if all three are no does a model start to make sense.
Things that should not be a model
From actual projects, actual proposals, and one very memorable pitch deck.
Validating input formats. Emails, phone numbers, postcodes, VAT numbers, IBANs. These have specifications. Some of them have checksums. A model gives you a probabilistic answer to a question that has a definitive one.
Extracting anything with a fixed shape. Order IDs, dates in a known format, prices from your own templated emails. If you generated the format, you can parse the format.
Deduplication by exact match. Normalise and hash. Embeddings are for spotting that “cannot log in” and “login broken” are the same complaint, not for spotting that two strings are identical.
Routing with fewer than about eight outcomes and clear keywords. A lookup table with a handful of trigger phrases is transparent, instant, and editable by the support lead without a deployment. Start there. Move to a model when the table stops holding.
Search, as a first attempt. Postgres full-text search is excellent, and you already have Postgres. Reach for vectors once you have watched real queries fail — and you will then know exactly which ones.
Summarising something nobody reads. The cheapest summary is deleting the field.
The costs people forget to count
The API bill is the visible cost, and it is usually the smallest one.
Latency becomes a design problem. A regular expression is microseconds. A model call is hundreds of milliseconds to several seconds. That difference is the difference between validating a field on blur and building a loading state, an error state, a retry, and a timeout.
Failure modes multiply. Your regex does not rate-limit you, go down during someone else’s incident, or deprecate itself with ninety days’ notice. Every model call is a dependency on a system you do not run.
Tests get harder. You can assert a regex. For a model you need a test set, a tolerance, and a decision about what counts as a pass. That is real ongoing work.
Nobody can debug it. When the regex is wrong, you fix the regex. When the model is wrong, you change the prompt, cross your fingers, and find out next week whether you broke three other cases.
It never stops being wrong sometimes. 97% accuracy sounds excellent until you do 10,000 a day and somebody has to look at 300 mistakes, and you have not staffed for that.
Where models earn their keep
I am not against any of this — I ship these features and they are genuinely good. The wins are consistent, and they look like this:
- Unstructured input from the outside world. Text people wrote in their own words, documents in formats you did not design, photos.
- Fuzzy human judgement at volume. Tone, intent, relevance, “is this the same complaint as that one”.
- Generating language. First drafts, rewrites, translation, summaries someone will actually read.
- The long tail after the rules. Handle the 80% you can state with code, and let a model take what falls through.
That last pattern is worth repeating. It is not model-or-code. It is code first, model for the remainder. You get a system that is fast and free for most traffic, debuggable where it matters, and smart exactly where being smart is worth paying for.
The question to ask in the meeting
Not “can we use AI for this” — you can use it for anything, which is what makes it a bad question.
Ask: what is the simplest thing that would work, and what specifically does it fail at?
If nobody can name the failure, you do not have a problem yet. If they can name it precisely — “it misses complaints that never use the word complaint” — then you know what the model is for, you know how to test it, and you know when it has earned its place.