Skip to content

A prompt is a spec, so write it like one

Prompt engineering has very little to do with magic words. It is the same skill as writing a clear ticket for a capable contractor who will not ask you any questions.

4 min read
Contents

The best prompt engineer I have worked with had never read a paper on the subject. She was a project manager. She was simply very good at writing briefs, and it turned out that was the whole job.

Think of a model as an exceptionally fast contractor who has read most of the internet, has no access to your codebase, cannot ask a follow-up question, and will confidently guess rather than admit confusion. Everything that makes a good ticket for that person makes a good prompt.

The anatomy

Almost every prompt I ship in production has the same five parts, in this order.

Role and context. Who is answering and on whose behalf. One or two sentences. You are the support assistant for ShipMyForm, a hosted form backend. You are talking to a developer who is integrating it. This is not flattery — “you are a world-class expert” does nothing — it is scoping.

The task. One sentence, imperative. If you cannot write it in one sentence, you have two tasks and you should make two calls.

The constraints. What must and must not happen. Length, tone, what to do when it does not know, what it is not allowed to say.

The output format. Exactly what shape you want back, shown rather than described.

The examples. Two or three input/output pairs. This is the highest-leverage part of the entire prompt, and it is the one people skip.

Show, do not describe

Compare these two instructions for the same job.

Return the result as JSON with the customer's name, their
email, the urgency, and a one-line summary.
Return only JSON in exactly this shape:

{
  "name": "Sara Iqbal",
  "email": "sara@example.com",
  "urgency": "high",
  "summary": "Card declined three times on checkout."
}

urgency is one of: low, normal, high.
If a field is not present in the message, use null.

The second one does not get argued with. It shows the casing, the key order, the fact that urgency is a fixed vocabulary, and what absence looks like. I have watched the first version produce customerName, customer_name and name across three consecutive calls.

Examples do the work your adjectives cannot

“Write in a friendly but professional tone” means nothing. Every team means something different by it, and so does the model.

Two real examples of your actual house style do more than a paragraph of adjectives ever will. This is the same reason we write tests instead of writing “the function should handle edge cases properly” in a comment.

Pick examples that sit near the boundaries of the task, not in the comfortable middle. The interesting example is the ticket that is angry but not abusive, or the message that contains two questions instead of one. Middle-of-the-road examples teach almost nothing.

Give it an exit

The most damaging default behaviour of a language model is that it never declines. Ask it something unanswerable and it answers anyway.

So write the exit in:

If the message does not contain an email address, return
{"error": "no_email"} and nothing else.

Now “I don’t know” is a valid, named, parseable outcome instead of a hallucination. This one line has saved me more debugging than any other prompt technique.

Let it think before it answers

For anything involving reasoning — classification with fuzzy boundaries, deciding between several actions, checking a document against a rule — ask for the reasoning first and the answer last:

{
  "reasoning": "The user mentions a chargeback and a deadline, so this is billing rather than technical.",
  "category": "billing"
}

Field order matters, because the model generates left to right and what it has already written conditions what comes next. Reasoning before the answer genuinely improves the answer. Answer before the reasoning just produces a justification for a guess it already made.

Strip the reasoning field before you show anything to the user. It is for you.

Treat prompts like code, because they are

This is the part that separates a demo from a product.

Keep them in version control, as files, not as strings pasted into a database by whoever was on call. prompts/support-triage.v3.md.

Write tests. Twenty inputs with expected outputs, run on every change. They will not all pass exactly — the output is text — so assert on what matters: the JSON parses, the category is in the allowed set, the summary is under twenty words, the refund amount matches.

Never edit one in production to fix a single bad case. That is the prompt equivalent of a if (userId === 4021) branch. Add the case to your test set, change the prompt, run the set, see what else you broke. Something always breaks.

Version them alongside the model. A prompt tuned on one model is not guaranteed on the next. When you upgrade the model, re-run the tests before you re-run the celebration.

What does not help

A few things I have watched people spend real time on, with no measurable effect: telling the model it is an expert, offering it a tip, threatening it, adding “IMPORTANT!!!” in capitals, and stacking six polite ways of saying the same constraint.

What does help is being specific about the task, showing the output, giving concrete examples, and providing an escape hatch. It is unglamorous. It is also just spec writing, which is a skill your team already has and probably already undervalues.

Share