Sign in with Google

Prompt Engineering Is Not the Skill You Need

The incantations are dying. What replaces them is an old, boring, extremely valuable skill: writing a spec.

Mochivia9 min read

There was a period, roughly two years long, when telling a model "you are a world-class expert" measurably improved its answers. There was a period when offering it a tip helped. When "take a deep breath and work through this step by step" was a real technique with a real effect size.

Most of that is gone, and it went quietly. The models got better at following plain instructions, the incantations stopped doing anything, and nobody circulated a correction to the people still copying them out of a listicle from 2023.

So here is the position, stated plainly: prompt engineering as a collection of phrasings is a decaying asset, and it was never the thing making the difference anyway.

What was making the difference — then and now — is much less exciting and much more durable. It is the ability to write down what you want with enough precision that a competent stranger could produce it without asking you a question.

That is not a new skill. It is technical writing. It is requirements definition. It is the thing that separates a project brief that works from one that generates six rounds of clarification. And it climbs in five rungs.

Why the Incantations Worked at All

Worth understanding, because the mechanism tells you which techniques will survive the next release and which will not.

A language model is trained to continue text plausibly. The architecture underneath nearly all of them is the transformer, introduced in Attention Is All You Need in 2017, and the practical consequence of the design is that everything in the context window influences what comes next. Early instruction-following was crude, so the text you supplied did double duty: it told the model what you wanted, and it also nudged the statistical neighborhood the continuation would be drawn from.

"You are a world-class expert" worked because it shifted the neighborhood toward text written by people who sound like experts. It was not persuasion. It was a retrieval cue for a region of the training distribution.

Then instruction tuning got dramatically better. A modern language model already treats your request as a request, so the nudge is largely redundant — the model was going to produce its best attempt regardless of whether you flattered it first. Techniques that worked by exploiting a weakness in instruction-following die when instruction-following improves. Techniques that supply information the model did not have cannot die, because information is not a trick.

That single distinction sorts the entire field.

What Genuinely Still Matters

Being contrarian about this only counts if you are fair about what survived. Four things did, and all four survive for the same reason: they add information or structure rather than pressing a behavioral button.

  • Examples. Showing two or three instances of what good output looks like remains the single highest-leverage move available, and it is not close. A description of your house style is ambiguous. An example of your house style is not.
  • Output format. Specifying the shape you want back — a table with these columns, JSON with these keys, six bullets of at most fifteen words — reliably improves usability and is not going anywhere, because the model genuinely cannot guess your downstream consumer.
  • Decomposition. Breaking a task into ordered steps still helps on genuinely multi-stage work, not because of a magic phrase but because you are supplying the plan. This one is fading for easy tasks and holding for hard ones.
  • Context placement and volume. What you put in the window matters enormously — the relevant document, the actual data, the real constraints. Provider guidance from Anthropic and OpenAI converges on this, and it is the least glamorous and most reliable advice in the field.

Notice that none of those four are phrasings. They are all acts of specification. Which is the whole argument.

The Specification Ladder

Five rungs. Most people live on rung one and blame the model for what happens there.

Rung 1 — the vague ask. "Write a summary of this report." A request this open has no single correct answer, so you get the statistical average of every summary that request could describe. The average of a wide range is always mediocre, and this is the origin of nearly every complaint about AI output being generic.

Rung 2 — constraints. Add the things that would make an answer wrong. Length. Audience. What to exclude. Tone. Reading level. Constraints do more work than instructions do, because they collapse the space of acceptable outputs rather than gesturing at a direction inside it.

Rung 3 — context. Supply what the model cannot infer: who the reader is, what decision this feeds, what was already tried, what the political sensitivities are, the actual source document rather than your paraphrase of it. Most output that is technically correct and practically useless died at this rung.

Rung 4 — evaluation criteria. State how you will judge the result before you see it. "Good means a controller could act on it without asking a follow-up question." This is the rung almost nobody climbs, and it changes both sides of the transaction: the model orients toward your bar, and more importantly you now have a bar, which means you can tell whether the output met it instead of reacting to how it reads.

Rung 5 — the verification loop. Specify how the claims will be checked and by whom. Which facts require a source. Which numbers get recomputed independently. What you will do with a claim you cannot verify. Generation is cheap now and checking is the bottleneck, so the spec is incomplete until it says how checking happens.

What Rung Five Looks Like Written Down

The same task, specified. Note that there is not a single clever phrase in it — no roleplay, no flattery, no incantation. It is a brief.

markdown
TASK
Summarize the attached Q3 vendor performance report.

AUDIENCE + DECISION
Our COO. She is deciding whether to renew two of the five vendors
at the November review. She has not read the report and will not.

CONSTRAINTS
- 250 words maximum, plus one table.
- No recommendation from you. She makes the call.
- Exclude anything about the vendors we already exited in Q1.
- Plain language. No vendor marketing terms.

OUTPUT FORMAT
- One paragraph: what changed since Q2.
- Table: vendor | on-time rate | defect rate | cost delta | open issues
- Three bullets: the open questions she should ask in the review.

WHAT GOOD LOOKS LIKE
She can walk into the review having read only this, and ask a
question that surprises the vendor. If she has to ask me a
clarifying question first, it failed.

VERIFICATION
- Every number must cite the page it came from.
- Flag any figure the report gives more than once inconsistently.
- If a metric is missing for a vendor, say so. Do not estimate it.

That takes four minutes to write and it is reusable forever, which is the part people miss. A clever phrasing is worth one session. A spec is an asset you edit monthly and hand to a new hire.

It also does something a prompt cannot. Half the time, writing rung four honestly reveals that you did not know what you wanted — and discovering that before you generate six drafts is worth more than any output the model could have produced.

The Cost of Optimizing the Wrong Layer

Spending a year collecting prompt techniques is a specific and common way to work hard on the wrong thing. It has all the surface features of skill development — there is jargon, there are practitioners, there is measurable improvement on toy tasks — and it produces almost nothing transferable, because the substrate it depends on is being actively engineered away by the model providers. That failure pattern has a general form we have written about in learning the wrong things.

The market has already priced this in. "Prompt engineer" as a standalone job title has largely been absorbed into ordinary engineering and product work; the prompt engineer role page is honest about how that consolidation went and what the work turned into. Specification, evaluation, and context design survived. Incantation did not.

Meanwhile the underlying skill is unusually portable. Writing an unambiguous brief works on models, contractors, agencies, junior colleagues, and your own future self at 8am on a Monday. It is the rare case where the AI-adjacent skill is also just a good professional skill.

There is a diagnostic that settles the argument for most people. Take a prompt you are proud of and delete every stylistic flourish from it — every "you are an expert," every "this is very important to my career," every please and thank you. Keep only the context, the constraints, the example, the format and the bar. Run both. In the overwhelming majority of cases the stripped version performs the same or better, and it is half the length. Whatever you deleted was not the skill.

Where the Ladder Sits in the Larger Skill

Specification is one of five verbs that make up practical competence with these tools — the others being verify, delegate, integrate and decide, each with its own drill in the AI skills every professional needs. And all five rest on a working mental model of the machine, which is the four-layer argument in how to learn AI skills that don't expire.

Mochivia teaches it in that order deliberately. A course that opens with prompting techniques produces someone who can reproduce techniques and diagnose nothing; a course that opens with mechanism produces someone who can derive the technique a novel task needs. If you want the narrow tactical version right now, the walkthrough in writing better AI prompts is the practical companion to this argument.

The ladder needs none of that. It is five rungs and a text file.

Climb One Rung

Do not rewrite how you prompt. Pick the one recurring task where the output is reliably almost-good, and add rung four to it: write down, in a sentence, how you will judge the result before you see it.

That one sentence outperforms every phrasing trick you have saved, and it will still work in three years, on a model that does not exist yet, made by a company that may not exist yet.

There is a nice second-order effect too. Teams that write specs instead of prompts end up with something they can review, argue about, and improve together. A prompt lives in one person's chat history and dies there. A spec sits in a shared document where a colleague can read rung four and say "that is not actually how we judge this," which is a conversation worth having and one a clever phrasing never triggers.

The people who look like prompt wizards are not doing anything mystical. They are just specifying, at a level of precision most people find slightly tedious to sustain — and tedious-but-precise has been the winning move in every specification-shaped craft for about a century.

Ready to start learning?

Mochivia turns your goals into personalized, AI-powered daily lessons. Start building your path today.

Try Mochivia Free

Related Articles