How to ask AI better questions
Everyone has a folder of prompts copied from somewhere, and it never works as well as they expect. The problem is the folder. Prompting is specification, not incantation, and four habits cover almost all of it.

Kingsley Ijomah
Founder, Codehance
The question I get straight after that one is always some version of: fine, so what should I be typing?
And the person asking almost always has a folder. Prompts copied from a thread, a template someone shared in Slack, a page of magic openings that are supposed to unlock the good version of the model. They show it to me slightly hopefully, and it never works as well as they expect, and they assume they are holding it wrong.
They are not. The folder is the problem. A prompt copied from somewhere is a solution to somebody else's task, tuned to a model that has since been replaced. It is a phrasebook when what you needed was the language.
There is a real field behind this. The Prompt Report, a systematic review out of Maryland, OpenAI and Stanford, catalogued fifty-eight distinct text prompting techniques and gave them a shared vocabulary. Fifty-eight. Nobody in a delivery role is going to learn those, and nobody needs to. What you need is much smaller: a sense of what kind of asking you are doing, and the habit of describing the job properly.
That is the whole of this post. Prompting is specification. You are writing a brief for somebody quick, widely read, and completely ignorant of your situation. Every technique below is a way of writing a better brief, which is why none of it will date the way the folder did.
One caveat before I start, and I would rather say it than have you find it. The studies I cite here mostly measure maths problems and classification tasks, not writing a steering pack. I use them because the direction is consistent and worth knowing, not because someone has proven a number for your job. Treating a benchmark as a promise is exactly the mistake the previous post is about.
Four kinds of asking
Watch someone use one of these tools for an hour and you will see one mode. They type a question, they read the answer, they type a follow-up. Whatever the work is, the opening move is the same.
But the work is not the same, and four kinds cover almost all of it. Knowing which one you are in tells you what to do first, and doing the wrong first move is the most common reason a session goes nowhere.
Analysis. You already have the material and you want it understood, checked, or summarised. A contract, a set of notes, a thread, a spreadsheet of last quarter's numbers. Here the material comes first and the question comes second, and most people do it the other way round. Paste the thing, then ask. If you ask before you have given it anything to work with, it will answer from general knowledge and produce something that sounds sensible and is about nothing.
Research. You do not have the material and you want facts. This is the mode that should make you most careful, for the reason the previous post covers: when the model has nothing to recall, it produces something plausibly shaped instead. Most assistants now have a search button, and turning it on genuinely changes what you are dealing with, because the answer can cite a page you can open. Use it for this mode, and treat anything unsourced as a lead rather than a fact.
Drafting. You want text produced. An email, a summary, a first version of a section. This is the mode that lives or dies on examples and constraints, which is the last section of this post.
Brainstorming. You want options, and you want more of them than you would generate alone. Volume is the point, so ask for fifteen and throw away twelve. This is also the mode where the agreement problem bites hardest, because you are asking about something that is already yours. Ask for options against your instinct, not options that support it.
None of that is a technique. It is a question you ask yourself before typing, and it takes about two seconds.
Worth adding, because it shows how little the tips transfer: the same review that catalogued the fifty-eight techniques found that zero-shot chain of thought, the "think step by step" line that everyone has been told to append, underperformed expectations. Nothing in this field is a spell.
When to let it interview you
This is the one I would teach first if I could only teach one, and there is a finding behind it that changed how I work.
A paper called Knowing but Not Showing put ten models against a thousand questions and tested two things separately: can they tell when a request is ambiguous, and do they say so. They can. Asked directly to judge, they identified ambiguous questions with 60% to 80% accuracy. Asked the question normally, they raised a clarifying question less than 5% of the time. Between 80% and 95% of the time they simply answered, picking one interpretation silently and proceeding as though it were the only one.
Put that next to the previous post and it is the same behaviour from another angle. It always answers. What that paper adds is that the model often knows its answer is built on a guess about what you meant, and it says nothing.
So make it say something. Before it starts, tell it to ask you questions first. Something as plain as: before you write anything, ask me up to five questions that would change how you approach this. Then answer them. The questions are usually better than the ones I would have thought to specify, because they are about the gaps in what I said rather than the gaps I already knew about. The same study measured what the silence costs: accuracy on ambiguous questions ran 10 to 15 points below accuracy on unambiguous ones. My own experience of it is less about accuracy than about not discovering on the third attempt that we were building different things.
One detail from that paper is worth carrying back to the previous section. Giving the model retrieved material made it less likely to ask for clarification, not more. Turning search on does not make it more careful about what you meant.
Now the part that is usually left out. This is wrong for most tasks.
If the job is well specified, interview-first wastes turns and fills the conversation with material you then have to read past. If you want a quick fact, it is absurd. If you already know exactly what the output should look like, you are better off describing it than being questioned about it. And every one of those exchanges is now permanently in the context, which the previous post explains the cost of.
The rule I use is simple. If I could not write the acceptance criteria myself, I let it interview me. If I could, I just write them.
One question is rarely one question
Write me a project plan is not one request. It is a scope, a breakdown, an estimate, a set of dependencies, a risk list and a format. Six requests wearing one coat. Ask for all six in one go and you get a plausible document that is shallow in all six places, and shallow in a way that reads as complete.
The evidence for splitting work up is unusually strong. In the Decomposed Prompting work, breaking a task into sub-tasks scored 50.6% against 36% for chain of thought on GSM8K, and 95% against 78% on MultiArith. Those are maths benchmarks, with the caveat I gave at the start.
The pattern reported for writing tasks is the same shape: separate passes for drafting, critique and refinement tend to beat one instruction asking for all three at once. That is the exact thing people do when they ask for a good draft and expect the model to have edited it too.
Three shapes, in rising order of effort.
Split it. Ask for the breakdown. Look at it. Then ask for estimates against the breakdown you accepted. You are still doing one thing at a time, you are just no longer pretending it is one thing.
Chain it. The output of one step becomes the input of the next, with you in between. The part in between is not optional. If you pass forward something you have not read, you have automated the propagation of an error rather than saved yourself a step.
Pass over it. Draft, then critique, then revise, as three separate asks. The critique pass works far better when it does not know it is criticising its own work, so give it the text without the history where you can.
The honest counterweight: do not do this by reflex. There is a good argument, made more often in 2026 as context windows have grown, that many people chain by habit without ever checking whether one well-specified request would have done. Splitting a simple task adds turns, adds cost, and adds pile to the conversation. Split when the output came back shallow. Not before.
Show, don't adjective
One habit separates the people who get good results from the people who do not, and it has nothing to do with cleverness.
Make it more professional. Make it punchier. Make it sound like us. Those transmit almost nothing. Professional to you might mean shorter, or more hedged, or less hedged, or no exclamation marks, or the particular flatness your legal team insists on. The model has no idea which, so it produces the statistical average of professional, which is why so much of this output has the same texture.
One paragraph of something you consider professional transmits all of it at once.
This is the best supported thing in the whole post. Across the empirical work, adding examples is repeatedly identified as the single largest driver of improvement, ahead of rewording the instruction, adding a role, or any of the other things people spend their time on. The lever is showing, not describing.
Two things about that which are worth more than the headline.
The gains plateau fast. The reported pattern is that most of the benefit arrives with the first two or three examples, and piling on more does not keep helping. So this is not a research project. It is finding two things you already have that are good and pasting them in.
And on newer reasoning models, examples can make output worse. Which is a decent reminder that none of this is a law, and that the way to find out is to try it without them first and add them when the result disappoints.
Beyond examples, three more things belong in a brief, and they are the ones people leave out.
Constraints. Length, audience, what to avoid, what must appear. What not to do is worth as much as what to do, and almost nobody writes it down. If you never want it to open with "In today's fast paced world", say so once and it will stop.
The shape of the output. Say whether you want prose, a table, five bullets, or a one-page memo with headings. If you do not, you will get whatever shape the training data most often took, and then spend your time reformatting.
Labels on the parts of your message. Once a prompt gets long, it contains different kinds of thing: your instruction, the source material, an example, some background. Run them together and the model has to guess which is which, and it guesses wrong in a specific way, treating your material as an instruction or your instruction as something to summarise. Put a plain label above each part. INSTRUCTION, then SOURCE, then EXAMPLE. That is all it takes. If you have seen people wrapping sections in angle brackets like tags, that is the same idea in tidier clothing, and it is worth doing once your prompts get long enough to need it.
None of this is a trick, and that is the point. A brief that specifies the job, shows an example, states the constraints and marks its own parts is just a good brief. It would work on a contractor. It works here for the same reason.
Which leaves the thing none of it solves. A better brief gets you a better draft. It does not tell you whether the draft is right.
See it for yourself
The argument in this post is a comparison, so run the comparison.
Find a prompt you actually used this week. Something real, from your own work, not an example. Open a new conversation and run it again exactly as you wrote it. Keep the output.
Now rewrite it as a brief. Decide which of the four kinds of asking it is. Put the material first if you have any. Add one example of what good looks like, state two or three constraints, and say what shape the output should take. Open another new conversation, because you want the brief to be the only thing that changed, and run it.
Put the two outputs side by side.
The gap between them is not the model getting lucky. It is the difference between a question and a brief, and once you have seen it in your own work you will not go back to the folder.
Sources
Linked so you can check them, and so you can see where the evidence is thinner than the claim.
- Schulhoff et al., The Prompt Report: A Systematic Survey of Prompt Engineering Techniques, 2024. The taxonomy of 58 text prompting techniques, and the finding that zero-shot chain of thought underperforms.
- Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions, 2026. Ten models on a thousand AmbigQA questions: 60 to 80% recognition when asked to judge, under 5% clarification when simply asked, and the finding that retrieved context suppresses clarification.
- Khot et al., Decomposed Prompting: A Modular Approach for Solving Complex Tasks, 2022. The GSM8K and MultiArith figures.
Two claims here are weaker than the rest and I would rather say so than let them pass as settled. That examples are the largest single lever is consistently reported, but I have not found one clean headline number I am willing to quote for it, and the same goes for where the gains plateau. Treat both as the direction the evidence points rather than as measurements.
Want to build this, not just read about it?
The free 3-day Codehance challenge teaches the architecture-first method hands-on. No coding background needed.
Start the free challenge