When not to use AI
Everybody asks what they should use AI for. Almost nobody asks what they should not, because nobody selling it is incentivised to answer. The work it is worst at, why you cannot feel the difference, and what must never go in at all.

Kingsley Ijomah
Founder, Codehance
Almost every conversation I have about these tools starts in the same place: what should I be using this for? It is the wrong first question, and it is the one everybody asks.
The right one is narrower and much less popular. What should I not be using this for?
Nobody is incentivised to answer it. The people selling courses are not going to tell you the tool is wrong for a third of your work. The vendors certainly are not. And the person in your organisation who has gone all in on it has staked something on the answer being nothing. So the question goes unasked, people find the boundary by crossing it, and the crossing is usually expensive and quiet.
This is the answer I give when somebody asks me privately, and none of it is an argument against the tool. It is the boundary that makes the rest of it trustworthy.
The work it is worst at
Start with the shape of the thing. It predicts text that plausibly follows other text. That tells you immediately where it will be strong and where it will not, and the boundary is not about difficulty.
It is good where a good answer looks like a well-formed piece of writing about something widely discussed. Restructuring your rambling notes. Explaining a concept you half know. Turning bullets into prose. Giving you fifteen options when you would have thought of four. In those, plausible and useful are close to the same thing.
It is weakest where the value of the output depends on something the text cannot contain.
That covers more of your job than it sounds like. It cannot know what was actually agreed in the room, only what the notes say was agreed. It cannot know that the client is quietly furious, that the estimate is political, that the risk everyone is worried about is the one nobody has written down. Ask it to draft a status update and it will produce a good one about the project as described to it, which is not the project.
There is a cost to getting this wrong, and it has been measured. A BetterUp Labs and Stanford study of 1,150 US workers named the output that results workslop: material that has the shape of finished work without the substance to move anything forward. 41% of them had received some in the previous month, each instance taking an average of one hour fifty-six minutes to sort out.
Notice where that time went. Not to the person who generated it. The cost of handing over work that should not have been handed over lands on whoever receives it, which is why it is so easy to keep doing. The same study found 42% of recipients thought less of the sender's trustworthiness afterwards. You can spend your reputation this way without ever seeing the bill.
You are a poor judge of whether it helped
This is the finding that changed how I think about the whole question, and I would put it in front of anyone rolling these tools out to a team.
METR ran a randomised controlled trial with sixteen experienced open-source developers across 246 real tasks in repositories they had worked in for years. Before starting, they expected AI to make them about 24% faster. Afterwards, they reported it had made them about 20% faster.
They were 19% slower.
Two caveats, and I would rather give them than have you find them. Sixteen developers is a small study in one setting. And METR themselves now label the result historical, saying it no longer reflects the current impact of AI models and pointing at newer numbers. If someone quotes the 19% at you as today's truth, they have not read the source.
But the part that transfers is not the slowdown. It is the gap. These were expert practitioners, working in code they knew intimately, and they were out by roughly forty percentage points about their own recent experience. Nothing about the newer models makes that gap smaller. If anything a smoother tool makes it worse, because the feeling of progress is exactly what is being mismeasured.
Which means the sentence "it definitely speeds me up" is not evidence. It is the thing the study found to be unreliable. If you want to know whether this is helping on a particular kind of task, you have to measure something, even crudely, rather than consult your impression of it.
What must never go in
Everything above is a judgement call. This part is not.
Cyberhaven's 2026 telemetry across its customer base found that 39.7% of all interactions with AI tools involved sensitive data, counting prompts, pastes and file uploads, and that the average employee puts sensitive data into one of these tools about once every three days.
The number I would actually put in front of a leadership team is the one about accounts. In that same data, 58.2% of Claude usage and 60.9% of Perplexity usage happened through personal accounts rather than company-managed ones. For ChatGPT it was 32.3%. Whatever your organisation has agreed about data handling, on those numbers a large share of the traffic is not passing through it.
Treat this as vendor telemetry rather than peer-reviewed research, because that is what it is, and the report does not publish its sample size. The direction is not in serious dispute though, and it matches what I see.
The working rule is simpler than any policy document. If it would be a problem for this to appear in a screenshot on somebody else's screen, it does not go in. That covers client contracts and anything under an NDA, personal data about identifiable people, unreleased financials, security details and credentials, and anything about a colleague, particularly performance and health. A performance review is the one I see most, and it is the one that would be hardest to explain.
Two things people get wrong here. A conversation feels private because it looks like a chat, but you are sending it to a company, and what happens to it depends on which account you are signed in to and what that account's terms say. And a personal account is not a discreet workaround for a blocked tool. It is the same disclosure with the company's controls removed.
If your organisation has an approved tool, use that one, even where it is worse. That is the entire point of it.
Whose name is on it
The last one is not about the machine at all.
When something you produced this way turns out to be wrong, the sentence you will want to say is that the AI got it wrong. It is true, it is useless, and it will not survive the room. You sent it. Your name is on it. The tool has no professional standing to lose and you do.
That is the accountability question and it has a clean answer. Nothing changes about who is responsible, so nothing should change about what you are willing to put your name to. The test is the same one you would apply to work from a contractor you had never met: would you stand behind it if asked to explain it line by line?
Disclosure is genuinely less settled, and I would not pretend otherwise. Norms differ by organisation and by field, and they are moving. What I would hold to is narrower: say so where the reader's judgement of the work depends on how it was made. Nobody needs telling that you used a spellchecker. If someone is relying on your independent analysis, or on the fact that a human read every line, and that is not what happened, they need to know.
And if you find you would rather nobody knew how a piece of work was made, that reluctance is information. It usually means you already know it was the wrong thing to hand over.
See it for yourself
The judgement in this post is hard to develop in the abstract, so make it concrete against your own week.
Write down the last ten things you used one of these tools for. Actual tasks, from your calendar or your sent folder, not the ones you think are representative.
Now sort them into three piles. The first is work where the output only had to be well formed to be useful: rephrasing, summarising, drafting from material you supplied, generating options. The second is work whose value depended on something the model could not have known: the politics, the history, the unwritten context. The third is anything you would not have wanted read aloud, because of what was in it.
Most people find the first pile is where the real wins are and the second pile is where the frustration has been. The third pile is the one worth sitting with, because it tends to be small, specific, and already sent.
Sources
Linked so you can check them, including the one that comes with a warning label.
- Niederhoffer, Kellerman, Lee, Liebscher, Rapuano and Hancock, AI-Generated "Workslop" Is Destroying Productivity, BetterUp Labs and Stanford Social Media Lab, Harvard Business Review, 2025. 1,150 US workers, the 41% figure, the one hour fifty-six minutes, and the trust effects.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 2025. Sixteen developers, 246 tasks, 19% slower against an expectation of 24% faster and a self-report of 20% faster. METR label this result historical and say it no longer reflects the current impact of AI models. The perception gap is what I have used it for, not the slowdown.
- Cyberhaven, 2026 AI Adoption and Risk Report. The 39.7% figure and the personal-account shares. This is vendor telemetry from one company's customer base and the published summary does not give a sample size, so treat it as an indication of direction rather than a measurement.
Want to build this, not just read about it?
The free 3-day Codehance challenge teaches the architecture-first method hands-on. No coding background needed.
Start the free challenge