jonas.simonsson
essayAI and data-driven operations

Fluent is not the same as correct

Jonas Simonsson/May 2026/5 min

Someone handed me a recommendation a while ago that was better written than most things people spend a week on. Clear argument. Clean structure. The kind of document you read and feel your shoulders drop a little, because finally, someone has done the thinking. At the centre of it sat a single number, the kind that decides real spend and is hard to walk back once it lands in a slide in front of the people who sign things.

The number was wrong. Not rounded wrong, not optimistic. Invented. It had the shape of a figure that came from somewhere, and it came from nowhere. The part that stayed with me is that nothing in how it was presented gave that away. The confidence was identical to the confidence a real number would have carried.

Fluency used to be a signal

For most of my career, a precise figure was a proxy for effort. If someone wrote a specific number into a document, you assumed a person had sat with the inputs, made assumptions they could name, and could defend the result if you pushed. The precision was a tell. It meant work had happened upstream of the sentence.

That proxy is broken now. A figure can be specific, well formatted, and completely fabricated, all at once, and produced in seconds. The cues most of us evolved to trust are fluency, structure, a confident tone. They now fire just as readily on output that has nothing underneath it. The surface and the substance have come apart, and we are still reading the surface.

Confidence and correctness are different properties

This is the whole thing, so it is worth saying plainly. Confidence is a property of the writing. Correctness is a property of the world. They are not the same measurement, and they are not even taken with the same instrument.

Confidence is a property of the writing. Correctness is a property of the world. The model is very good at one of them.

A language model optimises for the first. It is built to produce text that reads as though a competent person wrote it, and it is extremely good at that. Whether the claim matches reality is a separate question the fluency cannot answer, and the better the writing gets, the easier it is to forget there were ever two questions.

Being directionally right is the dangerous part

The uncomfortable version of this is that the recommendation is often not even wrong. It points the right way. The strategy is sound, the framing is reasonable, the conclusion is one you might have reached yourself. That is worse, not better, because being right about the direction earns a trust the specific numbers have done nothing to deserve. You nod along with the argument, and the fabricated figure rides in on the back of it.

The job did not disappear. It moved.

Here is where I have landed. The model can generate the analysis. It cannot be accountable for it. Accountability is a human function, and it needs something the model does not have, the ability to look at a claim and independently check the foundation underneath it. To know a number is wrong because you understand, at least roughly, how that kind of number is built and what it should look like when it is right.

That is the work. It did not get automated away. It got more important, because there is now far more confident output to check and the same number of hours to check it in.

Range stopped being the lesser profile

Look at what the checking actually requires. To catch a fabricated number you have to know how that kind of number is built, what it should look like when it is right, and how it is supposed to behave inside the system it claims to belong to. A specialist has that understanding inside one lane and nowhere else. Confident fabrication does not respect lanes. It lands across infrastructure, data, cloud, security, operations, and the management decisions sitting on top of all of them, anywhere a question was asked and a fluent answer was accepted.

So the scarce competence is not depth in one place. It is the ability to judge across places, to carry enough of how each kind of claim behaves that you notice when one does not fit, in more than one domain at once. For years that profile was treated as the lesser one. Depth was the prestige path, and breadth was what you called someone politely when they had not committed. The tooling quietly inverted that. Producing the analysis is no longer the work, because a model does it in seconds. Checking it is the work, and checking it across a wide range is rare.

I should be honest that this is a convenient conclusion for someone in my position to reach, which is exactly why it has to meet the same test as everything else here. The claim is not that breadth feels valuable. It is narrower and more checkable than that. Accountability has a mechanism, the mechanism is independent verification, and verification has a reach limited by what the person doing it actually understands. If that holds, breadth is not a trait to be flattered. It is simply how far the checking can extend.

The test

Take the last AI-assisted recommendation you acted on. Find its single most important number. Now say where it came from. Not where it could have come from, not where it plausibly came from. The actual input that produced it.

If you can trace it to ground, you had a recommendation, and the tool did you a real service. If you cannot, you had a very well written guess, and what you acted on was how sure it sounded.

AI did not remove the need to check the numbers. It made checking them the job. The draft arrives at speed now. The reading of it is still yours.

Written by Jonas Simonsson.·more writing·connect on linkedin