A vocabulary invented faster than the practice underneath it.
A category appeared over the last two years selling optimisation for AI answer engines. It has its own acronyms, its own conference talks, and a growing number of agencies who have repositioned around it entirely.
Some of it is real. Most of it is a rename, and a meaningful portion is unfalsifiable by construction.
Worth being precise about which is which, because the underlying shift is genuine — and that is exactly what makes the category attractive to people selling certainty they do not have.
Renamed technical SEO. Clean semantics, sensible heading structure, fast pages, structured data, sane internal linking. This is legitimate and it does help. It is also the same list from a decade ago with a new cover. Nothing is wrong with the work; something is wrong with charging a novelty premium for it.
Unfalsifiable claims about model behaviour. This is the problematic tier. Assertions about how a particular model weighs sources, what it “prefers,” how to be favoured in its answers. These systems are closed, non-deterministic, change without notice, and personalise. A claim you cannot test is not expertise. It is a posture that happens to be unrefutable, which is not the same thing.
Genuine retrieval work. A real and much smaller category: making content extractable in chunks, declaring entities so they resolve consistently, answering a question directly rather than burying the answer under context, keeping things current. Boring, mechanical, and it works.
The clearest signal that a category is running ahead of its evidence is how its numbers behave.
The projections in circulation — traffic collapse percentages, conversion multiples, adoption curves — are quoted with confidence and cited everywhere. Follow enough of them back and a striking share converge on a small number of vendor-published reports with undisclosed methodology, or on nothing citable at all.
A statistic whose origin you cannot reach is not evidence. It is a number that has been repeated until it sounds like one.
The irony is that these numbers are unnecessary. The mechanism argues for itself: if a buyer asks a system for a shortlist, and that system builds its answer from sources it can parse, then being parseable determines whether you are eligible. You do not need a projected percentage to justify that. You need the buyer’s behaviour to have changed, which it observably has.
An honest account has to include the boundary.
Knowable and controllable: whether your entity is declared in structured data; whether your headings answer real questions; whether each section stands alone when extracted; whether your claims are stated as text rather than baked into images; whether the content is current; whether a machine can reach the page at all.
Not knowable: the weighting inside any given model, why one source was cited over another in a specific answer, and what any of it will look like in a year.
The correct posture is to do the first list thoroughly and decline to make claims about the second. That is less impressive in a sales meeting and considerably more defensible eighteen months in.
This argument cuts toward anyone selling this work, including me.
So, plainly: the visibility instrument on this site scores structural legibility — entity declaration, credential legibility, service legibility, heading structure, language separation. Every one of those is checkable by the person running it, on their own site, without me. That is deliberate. It is the only version of this work I can defend, because it is the only version where the client can verify the finding themselves.
What I cannot tell you is that a particular score produces a particular placement in a particular model. Nobody can, and the ones who say otherwise are describing a system they do not have access to.
Next time someone pitches AI optimisation, ask one question:
Which part of this can I verify myself, and which part am I taking on trust?
A good answer separates the two without prompting and is comfortable being specific about the second. A weak answer treats the question as scepticism to be managed. That reaction tells you which tier of the category you are dealing with, and it tells you before you have paid for it.
Three instruments on this site score exactly this — free, ungated, no email required to see a result.
Run the diagnostics →