How Perplexity, ChatGPT And Gemini Pick Their Sources
A Numeric Name Is an Entity Problem Names beginning with digits behave differently across the web than names beginning with letters. They get written several ways, they sort strangely in directories, and they collide with unrelated numeric strings in ways that letter based names do not.
This variability is the main practical trap. Testing without web access and concluding you are invisible measures the training corpus rather than current retrieval, and the two can disagree sharply. Record which mode you used with every run.
Blocking these is therefore not one decision. Turning away a training crawler is a defensible editorial position. Turning away the agent that fetches pages at answer time removes you from answers entirely, and the two are frequently confused.
This entire area usually amounts to a day of work. It is routinely the difference between a brand that appears in answers and one that does not, and it is worth doing before anybody writes a single word of new content. ai visibility agency
How to Test Rather Than Trust Everything above is a starting hypothesis. Run twenty prompts in your own category across all three, from signed out sessions, recording the mode and the date, and count the cited domains for each.
The complication is that AI systems use several distinct agents for different purposes. One may crawl for training corpora, another may fetch pages live when composing an answer, and a search provider's traditional crawler may feed both search results and an AI summary.
One local specific worth checking is how your opening hours and availability are stated across every listing. These are among the details most frequently quoted in local recommendations and among the most likely to be wrong, because they change seasonally and get updated in one place. An assistant confidently telling somebody you are closed is a lost job that leaves no trace in any report.
The Blocks Nobody Chose Most blocking discovered during audits was never a decision. A disallow copied from a template. A staging rule that survived a migration. A security plugin with an aggressive default. A content delivery network setting labelled bot protection with a switch nobody has looked at since launch.
Track three things over time: how often you are named, which sources get cited when you are, and which competitors appear alongside you. Movement in the second of those usually predicts movement in the first.
What Transfers to an Ordinary Business Three things, and they are the three that most small operators skip. Check that you are readable before assuming you have a content problem, since on a small site an access failure is total rather than partial.
Put someone's name against this. Crawler rules sit between marketing, development and whoever administers the content delivery network, which in most organisations means nobody checks them. The failures documented here are not difficult to find, they are simply nobody's job, and a quarterly review taking half an hour prevents the most complete form of invisibility available.
Broad sites are forgiving. A blocked section or a badly rendered template still leaves a hundred other pages describing the organisation. A small site with five pages has no such buffer, which makes the mechanical checks disproportionately important.
It is also worth recording the reason for every rule you keep. A disallow line with no explanation gets preserved indefinitely through migrations and redesigns because nobody dares remove something they do not understand. A one line comment saying who added it and why turns a permanent mystery into a decision that can be revisited.
And pick a narrow enough definition of what you do that the existing coverage is thin. Competing to be the best documented answer to a specific question is a solvable problem. Competing for a broad category against everyone is not, and the small operators who do well here are almost always the ones who narrowed first. ai visibility agency
What robots.txt Controls It is a request, honoured by mainstream crawlers, that certain user agents avoid certain paths. It has no enforcement behind it and it does not secure anything, but the major providers respect it.
You will find your own category's pattern, which frequently contradicts the general one. Some industries are dominated by a single trade directory. Others are dominated by one forum. That specific finding is worth more than any general description of how these systems behave.
The Shared Architecture All three now commonly retrieve live sources rather than answering purely from training. Your question becomes one or more searches, a set of pages is fetched and read, and the answer is composed from what was read.
What a Local Business Should Do This Month Run five prompts asking for a business like yours in your town, from a signed out session, and record who gets named and what gets cited. Then fix every listing on the sources that appeared, starting with the phone number and address.