Search Visibility

How to test whether the assistants name you, every month

A client asks whether the assistants mention their company, and someone opens ChatGPT, types the company name, gets a flattering paragraph, and everyone relaxes. That test proves nothing. You asked a system that had your name in the prompt whether your name existed.

The question worth answering is different, and harder: when a buyer describes their problem without naming anyone, does your company come up? That is a measurable thing, and it changes month to month, and almost nobody is tracking it.

This is the test I run, once a month, in about forty minutes. It requires no tools you have to buy and no access you do not already have.

First, be clear about what you are measuring

Being ranked and being recommended are separate jobs — I have argued that at length in ranking and being recommended are not the same job, and this post assumes it rather than repeating it. What follows is the procedure, not the theory.

One thing worth settling before you start, because it saves a great deal of wasted effort. Google’s own documentation on its AI features is explicit that there are no additional requirements to appear in AI Overviews or AI Mode, no special optimisations, and no new machine-readable files or markup to add. Anyone selling you a package to fix that is selling you something that does not exist. What you can do is be findable, be clear, and be worth quoting — and then measure whether it worked.

The prompt set, written once

You need a fixed list of questions that a real buyer might ask, none of which contains your name. Ten to fifteen is enough. Write them once and never change them, because the point is comparison over time, and a prompt you rewrite is a measurement you have thrown away.

Build them from four angles:

  • The problem, stated plainly. “Our website takes three seconds to load on mobile and we do not know why.” No industry, no location.
  • The problem plus the market. The same question with “in the UAE” or “for a company in Dubai” appended.
  • The category. “Who can help a mid-sized company plan an AI project” — the phrasing a buyer uses when they do not yet know what to search for.
  • The comparison. “What should I ask before hiring a technology consultant” — where you are hoping to be cited as the source of the answer rather than named as a supplier.

Write half in English and half in Arabic. They will not return the same results, and the gap between them is one of the more useful things this exercise tells you.

Run them clean, every time

The single most common mistake is running the test in an account that knows you. A logged-in assistant with memory of your previous conversations is not answering the question a stranger asked.

So: a signed-out session, or a private window, with memory and personalisation off. No account history, no earlier messages in the same thread, one question per conversation. If the tool offers a memory setting, confirm it is disabled rather than assuming.

Bing’s own webmaster guidelines are worth reading once for the same reason: the fundamentals it asks for — crawlable pages, clear titles, content that answers the query — are the fundamentals every assistant inherits from the index underneath it. There is no separate track.

Run the same set across whichever assistants your buyers actually use. In this market that realistically means ChatGPT, Google’s AI mode, Copilot and Gemini, plus Perplexity if your buyers are technical. Do not add a fifth for completeness. Four consistently beats six inconsistently.

Four cards setting out the monthly assistant visibility test: twelve fixed prompts, signed-out runs, four fields recorded, and reading the pages cited instead of yours
Forty minutes a month, and no tools to buy.

What to write down

For each prompt, on each assistant, record four things and nothing else. The discipline of a small fixed schema is what makes the numbers comparable next month.

  1. Named or not. Did your company appear at all. A yes or no, not an impression.
  2. Position. First mention, mentioned among others, or mentioned only after a follow-up question.
  3. Cited or not. Was a link to your site attached to the answer. This is the one that separates being known from being used, and the two are not the same.
  4. Described correctly. If it named you, did it get what you do right. A wrong description is worse than no mention, and it is the finding that most often needs acting on.

Google publishes which result types can appear with rich presentation in its search gallery, and it is a useful reality check when someone claims a format will get you quoted: if it is not in there, it is a claim rather than a feature.

Keep it in a spreadsheet with one row per prompt per assistant per month. Twelve prompts across four assistants is forty-eight rows a month, which is twenty minutes of typing and the only honest trend line you will get.

The second half: who it names instead of you

The part most people skip, and the part that produces the actions.

Whenever an assistant names someone else, write down who, and then go and look at the page it is drawing on. Do not skim it for quality — look for structure. Does it answer the exact question in its first paragraph. Are the specifics in the prose or hidden in a PDF. Is there a passage that could be lifted whole and still make sense.

It is also worth knowing how the answer was assembled. Several assistants now search the live web mid-answer rather than relying on training data — OpenAI documents its web search tool doing exactly that. When an answer carries links, you are looking at something closer to a search result than a recollection, and that is the version you can influence.

That last quality is what actually gets quoted. A page that requires the reader to have read the preceding section will not be extracted, however good it is. This is also why the FAQ blocks on every page of this site are real markup rather than decoration — a self-contained question and a self-contained answer are the smallest quotable unit there is.

What to do with what you find

The results sort into four situations, and each has a different fix.

Not named, and the winner has a page you do not. The most common and the easiest. Write that page. Not a longer version of theirs — the version that answers the question in its opening lines.

Not named, but the answer cites nobody at all. The assistant is answering from general knowledge, which means the question is not yet a citation opportunity for anyone. Deprioritise it. This is the category people waste the most effort on.

Named but described wrongly. Your own pages are ambiguous about what you do. Fix the descriptions on your site first, and make sure the same description appears consistently — the entity properties in schema.org’s Organization vocabulary exist precisely so a machine can connect the same organisation across places, and inconsistency is what breaks that connection.

Named and cited. Note which page earned it and what shape that page has, then write more pages with the same shape. This is the only reliable signal you will get about what works for your specific domain.

The bilingual finding nobody expects

Run the Arabic half properly and you will usually find one of two things.

Either the Arabic answers cite almost nothing local, in which case the space is open and a genuinely Arabic page — written, not translated — has an unusually cheap route to being the cited source. Or they cite a small, fixed set of the same sites, in which case you are looking at exactly what to compete with.

Run the prompts about connecting systems too, if that is your market — the questions in the four questions before you connect a model to a live system are exactly the shape a buyer types when they are looking for someone to ask.

Either result is more actionable than anything the English half tells you, because the English half is crowded and the Arabic half generally is not. That is the same argument as Arabic-second always shows, arriving from the measurement side rather than the editorial one.

How long before it moves

Longer than anyone wants. A new page that deserves to be cited will not be, for weeks, because it has to be crawled, indexed, and then judged useful enough to draw on. Two to three months before a change in the numbers means anything is realistic, and one month of data means nothing at all.

Which is the actual argument for making this a monthly habit rather than a project. The value is not in this month’s reading — it is in having twelve of them when someone asks whether any of this is working. That is the same discipline as every other measurement worth keeping, and it is part of what I do as an SEO, AEO and GEO consultant and inside fractional CTO work, where the question is usually whether the marketing spend is buying anything at all.

Running this test every month, and acting on what it returns, is work my company takes on: AI search visibility at Tothiq.

Frequently asked questions

Can I automate this instead of doing it by hand?

Partly, and the tools that do it are improving. Two cautions. Automated runs through an API often behave differently from the consumer product a buyer actually uses, so you may be measuring a system nobody is asking. And the second half — reading the pages that got cited instead of yours — is the part that produces the actions, and it does not automate. Run it by hand for six months first, so you know what the numbers mean before you trust a tool to produce them.

Does structured data make an assistant more likely to quote us?

Not on its own, and Google’s documentation says so directly — there are no additional markup requirements for its AI features. Structured data does something narrower and still worth having: it removes ambiguity about who you are and what a page is, which helps a machine connect mentions of you across different sources. Treat it as hygiene, not as a lever.

Our name never comes up. Where do we start?

With the prompts where the assistant already cites somebody. Those are proven citation opportunities — the question is settled, only the source is in play. Ignore the prompts where nothing is cited, however important they feel, until the first group is working. Starting with the hardest question is why most of these efforts produce nothing in the first year.

How is this different from checking our search rankings?

A ranking tells you where a page sits in a list a person still has to choose from. This tells you whether a machine, answering on your buyer’s behalf, considered you worth mentioning at all. You can rank first and never be named, and you can be named without ranking anywhere. They need separate measurement because they are separate outcomes.

Is forty minutes a month really enough?

For twelve prompts across four assistants, yes, once the prompt set is written and you are not re-deciding what to ask each time. The setup — writing the questions and building the sheet — takes an afternoon once. If it is taking three hours every month you have too many prompts, and you will stop doing it by month four, which costs more than the extra coverage was worth.

Have a project, problem or idea?

Let's discuss what you're trying to build, improve or grow — and whether I can help.

Discuss Your Project