Insights
    ChatGPT
    Prompt Testing
    AI Visibility

    The Questions Every Brand Should Test in ChatGPT

    Avisible TeamJune 16, 20266 min read

    Most companies have never asked an AI system what it says about them. The few that have usually asked the wrong question — typing their own brand name and reading the answer with relief.

    That test proves nothing. You named yourself in the prompt, so of course you appear. The questions that reveal your real position are the ones your buyers ask before they know you exist, and building a set of them is most of the work in any honest AI visibility measurement.

    Why question design is the whole exercise

    AI systems do not behave like search engines. Which brands get named depends heavily on wording, framing, category association and how much corroborated authority a company has in the specific area being asked about. Change best X for enterprise to affordable X and the recommended set can turn over completely.

    This makes the question set a design problem rather than a list-writing exercise. A badly designed set produces a flattering, useless picture. A good one maps where in the buying journey you disappear — which is the information you can act on.

    The four types worth covering

    We organize question sets across four stages, because visibility failures cluster by stage rather than spreading evenly.

    Discovery — the buyer does not know the category yet

    Broad, informational prompts: what is GEO?, how do brands appear in ChatGPT?, how do I improve AI discoverability? Nobody is close to buying here, but these answers set the vocabulary and the mental model — and often the initial shortlist. Brands that skip explanatory content lose this stage without noticing, because it produces no leads to miss.

    Awareness — the buyer is choosing a category, not a vendor

    Category-level comparison prompts: best AI visibility tools, top GEO agencies, alternatives to a traditional SEO agency. This is where share of answer is won or lost, and where most brands discover a competitor they had not been tracking. If you appear nowhere in this stage, later-stage visibility rarely rescues you.

    Reputation — the buyer has your name and is checking it

    Brand-specific prompts: what does this company do?, is it reputable?, how does it compare to a named competitor? The risk here is not absence but inaccuracy — outdated pricing, a service you no longer offer, a confident description of the wrong company. Being described wrongly is more damaging than not being described at all, because it is repeated with authority.

    Conversion — the buyer is ready to act

    Decision-stage prompts: who should I contact for X?, what does an audit cost?, how do I get started? Low volume, high intent. Brands that publish prices and process tend to do well here simply because most competitors hide both, leaving the model with nothing concrete to quote.

    What makes the set reliable

    Three properties separate a set you can trust from a list of prompts:

    1. Fixed. The same questions every time, so results are comparable across runs. Changing the questions between measurements makes the trend meaningless.
    2. Repeated. Each question asked several times, because AI answers vary between runs. One answer is an anecdote; three is a rate.
    3. Multi-system. Run across more than one model, because they disagree with each other about who is worth naming.

    Scale follows from those properties rather than being a goal in itself. A hundred questions run three times across three systems produces nine hundred answers — not because volume is impressive, but because that is what it takes for the numbers to be stable enough to compare.

    What brands usually find

    The results are rarely what teams expect. The recurring findings:

    1. A competitor dominates category questions the brand assumed it owned.
    2. The brand is visible at conversion stage and absent at discovery, so it only appears to buyers who already knew the name.
    3. AI describes the company accurately but blandly, with none of the differentiation the marketing team spent a year on.
    4. The sources shaping the answer are third-party pages nobody internally had looked at.

    Each of those points at a different fix, which is precisely why the measurement has to come before the strategy. The underlying causes are usually structural, and the metrics that read them are worth understanding before commissioning any work.

    How many questions is enough?

    Fewer than most vendors imply, more than most teams first try.

    Thirty well-chosen questions across the four stages will tell you whether you have a visibility problem and roughly where it sits. That is enough to make a decision about whether to invest further, and it is achievable manually if someone is willing to spend an afternoon on it.

    Around a hundred is where the numbers become stable enough to compare over time and to break down by stage without the segments getting too small to mean anything. Below that, a single unusual answer can swing a stage-level rate noticeably.

    Two hundred and upwards buys resolution rather than reliability: multiple markets, multiple languages, per-competitor breakdowns, and enough coverage to say something confident about a sub-category rather than the category as a whole. It is the right scale when the decision it informs is large, and overkill when it is not.

    The number that is always wrong is one. A single question asked once, in one system, from a logged-in account, is the test most companies have actually run — and it is the one that produces false comfort most reliably.

    What a question set looks like in practice

    To make this concrete, a hospitality brand's set might allocate roughly half its questions to awareness, a third to reputation and the rest to conversion — because that is where a hotel's discovery actually happens.

    Awareness questions would name the city and the traveller rather than the hotel: where should I stay in Helsinki for a business trip?, which hotels near the airport have good meeting facilities? Reputation questions would test the brand directly for accuracy: is this hotel good for families?, how does it compare with a named competitor? Conversion questions would test the last step: what is the cancellation policy?, how do I book a group rate?

    Notice that the brand name appears in only one of the three groups. That imbalance is deliberate and it is the part most self-run tests get wrong — a set weighted toward brand questions measures how well AI knows a company that has already been named, which is not the thing that determines whether anyone finds you.

    Where the questions come from

    The best source is not a keyword tool. It is your own sales calls, support tickets and lost-deal notes — the questions buyers actually asked, in their words. Add your competitors' comparison pages, the questions your category's forums keep repeating, and the ones your own site fails to answer.

    At Avisible this becomes the Question Matrix: a structured set covering all four stages, agreed before anything is measured, so the result is a picture of your market rather than a picture of whichever prompts someone thought of on the day. If you want one built and run for you, that is what the Visibility Snapshot does.