To check whether AI recommends your business, you need to ask five engines the same buyer questions in clean sessions and score the results. A single question in one engine tells you almost nothing. AI recommendations change between sessions, vary by engine, and shift when models update. A structured audit across ChatGPT, Gemini, Perplexity, Claude, and Grok gives you a number you can track and a baseline you can measure progress against.
Most businesses haven't done this, even though the stakes are already measurable. The 6sense 2025 B2B Buyer Experience Report found that 94% of buying groups ranked their preferred vendors before their first sales conversation, and 94% of those buyers used large language models to organize that research. If you don't know what those models say about you, you are flying blind through the phase where the shortlist gets made.
What follows is the actual procedure, written so you can run it yourself in under an hour.
Why One Engine Is Not Enough
Each AI engine draws from different sources, weights evidence differently, and updates on different schedules. Research from MIT and Stanford has shown that large language models vary significantly in which sources they retrieve and how they weight that evidence when generating recommendations — a finding that holds even when models are given identical prompts. As of September 2026:
- ChatGPT retrieves from Bing's index and its training data, with optional web browsing.
- Gemini pulls from Google's search index, Business Profiles, and Knowledge Graph.
- Perplexity runs real-time web searches and shows its cited sources.
- Claude can search the web when enabled, and draws from training data.
- Grok searches X alongside web sources.
These capabilities change with model updates, so note the date of your audit.
A business that appears in ChatGPT's answer may be absent from Gemini's. A business that Perplexity cites with a source link may go unmentioned by Claude. The differences aren't random. They reflect which evidence each engine can access and how it weighs that evidence against alternatives.
Checking one engine is like checking one review site and assuming it represents all of them. Your buyers use different engines. Your audit should too.
For a detailed look at how ChatGPT specifically makes its decisions, read how ChatGPT decides who to recommend.
The Ten-Question Protocol
Start with ten questions that your actual buyers ask when they are looking for a business like yours. Not ten questions about your industry. Ten questions a person in the middle of a hiring or purchasing decision would type into an AI assistant.
The format matters. Frame them the way a buyer would: "Who should I hire for [service] in [location]?" or "What should I look for in a [provider type]?" or "Is [specific concern] something I need to worry about when choosing a [provider]?"
Here is a starter list you can adapt to any professional service:
- Who is the best [your service] in [your city]?
- Who should I hire for [your service]?
- What should I look for when choosing a [your provider type]?
- Is [your company name] good at [your service]?
- Who are the top [your provider type] companies?
- [Your company] vs [competitor]: which is better for [service]?
- What does [your service] cost?
- How do I know if a [your provider type] is legitimate?
- Can you recommend a [your provider type] that specializes in [your niche]?
- What questions should I ask before hiring a [your provider type]?
Write your ten questions down before you start. You will use the same list across all five engines and again next month.
How to Run a Clean Session
AI engines personalize responses based on conversation history and, in some cases, your account activity. To get a baseline that represents what a new buyer would see, you need clean conditions.
For ChatGPT: open a new conversation. Do not continue a previous thread. If you are logged in, your history may influence responses. For the most neutral result, use a private browser window or log out.
For Gemini: open a new conversation at gemini.google.com. The same rules apply: fresh thread, no continuation from a previous session.
For Perplexity: open perplexity.ai in a new tab. Perplexity shows its sources, which makes it the easiest engine to audit. Note not just whether you appear in the answer but whether you appear in the cited sources panel.
For Claude: open claude.ai and start a new conversation. Claude can search the web when the feature is enabled, but it also draws from training data. Try your questions with web search on to see what a buyer with that configuration would get.
For Grok: access through X or grok.com. Grok's answers incorporate social media alongside web sources, so your X presence may influence results here more than on other engines.
Ask each of your ten questions in each engine. That is fifty prompts. It takes forty-five minutes to an hour, depending on how quickly you read and score the responses.
How to Score What You Find
For each response, record one of four outcomes:
Recommended. Your business is named as a specific recommendation. The engine says something like "I would suggest [your company]" or "a strong option is [your company]." That's the highest-value outcome.
Mentioned. Your business appears in the answer but is not singled out as a recommendation. You are listed alongside competitors or referenced as one of several options. Valuable, but less decisive than a recommendation.
Cited. Your content is used as a source but your business is not named as a recommendation. Perplexity shows this most clearly with its source panel. Your article or page provided the evidence, but the engine recommended someone else or no one.
Absent. Your business does not appear anywhere in the response. Not named, not cited, not mentioned. That's the outcome that costs you buyers you never know about.
Record these in a simple grid: ten questions across the top, five engines down the side, and a score in each cell. Your appearance rate is the percentage of cells where you scored Recommended or Mentioned out of the total fifty.
What Your Score Means
There is no industry benchmark for this number yet. What matters is the baseline itself, because everything that follows is measured against it.
If you appear in fewer than 10% of responses (under 5 out of 50), AI assistants are not recommending your business to your buyers. That's more common than you'd think. When we ran this exact audit on ourselves in July 2026, we were named in zero of fifty responses, and forty-five of those fifty named nobody at all.
If you appear in 10% to 30% of responses, you have some visibility but it is inconsistent. You likely appear in engines that have indexed recent content but are absent from engines that rely more heavily on training data or structured signals.
If you appear in more than 30% of responses, you have meaningful AI visibility. The question shifts from "how do I get visible" to "how do I maintain and expand this position as competitors enter the space."
The number matters less than the trend. Run this audit monthly with the same ten questions. A business that moves from 4% to 22% over three months has measurable evidence that its visibility work is producing results. For a complete walkthrough of measurement methodology, read how to measure whether AI visibility work is actually working.
What to Do with the Results
The audit tells you three things no vendor pitch can substitute for.
First, it tells you where you stand. Before any conversation with a provider or any internal decision about investment, you have a number. That number is yours, not a vendor's proprietary score.
Second, it tells you which engines see you and which do not. The pattern reveals where your evidence trail is strong and where it is thin. A business that appears in Perplexity (which searches the web in real time) but not in ChatGPT (which relies more on training data) likely has recent content that has not yet been absorbed into model training.
Third, it shows you what your competitors' AI visibility looks like. The same audit that checks your appearance also surfaces who the engines recommend instead of you. That competitive map is information you cannot get from any other source.
If your appearance rate is low and you want to understand what to do about it, start with the mechanism: how ChatGPT decides who to recommend explains the evidence trail that earns a recommendation. To understand what a program to build that evidence looks like, visit how it works.
If you'd rather skip the DIY version and get a complete picture, the 109-Point AI Visibility Diagnostic runs the audit across eight zones and scores your position with a methodology you can see. It's free, and it gives you the baseline this article describes without the forty-five minutes of prompt-by-prompt work.