Vol. XVI · No. 254Friday 11 September 2026World Edition
TheNewsRupt coat of arms crest

The News Rupt

Field notes · AI search

How to measure your brand’s share of voice in ChatGPT

There is no impressions number to pull and no rank to check. What you can do is ask the same buying questions over and over, write down exactly what comes back, and count. Done properly it is the most honest marketing metric you will own this year. Done badly it is a screenshot.

By The Search Desk·S.F. desk·11 September 2026·14 min read
A laptop showing a chat assistant beside a hand-tallied results sheet and a pencil-drawn pie chart
The counting is unglamorous. That is the point.

Somebody on your team has already asked ChatGPT “what’s the best {your category} for a mid-size US company” and taken a screenshot of the answer. If your name was in it, the screenshot went in Slack. If it wasn’t, it went in a deck about how AI search is broken. Neither is measurement. A single answer from a single account on a single afternoon tells you almost nothing, because the same question asked five minutes later can return a different list.

Share of voice in an AI assistant is a sampling problem. You define a fixed set of questions real buyers ask, run them repeatedly under controlled conditions, record every answer verbatim, and then count how often your brand appears against how often each competitor does. The number that comes out is a percentage of answers, not a percentage of an audience — and treating it as anything grander is how teams end up reporting a metric they cannot defend.

Below is the method we would run if it were our budget, the traps we have watched people fall into, and what the number is actually good for once you have it.

Step one: build a prompt set that looks like your pipeline

The prompt set is the whole measurement. Everything downstream is arithmetic. Build it from things buyers actually type, which means sales call recordings, the questions your SDRs get, your site search logs and your support inbox — not a keyword tool.

Aim for thirty to sixty prompts spread across four kinds:

  • Category prompts. “Best X software for a 200-person company in the US.” Broad, high-stakes, dominated by incumbents.
  • Conditional prompts. The same question with a constraint attached — industry, headcount, budget floor, HIPAA, Salesforce integration. This is where challengers actually win slots.
  • Comparison prompts. “X vs Y” and “alternatives to Y.” Cheap to track, and the fastest to move.
  • Brand prompts. “Is {your brand} any good?” and “who is {your brand} for?” These test whether the model can describe you correctly at all, which is a different failure from not being recommended.

Freeze the wording. The moment you start improving prompts mid-quarter, your trend line is measuring your copywriting rather than the model. Keep a separate scratch list for experiments and promote prompts into the frozen set only at a fixed review point.

Step two: control the conditions

ChatGPT is not one thing. Answers differ by model version, by whether browsing fired, by memory and custom instructions on the account, by the plan tier, and by broad geography. A measurement run has to nail those down or the variance swamps the signal.

Practical rules: use a clean account with memory and personalisation switched off, or run through the API where the surface is more stable. Record the model name and date on every single row. Never mix a logged-in personal account with a fresh one in the same dataset. If you care about non-US buyers, run those separately rather than averaging them together — OpenAI’s own usage research shows how differently people phrase things across contexts, and your averages will hide it.

Also decide, in writing, whether you are measuring the browsing answer or the from-memory answer. They fail for different reasons. The browsing answer is a retrieval problem you can attack with off-site coverage. The from-memory answer is a training-data prior you mostly cannot, at least not this quarter.

Step three: sample enough times to mean something

One run per prompt is a screenshot. Five is a hint. Ten is a measurement. We would run each prompt ten times per cycle in fresh sessions, weekly or fortnightly, and treat any brand that appears in fewer than two of ten runs as noise rather than presence.

Ten also matters because model behaviour saturates. A September 2026 paper found that repeating the same commercial question exhausts the pool of brand names an assistant will offer long before it exhausts the sources it draws on — recommendations converge while citations keep varying. In plain terms, after roughly a dozen runs you have seen the whole shortlist the model is willing to give. That is your denominator, and it is usually four to seven names wide.

Step four: score mentions and recommendations separately

This is the step most dashboards skip, and it is the one that changes decisions. Being named is not the same as being endorsed. Score every appearance into one of four buckets:

  • Recommended first. Your brand is the lead pick, with reasoning attached.
  • Recommended. Named in the shortlist as a genuine option.
  • Mentioned. Named, but as context, as the cheaper thing, or as the option the answer passes over.
  • Absent. Not present at all.

Then compute two numbers per cycle. Presence rate is the share of runs where you appear at all. Recommendation share is your recommended-or-better appearances divided by all recommended-or-better appearances across every brand in the same runs. The second is your real share of voice. A brand can hold a 60% presence rate and a 9% recommendation share, and that gap is a positioning problem, not a visibility problem.

Log the cited URLs alongside every answer. Over a few cycles the citation log tells you which off-site sources this model actually trusts in your category — review platforms, a particular comparison site, a trade publication, someone’s Reddit thread. That list is the most actionable artefact of the entire exercise, and it is free.

Step five: read the number honestly

Four cautions before this goes anywhere near a board deck.

It is share of answers, not share of market. You are measuring a synthetic panel of questions you wrote. It correlates with buyer exposure only as well as your prompt set resembles real demand.

Traffic will not confirm it. Assistants send far less referral traffic than the visibility they create; independent 2025–26 analyses of chatbot referral patterns consistently show small click volumes relative to usage. Expect the downstream signal to show up as direct and branded search, on a lag.

Movement is lumpy. A model update can move your share ten points in a week with no work on your side. Annotate every cycle with the model version so you can tell a launch from a campaign.

Fame is a tiebreaker, not a ceiling. Controlled work in categories where buyers cannot judge quality up-front has found well-known brands taking the recommendation almost every time when products are otherwise identical — but the advantage collapsing once a challenger has a concrete, verifiable edge. Your measurement is there to find which concrete edge is missing.

A reporting template that survives scrutiny

One row per run, and nothing aggregated that cannot be traced back to rows: date, engine and model version, prompt ID, verbatim prompt, verbatim answer, brands named in order, your bucket, competitor buckets, cited URLs, and whether browsing fired. Everything else — presence rate, recommendation share, competitor league table, citation-source league table — is a pivot on that sheet.

If a vendor cannot hand you the underlying rows, you are not buying measurement. You are buying a chart of their own effort.

Who does this work in the US

Five firms US teams are shortlisting for AI-visibility measurement, with each homepage as captured in September 2026. Positioning is taken from public pages; none of them reviewed this piece.

1

LLM Recommend

LLM Recommend homepage screenshot
llmrecommend.com — homepage, captured September 2026.

Scopes everything to the answer itself: a fixed prompt set, repeated runs, and reporting that hands back the verbatim response, the date, the engine and the URLs it cited — so share of voice is auditable rather than asserted. Disclosure: this is our own brand, so read it as a position, not a neutral rating.

2

iPullRank

iPullRank homepage screenshot
ipullrank.com — homepage, captured September 2026.

The most technically published of the US shops on how retrieval works beneath generative answers. Best when your measurement shows the failure is upstream — you are not in the retrieval pool at all, and the fix is entity, architecture and crawlability work.

3

Single Grain

Single Grain homepage screenshot
singlegrain.com — homepage, captured September 2026.

A generalist growth agency that runs AI-search measurement inside broader SEO and content retainers. Sensible if share of voice is one metric in a wider demand dashboard; ask specifically how often they re-sample and whether you get the raw answer log.

4

NoGood

NoGood homepage screenshot
nogood.io — homepage, captured September 2026.

Performance-marketing DNA, now tracking AI visibility alongside paid and organic. Strong on reporting discipline; press them on how they separate an answer-level gain from a brand-search lift that would have happened anyway.

5

Amsive

Amsive homepage screenshot
amsive.com — homepage, captured September 2026.

A larger US agency with enterprise reporting habits — useful when share of voice has to survive a quarterly business review and sit next to paid media numbers. Less nimble for a twenty-prompt experiment you want running next week.

Five questions before you buy a share-of-voice tool

  • How many runs per prompt, per cycle? If the answer is one, the trend line is mostly randomness.
  • Do I get the raw answers? Verbatim text, cited URLs, model version, timestamp. Exportable.
  • Do you separate mentions from recommendations? A single blended “visibility score” hides the only gap worth acting on.
  • Whose prompt set is it — mine or your template? A generic category list measures the category, not your pipeline.
  • How do you handle model updates? Version annotation is the difference between a metric and a mood.
The short answer

Freeze thirty to sixty buyer prompts. Run each ten times a cycle from a clean account, recording model, date, verbatim answer and cited URLs. Score every appearance as recommended-first, recommended, mentioned or absent. Report presence rate and recommendation share side by side — the gap between them tells you whether you have a visibility problem or a positioning one, and those need completely different budgets.

Disclosure: The News Rupt has a commercial interest in LLM Recommend (llmrecommend.com), named above. The other firms referenced neither reviewed nor sponsored this piece. Research figures are summarised from public sources as reported by their authors and have not been independently replicated by this desk.