Someone asked ChatGPT which vendor to shortlist last Tuesday. Your competitor got named. You didn't.
The complete guide to being found inside the answer, not ranked underneath it.
Ten chapters, in the order the work has to happen. Each one opens with a diagram, because the mechanism is easier to see than to read.
Whose buyers now get an AI shortlist before they ever reach a website.
The scene carries the argument. The text under it is what you need to brief someone.
Ranking correlates with citation. It does not cause it. Chapter 3 has the receipts.
Before you start
Your buyer never sees the search. They see a finished answer naming three vendors, with no trace of who was considered and dropped. If you were the one dropped, no report tells you, and the deal is shaped before anyone visits your site. That matters because the dropping usually happens for mechanical reasons rather than competitive ones. The two panels below are one answer seen twice: what your buyer reads, and what the machine had to work with.
This was not a product loss. All three cited competitors returned specific, parseable claims. Your page returned a shell, and the model moved on.
From my experience, this is the most common failure and the hardest to spot. Rankings hold, traffic drifts, and the shortlist quietly stops including you.
The mechanics of retrieval, the structure of extractable proof, and the measurement discipline that tells you whether any of it worked. Not recycled SEO advice with "AI" pasted on top.
Chapter 01 · The pipeline
Teams treat AI visibility as one thing you either have or you do not. It is really five filters in sequence, so a company can pass four of them, never appear once, and spend the next quarter fixing the wrong one. That matters because the stage where you drop out decides the fix, and nothing in your analytics tells you which stage it was. The funnel below follows a hundred buying prompts and marks where each loss happens.
"We publish a lot of content" addresses one stage. A page that cannot be retrieved is never extracted, and one that cannot be extracted is never synthesized.
The shape is illustrative. Your own curve depends on your category and your site. What holds everywhere is the direction: it only narrows.
Query reformulation. The model does not search your prompt. It rewrites it into queries of its own, then searches those. Ask it to explain TCP handshakes and it answers from memory. Ask which vendor supports Snowflake and Okta and it searches, because that answer is specific, current and checkable.
Chapter 01 · The multiplicative problem
So which stage do you fix first? Most teams pick the one they are already strongest at, because that is the team they have. That matters because the five gates multiply rather than add, which makes the instinct backwards: your worst gate sets your ceiling no matter how good the other four are. Both channels below carry the same water through the same five gates, with one difference.
Read the top channel first. Both carry the same water through the same five gates, and the only difference is the coral gate near the end. Channel B beats A at four of the five and still passes a fifth as much water.
Which is why the order in Chapter 8 is not a suggestion. Twelve months of publishing multiplied by a broken extraction gate is still close to zero.
The content team writes more content. The PR team gets more mentions. Meanwhile the site returns JavaScript shells to the fetcher and nothing is extracted at all. Find your weakest gate. Fix that one.
Chapter 02 · The reformulation layer
You picked a target keyword, wrote the page, ranked for it, and still were not cited. The model never searched your keyword. It rewrote the buyer's question into several queries of its own and searched those instead. That matters because your coverage is being decided by queries you have never seen and cannot look up. The prism below takes one real prompt and shows the six searches it became.
What our own study showed. Across 8,616 API runs we found that on community-style questions, ChatGPT's larger models cited Reddit in 298 of 300 cases without being asked to. Then we tried to break the result, and the follow-up tests were the real finding: the large models add platform names and site operators to their own queries unprompted. GPT-4.1 mini does not.
So reformulation is a property of the model, not of your prompt. The more capable the model, the more aggressively it rewrites, which means your visibility shifts on someone else's release schedule. Log the model version with every measurement you take.
Source: Growtika, "AI Thinks the Internet Is Reddit." 8,616 OpenAI API runs, July 18 to 19, 2026. Full study and method at growtika.com.
You cannot see the reformulations. Generate the ten to twenty queries a capable model would plausibly write, then check coverage one by one.
If a whole cluster returns nothing of yours, that cluster is your roadmap.
Chapter 02 · Zero volume, high value
If the model writes its own queries, your keyword tool is measuring the wrong thing. It reports volume for phrases humans type, so multi-constraint buying questions come back as zero and never make it into a brief. That matters because the filter hiding them from you is hiding them from your competitors too. The map below shows which ground is crowded and which is still open.
The test is not how many people search it. Ask instead: would a model write this query to answer a real buying question, and is the answer verifiable from published text? If both are yes, the volume number is irrelevant.
Why the ground is still open. Keyword tools return zero, so the brief never gets written, so nobody publishes the page. The filter that hides these queries from you hides them from your competitors too.
An empty plot is only worth claiming if you can put a real number on it. A page that restates the question without answering it still gets retrieved, and the model then has nothing to quote from it, so the slot goes to whoever did answer.
Chapter 03 · With receipts
The comfortable assumption is that AI search is a Google derivative: rank well, get cited, keep one dashboard. If that is wrong, every AI visibility number you have reported is describing a different channel. That matters because it is testable in about a week. Build a thousand matched pairs, run each through its own system, and compare where the domains land. Below is the shape the answer takes.
The shape is stable. The decimal places are not. Percentages shift as models and indexes update, so treat any published figure as directional.
The method is cheap. A thousand pairs is a week with a script. Two hundred already show you the shape, and it will be your shape rather than someone else's blog post.
On this method: the matched-pair test is a procedure you run on your own category, not a published Growtika result. The ribbon widths above are drawn to illustrate the shape, and your numbers will differ.
Chapter 03 · Four structural reasons
So why do ranking and citation come apart? Because the two systems are built to make different things: one returns an ordered list of pages to visit, the other assembles a paragraph out of borrowed passages. That matters because a page can be the best destination for a query and a poor source for a sentence. Once you have seen both machines below, the four reasons in the table stop being surprising.
Left, the ranking engine. Whole pages go in, a sorted list of ten comes out, and the crawler feeding it had time to render and come back.
Right, the answer composer. Passages go in, one paragraph comes out, and anything that returned nothing or contradicted the others is simply not in it.
| Reason | ChatGPT, and why it matters to you | |
|---|---|---|
| Different objectives | Ranks documents by likely relevance. | Picks passages it can build a paragraph out of. A page can be the best place to send someone and a poor place to quote from. |
| Reformulation drift | Matches your keyword, more or less as typed. | Searched three other things. Your rank for the original term never entered the calculation. |
| Extraction sensitivity | Patient crawler: renders, waits, retries, revisits. | Impatient fetcher, often without rendering. A page that ranks first can return nothing usable and be silently skipped. |
| Corroboration weighting | No equivalent mechanism. A strong single page can rank alone. | Favors claims that appear in several independent sources. A unique claim on one authoritative page often loses to a consistent claim across four mediocre ones. |
Chapter 04 · The extraction layer
Now to the gate that fails without telling you. Your page ranks, the fetcher opens it, nothing usable comes back, and no report anywhere records the miss. That matters because you will read the flat results as a content problem and brief more content. The two panels below are the same URL: the version your team signed off, and the only version that can ever be cited.
The diagnostic takes one minute. Fetch the page with curl, JavaScript disabled. If your differentiating claim is not in that raw HTML, it does not exist to the model. Do this before you brief anyone on content.
There is a second failure here. Even in server-rendered markup, the claim sits inside a tab. Content behind interaction is often invisible, so your specifics hide exactly where the model was looking.
Chapter 04 · Chunk-level thinking
Suppose the fetcher does get your HTML. It still does not read your page, it reads fragments of it, and it judges each fragment with no memory of the paragraph above. That matters because most marketing sentences stop meaning anything once they are alone. Below are four sentences from one page, after the page is gone.
The habit to build. No orphan pronouns. No "as mentioned above." No numbers without units and baselines. If a sentence needs the paragraph above it to make sense, rewrite it or accept that it will never be cited.
It doubles as a review test. Take any draft and cover everything except one sentence. The sentences that still say something with their neighbours hidden are the ones doing the work.
| What helps | Why it works once the page has been split up |
|---|---|
| Question-shaped headings | They match reformulated queries directly, which is what retrieval compares against. |
| Answer-first paragraphs | First sentence answers, the rest supports. If the chunk truncates, the answer already landed. |
| Tables for comparison | Structurally unambiguous. Row and column relationships survive extraction intact. |
| Mission-statement openings | The first chunk, the one most likely to be read, contains nothing verifiable. |
| Key facts in images and PDFs | Often unreadable, always more expensive to extract than the HTML page you could have published instead. |
Chapter 05 · The micro-feature gap
You lose a recommendation to a competitor with a weaker product, and it reads as a brand problem. It is almost always a publishing problem. Models check a specific set of small, checkable facts against public text, so the vendor who wrote the number down takes the slot. The console below runs five of those checks across three vendors, including you.
Run the audit yourself. Pick a category and twenty representative prompts. Record which vendors get recommended and, critically, why the model says it recommended them. Those stated reasons are your row labels.
We first ran this for a clinical AI scribe client. The models kept naming two competitors. Neither had a stronger product. Both had published proof the client had only ever said out loud on sales calls.
Where this came from: first run for a clinical AI scribe client in 2026, on a category where two competitors were being named repeatedly. Output was a fifty-article roadmap with one target claim per article.
Does a public, specific claim exist? Not "do we do this," and not "could we do this." Is the proof sitting in text a fetcher can read?
Chapter 05 · From matrix to roadmap
The audit hands you a list of gaps, and the temptation is to brief an article for each one. Some of them cannot be closed by writing at all, and briefing those burns a quarter for nothing. That matters because sorting them takes an afternoon and getting it wrong costs the quarter. Every gap leaves the hopper below on one of three rails.
Each winnable gap becomes one article with one target claim. Not "write about integrations." Write a page that states, specifically and extractably, which systems you integrate with, what the setup time is, and what a customer measured.
It inverts the usual brief. Instead of starting from a keyword and hoping proof emerges, you start from the proof the model wants and build the smallest page around it. I have used this to produce fifty-article roadmaps, one named claim per article.
Target keyword, word count, three competitors to beat, tone guidance.
Target claim: the exact sentence that must appear, with its number, baseline and scope.
Micro-parameter: which one this closes.
Proof source: where the number comes from and who signs off.
Extraction check: visible in raw HTML, early in the chunk.
Chapter 06 · The off-site layer
You can publish the perfect claim on the perfect page and still not be cited, because a model also weighs how many independent sources agree with you. Forty mentions in forty places you control is one source with forty URLs. That matters because most link budgets are spent on exactly those forty. What earns weight is editorial distance, which is what the balance below measures.
One trade-publication mention outweighs a year of guest posts. From my experience teams spend more on the bottom of this ladder than the top, because the bottom is easier to buy and easier to count.
Audit your own mix once. Take your last twenty mentions and mark the ones that pass all three tests. The surviving number is usually a lot smaller than the one in your board deck.
| Source tier, heaviest first | Weight | What earns the weight |
|---|---|---|
| 1 · Independent editorial and analyst notes | High | A publication with its own standards decided the claim was worth printing. |
| 2 · Well-sourced technical writing citing you | High | Practitioners with reputations attached to being right. |
| 3 · Community discussion by real users | Mid | Real phrasing, real specificity, no marketing filter. |
| 4 · Review platforms with written detail | Mid | The prose is what gets extracted. Star ratings are not text. |
| 5 · Directories and listicles you did not write | Low | Useful for corroboration, weak on their own. |
| 6 · Your own content | Low | Necessary, never sufficient. This is the assertion, not the evidence. |
Chapter 06 · The entity record
Independence is only half of it. The sources also have to agree, and most companies have quietly accumulated three or four different descriptions of what they sell across profiles, partner pages and old press. That matters because conflicting descriptions do not average into a fuzzy answer. They cancel. The tally below counts the votes for one company against a competitor with one story.
The stale partner page from 2023 is not a cosmetic problem. It is an actively retrieved contradiction, and you cannot delete someone else's page. You can only give them a reason to update it.
Write the canonical claims down once: the exact name, one category descriptor used everywhere, the differentiator in a single quotable sentence, and the three to five numbers you want repeated back to you.
Community content gets retrieved far above what its authority scores predict, because it carries the phrasing real people use. Show up as a practitioner and answer specific questions with specific detail, including where your product is the wrong fit. Stated limitations are among the most trusted content you can publish, because almost nobody publishes them.
Not astroturf. Models and communities both surface fake engagement, and once a thread about manufactured reviews is part of your record, the model synthesizes that too.
Chapter 07 · Measurement
Everything so far is a hypothesis until you can show it moved. The trap is that answers are stochastic, so two runs of the same prompt in the same week can differ by eight points and a single run lets you call either one progress. That matters because it is the easiest way to report a win that never happened. Below is eight months of a fixed prompt set with every sample plotted, not just the average.
Two things separate a measurement from an anecdote. Fifty to three hundred real buying questions, held stable so results stay comparable, and a fixed protocol: same model, same settings, three to five samples per prompt.
Change the prompt set and you lose the series. Add prompts to a second set if you must, but never edit the baseline mid-program.
| Metric | What it tells you | Common mistake |
|---|---|---|
| Mention rate | How often you appear at all. | Treated as the goal rather than the floor. |
| Citation rate | How often you are linked, not just named. | Conflated with mention rate, which moves independently. |
| Position and framing | Named first, listed, footnoted; recommended, alternative, or limitation. | Counted as a win regardless of which one it was. |
| Source attribution | Which of your pages actually got cited. | Skipped entirely. The most actionable number you have, and usually not the page you expected. |
Chapter 08 · The implementation sequence
You now have more work than a quarter holds, and the order is not free. Because the gates multiply, work done early against a broken gate is wasted, and the phase you cannot speed up is the one teams schedule last. That matters because publishing into a broken extraction layer is publishing into a void. The schedule below puts all twelve weeks in dependency order.
Read the hatched bar as the counterfactual. It is the same publishing work, done in the same weeks, against pages the fetcher cannot read. Same cost, no citations, and a report at the end of the quarter saying AI search does not work.
Phase six spans the whole quarter because editorial coverage runs on someone else's calendar, not yours.
Publishing sits at phase five because its results multiply against every gate above it. Corroboration starts in week one for the opposite reason: it is the only phase whose speed is not yours to set.
Chapter 09 · What does not work
Some of the most confidently sold AI-search tactics do nothing at all. They share one property: each tries to influence selection without producing proof, and the mechanism has no shortcut in it. The only thing that varies is how far each road below runs before it stops.
The lanes are not equally bad. Schema and llms.txt are harmless work that runs a long way before stopping, which is exactly why they get mistaken for strategy.
The unit of value is the citable sentence, not the published URL. Every road above skips that unit and hopes the mechanism will not notice.
For any tactic, ask which of the five gates it moves: retrieved, extractable, trusted, corroborated, synthesized. If the honest answer is "none, but it feels like AI work", it belongs on this page.
Chapter 10 · The uncomfortable conclusion
Two conditions make this the moment rather than next year: zero-volume queries are still uncontested, and most competitors still return empty HTML to fetchers. Both are temporary and neither will announce its closing. Do nothing and nothing visibly breaks, which is the failure mode of this channel: it is quiet.
The gap is the opportunity. The slab coming down is competitors fixing their extraction layer. The one rising is the open ground getting claimed. Everything you build now happens in the space between them.
Proof beats presence. The winners are not the loudest. Their claims were sitting in clean HTML when the model went looking. Every citation loss I have diagnosed came down to a sentence that either did not exist or could not be read.
Optimize for the invariants. Models, indexes and fetchers will all shift, but there will always be a retrieval step, an extraction step and a synthesis step. Get those right and the version numbers stop mattering.
Someone is asking ChatGPT about your category right now. Make sure the answer has your proof in it.
The checklist
The whole guide, compressed into fourteen lines and kept in gate order. Because they multiply, a tick in the second column is worth nothing while a box in the first is still empty.
First, the gates that block everything
Then, the work that compounds
"Show me the raw HTML of the page where we state our differentiator, and read me the one sentence a model would quote." If that takes longer than a minute to answer, you have found your weakest gate.
It runs a query, grabs what parses cleanly, and moves on. Every diagram in this guide follows from that one sentence.
Where this came from
Growtika is a GEO-first B2B SaaS agency. The frameworks here, the Selection Chain, reverse fan-out analysis and the micro-feature gap audit, come out of original citation research across thousands of prompts and applied work with security, fintech and developer tools clients.
13 chapters, 18 diagrams, and the exact playbook we use to get clients cited inside ChatGPT answers. Free.
One email. No spam. Unsubscribe anytime.
Helping SaaS companies and developer tools get cited in AI answers since before it was called "GEO." 10+ years in B2B SEO, 50+ cybersecurity and SaaS tools clients.
Most brands never appear in ChatGPT answers. Here is what separates the cited from the invisible.
Research1,000 B2B buyer questions. ChatGPT cited 3,971 domains. 2,842 appeared once. 48 were cited 11+ times.
Original Research8,616 GPT calls show models rewrite questions into Reddit searches, then cite Reddit even when users never name it.