A lot of the present recommendation round generative engine optimization (GEO) relies on concept, remoted screenshots, or a single marketing campaign. I wished measurable outcomes, so I ran two structured experiments again to again, tracked the outcomes manually, and documented what labored, what failed, and what modified between the 2 assessments.
The primary experiment, for an current model we consulted on, price hundreds of {dollars} and ran over a number of months. The second was a 30-day cold-start experiment a SaaS hyperlink constructing company with no measurable AI presence when the take a look at started.
Every take a look at tracked 15 commercial-intent key phrases. The primary coated 4 AI platforms, whereas the second coated six. I logged 775 quotation occasions throughout each experiments, and one in every of my unique conclusions didn’t survive the second take a look at. Let’s dig into how I set these up, what I noticed, and how one can apply it to your work.
Experiment 1: The marketing consultant take a look at
We tracked 15 commercial-intent key phrases throughout 4 platforms: ChatGPT, Claude, Gemini, and Perplexity. Each question was run manually, with and and not using a VPN, to account for doable location-based variations.
Our technique was to position listicles on sources that had been generally surfaced by LLMs for our key phrases, in addition to on the model’s web site. We supported these by way of PR, visitor posts, and natural LinkedIn exercise.
By the top, the model appeared for roughly 10 to 12 of the 15 key phrases. Peak key phrase presence reached 37.01% on April 29. The general development was upward, with significant week-to-week variation. The citations by platform had been as follows: ChatGPT 148, Claude 96, Gemini 87, and Perplexity 64.
The main supply for citations was a complete listicle on Certainly web optimization, with 190 mentions. MEXC and a GlobeNewswire launch adopted, however neither got here near the identical quantity. In all, listicles accounted for 72.4% of citations and PR 24.1%. Visitor posts, the owned web site, and LinkedIn break up what was left.
We additionally realized 4 main classes you possibly can apply to your subsequent mission.
Be the brand AI recommends.
See where your brand appears in AI search, where competitors are winning, and what it takes to become the answer AI recommends.
See your AI visibility
Listicle placement, PR, and visitor posts reinforce one another
On this experiment, listicles and PR appeared to work as one system reasonably than as separate channels. The listicles that earned citations had been often the identical ones being amplified by way of PR. The listicle launched the declare, the press launch bolstered it, and visitor posts referenced each. Every layer appeared to carry out higher when supported by the others.
We discovered securing a placement on a supply the mannequin already cites can compound visibility shortly. Right here, Certainly web optimization appeared round 10 instances in ChatGPT’s solutions earlier than we contacted the publication.
After our placement went reside, it remained our largest quotation supply for months. That placement generated extra citations than all of the others mixed.
Visitor posts appeared to increase the impact of the PR placements. Claude, specifically, cited visitor posts that referenced options in Yahoo and Enterprise Insider, even when it didn’t cite these publications instantly.
Supply decay is fast and authority issues
Roughly half of the sources stopped being cited inside 30 days. One MEXC placement fell from 29 mentions to 11 week over week, whereas a Triple Evaluation listicle that had been cited constantly declined with none direct intervention. The consequence means that one-time publication is unlikely to maintain visibility by itself.
In our assessments, supply authority and relevance appeared to matter greater than content material high quality alone. A sensible hierarchy was:
- Authorities and schooling sources.
- Information publications.
- Business-relevant websites.
- Common websites.
In Section 1, trade placements supported by PR produced robust outcomes. In Section 2, better-written listicles on normal websites produced little visibility. PR appeared able to amplifying a robust placement, however not compensating for a weak supply.
Depth appeared to matter greater than frequency. Our first 5 listicles coated the subject solely briefly, and most generated little visibility. The one complete piece continued to earn citations.
The peer set appeared to affect visibility. One listicle positioned us alongside Lily Ray and Aleyda Solis. After these names had been eliminated, efficiency declined inside days. This means that the fashions could consider the encompassing entities, not solely the person point out. An analogous sample appeared in PR, the place being named alongside acknowledged specialists outperformed a standalone characteristic on a stronger outlet.
SERP visibility nonetheless appeared to affect LLM visibility, particularly in ChatGPT. When the mannequin relied on internet search to resolve a question, manufacturers absent from the retrieved outcomes had been additionally absent from the reply.
Capitalization and question kind matter
The capitalization of queries appeared to have an effect on which sources had been retrieved. Capitalized and lowercase variations returned totally different citations in three repeated assessments, though this discovering requires additional validation.
Question kind appeared to affect the supply varieties chosen by the fashions. Software program and power queries favored high-authority evaluation websites, whereas service queries extra usually returned listicles. Matching the location kind to the question appeared to enhance the probability of being cited.
Actual-match key phrases and reply placement are vital
Actual-match key phrase concentrating on nonetheless appeared to matter. We ranked for “Greatest LLM web optimization Marketing consultant” however barely appeared for “Greatest AI web optimization Marketing consultant,” regardless of the same intent. The primary phrase had a devoted listicle, whereas the second didn’t. On this take a look at, broader semantic protection didn’t bridge the hole, suggesting that high-value business key phrases could require devoted belongings.
In our assessments, pages carried out higher when the reply appeared throughout the first 100 phrases. A key takeaway block close to the highest of the web page produced a bigger enchancment than some other on-page change we made.
Different content material findings embrace:
- FAQ content material carried out higher when it was seen by default reasonably than hidden behind expandable sections.
- Self-contained sections appeared to carry out higher. For instance, “What to search for when hiring an LLM web optimization knowledgeable” and “The place to rent one” labored higher as separate sections than as a single mixed part.
- Query-based headings additionally carried out higher in our assessments. For instance, “How is AI web optimization totally different from conventional web optimization?” outperformed “AI web optimization vs. conventional web optimization.”
- Freshness additionally appeared to matter. Current knowledge, present references, and visual publication dates had been related to stronger quotation efficiency.
We realized so much from this primary take a look at.
The unique plan for spherical two was an inventory of issues to check: Particular person schema, LinkedIn cadence, and a YouTube push. As an alternative, I acquired the possibility to run the entire playbook from zero on a special model in a special vertical, which is a much better take a look at of whether or not any of the findings are generalizable.
Dig deeper: 3 GEO experiments you should try this year
Get the publication search entrepreneurs depend on.
Experiment 2: The chilly begin take a look at
The take a look at concerned a SaaS hyperlink constructing company with no measurable AI presence when the experiment started. In the course of the baseline window from April 30 to Might 29, the model had no measurable presence on any tracked platform.
The experiment ran from Might 30 to June 28, and coated six platforms, including Google AI Mode and Grok to the listing from the primary experiment. We tracked 15 commercial-intent key phrases {that a} SaaS purchaser would possibly use whereas evaluating an company. Each platform and key phrase was checked manually.
The experiment resulted in 298 appearances in 30 days from a standing begin. The appearances by platform had been: Gemini 104, Google AI Mode 95, Claude 59, ChatGPT 32, Grok 4, and Perplexity 4.
Gemini and Google AI Mode accounted for roughly two-thirds of the platform totals listed above. This differed sharply from experiment one, the place ChatGPT led. The comparability must be handled cautiously as a result of AI Mode wasn’t tracked within the first experiment, and the 2 niches weren’t instantly comparable.
Enterprise impression: In the course of the 30-day window, 18.5% of latest customers arrived by way of referral visitors, and one other 3.25% by way of GA4’s AI Assistant channel. Collectively, these channels accounted for simply over one-fifth of all new customers. One Perplexity referral led to a prospect who later turned a paying buyer, regardless that Perplexity was the lowest-volume platform within the experiment.
Right here’s what this experiment taught me.
Give attention to noticed citations and earned placements
I realized you wish to construct the goal listing from noticed citations, not from DR alone. Earlier than starting outreach, I ran all 15 key phrases by way of each platform, logged the sources that appeared, and ranked them by quotation frequency. That ranked listing turned the outreach listing.
5 of eight targets appeared within the ultimate quotation combine. Indie Hackers elevated from 44 mentions throughout prospecting to 146 after our placement went reside, a 232% enchancment. Bruce Jones web optimization elevated from 26 to 69, up 165%, whereas TechBullion rose from 15 to 37, a 147% soar.
The strategy didn’t work in each case. RankTracker and HR.com confirmed fewer citations after placement than throughout prospecting, and two shortlisted websites hadn’t appeared in any respect by the top of the measurement window. The strategy appeared to enhance the hit charge, nevertheless it didn’t assure citations.
Placement focus was excessive. Three sources — Indie Hackers, Bruce Jones web optimization, and our personal listicle — accounted for 342 of 437 whole supply mentions, or roughly 78%. Seven different reside placements shared the remaining mentions. On this experiment, a small variety of sources drove a lot of the visibility.
Inside this experiment, earned placements outperformed owned content material by a large margin. Of the identical 437 supply mentions, third-party listicles generated 85.8%, our self-published listicle generated 14.0%, and PR generated 0.2%.
Measure citations over time
Time to quotation ranged from one to 18 days. Two placements had been cited the day after publication, whereas others took 10 or 11 days. 4 placements had been reside however hadn’t been cited by the top of the measurement interval. The variation means that checking solely as soon as, one week after publication, isn’t a dependable measurement strategy.
The slowest supply to be cited was the one we managed. Our personal listicle took 18 days, longer than each third-party placement, together with two that had been picked up in a single day.
Owned listicles are a basis, not a development lever
That is the place I needed to revise my unique conclusion.
After experiment one, I beneficial publishing listicles on an owned web site as a result of many top-ranking manufacturers appeared to profit from their very own content material. Experiment two instructed that an owned listicle is extra of a basis than a major development engine. It was the slowest supply to be cited and contributed 14% of mentions, whereas earned placements generated a lot of the visibility.
Dig deeper: How to know if your GEO is working
Comparative content material is efficient
The owned listicle improved when it turned much less promotional and extra comparative. On June 23, we up to date it to incorporate our main opponents as an alternative of presenting the model alone. Mentions of that supply rose from 4 to 49, a 12.25-fold enhance within the ultimate depend. The day by day visibility curve additionally elevated throughout the identical interval, from 22 on June 23 to a peak of 95 on June 27.
This mirrored the peer-set impact noticed in experiment one, however from the wrong way. Eradicating acknowledged names was adopted by a decline in spherical one, whereas including acknowledged opponents was adopted by a considerable enhance in spherical two. The identical sample appeared throughout two manufacturers and two verticals, making it one of many findings I’d prioritize for additional testing.
Comparative protection outperformed advocacy on this take a look at. The 12.25-fold enhance adopted an replace that made the web page extra helpful as a class useful resource reasonably than as a web page centered totally on our personal firm. The area, creator, and key phrase goal remained the identical. The primary change was the scope of the content material.
Citations don’t equal clicks
Probably the most-cited and most-clicked sources weren’t the identical.
Indie Hackers generated extra quotation quantity than some other supply, however its referral visitors remained flat. TechBullion produced fewer citations however elevated periods from one to 64. Claude.ai referral periods tripled, whereas ChatGPT referral periods elevated by 166%.
These outcomes counsel that quotation quantity alone isn’t enough for deciding which sources deserve additional funding.
Intent varies by mannequin
On this experiment, Claude concentrated extra closely on high-intent phrases. Gemini led in whole appearances, 104 to 59, however Claude led or tied for first on 5 of the ten best-performing key phrases.
Gemini’s quantity was distributed extra broadly, whereas Claude’s was extra targeting business queries. For a enterprise evaluating platform worth, that distribution could matter greater than the headline whole.
Actual-phrase belongings additionally carried out properly
Our two strongest key phrases had been “Greatest SaaS Hyperlink Constructing Company in USA” and the identical phrase with “2026” added, with 18 appearances every. This repeated the sample from experiment one in a special area of interest and means that devoted belongings should still be mandatory for high-value business key phrases.
If AI can’t find you, customers won’t either.
Track your visibility across AI search, uncover missed opportunities, and grow your presence where customers are asking questions.
See your AI visibility
What I’d inform somebody beginning at this time
These experiments supplied me with a stable listing of classes that may make it easier to prioritize your efforts.
- Run your goal key phrases by way of the related AI platforms earlier than investing in placements. The sources already being cited ought to inform the outreach listing, and that listing could differ considerably from a standard prospecting spreadsheet.
- Price range for upkeep, not solely preliminary publication. In our first experiment, roughly half of the sources stopped being cited inside 30 days.
- Take note of entity associations. Throughout each experiments, efficiency modified when acknowledged corporations or specialists had been added to or faraway from listicles.
- Don’t deal with quotation quantity as the ultimate enterprise final result. Probably the most continuously cited sources weren’t at all times the strongest referral sources, and one low-volume Perplexity referral resulted in a paying buyer.
Conventional SEO focuses largely on rating pages. Each experiments indicated that LLM visibility relies upon extra closely on associations: which sources point out you, who seems alongside you, and the way just lately these relationships had been bolstered. The primary disagreement was the position of owned content material, which the second experiment confirmed was much less highly effective than I initially believed.
LLM visibility continues to evolve, nevertheless it’s by no means too early to construct on these experiments and see what your knowledge tells you.
Dig deeper: GEO for people who have to hit revenue targets
Contributing authors are invited to create content material for Search Engine Land and are chosen for his or her experience and contribution to the search neighborhood. Our contributors work below the oversight of the editorial staff and contributions are checked for high quality and relevance to our readers. Search Engine Land is owned by Semrush. Contributor was not requested to make any direct or oblique mentions of Semrush. The opinions they categorical are their very own.
#GEO #experiments #problem #standard #visibility #recommendation
