Majestic

  • Site Explorer
    • Majestic
    • Summary
    • Ref Domains
    • Backlinks
    • * New
    • * Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Author ExplorerBeta
    • Summary
    • Similar Profiles
    • Profile Backlinks
    • Attributions
  • Compare
    • Summary
    • Backlink History
    • Flow Metric History
    • Topics
    • Clique Hunter
  • Link Tools
    • My Majestic
    • Recent Activity
    • Reports
    • Campaigns
    • Verified Domains
    • OpenApps
    • API Keys
    • Keywords
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
    • Link Tools
    • Bulk Backlinks
    • Neighbourhood Checker
    • Submit URLs
    • Experimental
    • Index Merger
    • Link Profile Fight
    • Mutual Links
    • Solo Links
    • PDF Report
    • Typo Domain
    • TLD Checker New
  • Free SEO Tools
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
  • Support
    • Blog External Link
    • Support
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • Style Guide
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • About Backlinks and SEO
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Sign Up for FREE
  • Plans & Pricing
  • Login
  • Language flag icon
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文
  • Get started
  • Login
  • Plans & Pricing
  • Sign Up for FREE
    • Summary
    • Ref Domains
    • Map
    • Backlinks
    • New
    • Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Summary
      Pro
    • Backlink History
      Pro
    • Flow Metric History
      Pro
    • Topics
      Pro
    • Clique Hunter
      Pro
  • Bulk Backlinks
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
      Pro
  • Neighbourhood Checker
    Pro
    • Index Merger
      Pro
    • Link Profile Fight
      Pro
    • Mutual Links
      Pro
    • Solo Links
      Pro
    • PDF Report
      Pro
    • Typo Domain
      Pro
    • TLD Checker New
      Pro
  • Submit URLs
    • Summary
      Pro
    • Similar Profiles
      Pro
    • Profile Backlinks
      Pro
    • Attributions
      Pro
  • Custom Reports
    Pro
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • Site Updates
    • The Company
    • Style Guide
    • Terms & Conditions
    • Privacy Policy
    • GDPR
    • Contact Us
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Blog External Link
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文

All Generative AI tools rely on traditional search results to answer questions

James Hocking

James Hocking discusses the evolution of SEO and explains the shift towards quantitative, concise content over qualitative, lengthy answers.

 
James Hocking 2025 Additional Insights podcast cover with logo
« Back to Additional insights
More Additional Insights YouTube Podcast Playlist Link Spotify Podcast Playlist Link Audible Podcast Playlist Link

James says: “All the AI – Perplexity, Copilot, ChatGPT – they all use Bing as their web search engine.”

Why do they use Bing as their search engine, and what do SEOs need to do about this?

“Originally (because I've been following it for a number of years now), they were using Google. As Gemini came along, and Google introduced the AI Overview, they slowly banned these other AIs from using the Google AI engine, is my observation, and they've moved over to Bing.

Just to digress ever so slightly, all of the large language models – the generative AI – they're essentially a language engine. So, they understand spoken language rules really, really well, but they're very expensive to train, and not just in money, but in time. It's a very time-costly activity. So, in order to keep them up to date, they have to use another source, which is ultimately a search engine.

When you write in your query and you want, ‘What's happening in the news today? Can you summarise it?’, it has to get that from somewhere. So they rely on a search engine such as Bing.

On their own, all they have is what they learned six months ago or a year ago, depending on which one you're using.”

Why do Google not want these LLMs, these AI answer engines, to use their results? And is Bing likely to go on the same train and not wish their results to be used?

“I think it's a very interesting question, especially because, as recently as one week ago, third parties such as Cloudflare have introduced a new mechanism where you ‘pay-per-scrape’, in a way, for these engines.

I've got my thoughts rather than definitive facts. For me, it's a bit like if you have a robots.txt file, you might wonder why your site is not being scraped or being indexed in the traditional search, and you realise your robots.txt file is preventing it. With Google, since about the late 90s, they've owned around 98% market share of the search. Now it's dropped down to 90%.

Now, you could argue that some of that drop is the same people using Google but relying on the AI overviews that Google creates. My supposition, or my belief, is that Google are trying to own this space. They have done with search, and so they have a very rich capital of information in their search results.

Bing, on the other hand, even with the introduction of ChatGPT around 2023 (where Copilot became richer), in the desktop market, they only grew to about 12% of the desktop market share, and they have even less of the mobile market share.

My personal belief is that Bing is probably more open to being scraped, because they're still cementing their position as the de facto search engine.”

If we're assuming that Bing remains open to being scraped, what do SEOs need to do (if at all) slightly differently to rank on Bing over Google?

“When I've met with people, and when I've spoken with my customers, Bing is something where they say, ‘Do you use Bing?’, ‘No,’ ‘So, do you want to invest in optimizing for Bing, or do you want to optimize for Google?’ The preference tends to be Google.

I think the lessons we've learned for Google are completely transferable over to Bing. I mean, they make rule changes just like Google does so, depending on which type of business you are…

If you're a small business and your main catchment area is the local area that you're operating within, then it's things like getting your business profile on Bing (Places for Business) as good as it possibly could be, and as relevant as it possibly could be, just as you may have done for your Google Business Profile.

Similarly, when you want to get optimized for agentive engine optimization, how do you make sure that, say, ChatGPT or Perplexity recommends your business in an AI response over someone else's? Then, the type of information that you share has to be in a form that works well for that type of tool.

With the large language models, despite their ability to understand spoken language, the AI – fundamentally, at its core – likes numbers over text. When you start trying to describe yourself in a business context, you can then say, ‘I service customers with a market cap of 1,000,000 annualised revenue.’ That's a number that the AI can filter on, and it's very good at doing that.

When you get into a more textual, or what would be more of an essay response, whilst the AI is very good at summarising that, if you're trying to stand out from your competition, you then get lost in the weeds, so to speak. You really want to make sure that your Bing profile is as good as it can possibly be; your business profile is as good as it can possibly be.

The facts that say, ‘This is me as a business. This is my company registration number, these are the things that stand me out as being a real business as opposed to something that isn't.’ You will gain prominence in that form, and that's before you even start thinking about your content.

With a lot of these AIs, they're trying to say, ‘Look, we are a place where you can get good facts,’ and that's a good starting place for Bing.

Similarly, with your keywords and the way you would compete for those, it's a very similar process to what you've done with Google – it's just you're going to do it with Bing as well.”

Obviously, it's the Google Business Profile, but for Bing, I believe it's called Bing Places for Business.

“I couldn't comment. My background is data and getting data ready for AI – and I've been in large language models since around 2018/2019. I came into this by being knowledgeable in large language models and data, as opposed to pure SEO. I then ventured into SEO out of coincidence, because it became popular. I believe you're right, but I wouldn't like to say I'm a Bing SEO expert.”

Yes, absolutely. It's available over at BingPlaces.com, as somewhere where you can actually submit your business details to the Bing search engine.

With regards to the information that the tools that you were talking about (Copilot, ChatGPT, Perplexity) take from Bing, what specifically would you say that they're likely to take from Bing?

“There's a process called RAG (Retrieval Augmented Generation) and it's a technique that these tools use.

What it means is, ‘I'm going to go off to a number of different sources.’ The LLM knows nothing, at this point, about what's authoritative and things like that. So, what you will find is that, when you type a prompt into, say, Copilot, it will go off and perform a number of web search queries.

It won't do a single one, it will do a number, and what it will do is it will take your original question, and it will create different versions of that question – a bit like the semantic search that you get with traditional search – looking for similarities. It will go off and run, let's say, 10 separate queries. It doesn't really care so much about the ranking of the result, only that it's looking for things that are semantically similar to what your question is.

When it finds them, it will collect those, and it will then go into a process of summarising that in order to work out how to make the best answer. This is why, sometimes, when you get a citation in a response, it may or may not be a top-ranking result if you did the search yourself.”

When you say it doesn't have to be a top-ranking result, I assume that there's a certain cut-off.

Are we talking about it has to rank within the top 10 or 20 results, or is there any particular number that you can share?

“I've found it's variable. In some of our experiments that we've run, it's been up in the top 50 results, so it's gone through the first couple of pages.

More often than not, it sticks with the first page, that we've found, but it doesn't really care whether it's rank 1, rank 2, or rank 3. It's just looking, based on the prompt that was entered by the user: What is semantically similar? This is where you get things like the cosine similarity, which some SEOs may have performed, to see whether this piece of text is similar to that piece of text. If they are, then there's an assumption that the piece of text you’ve pulled out of the webpage is going to be relevant to the answer.

Now, I was going to say the biggest challenge is that, with a large language model, you may phrase the question one way to get the same result, and I may phrase it a slightly different way, looking for the same result. This is why keywords matter less than in traditional search. The way the large language model works is that the way it interprets the question could be different, even though, to you and me, the question is ultimately asking the same thing.

What these large language models are good at is generating sequences and predicting sequences. If you think of it that way then, if you take your input question as a sequence of words, what it predicts next is very dependent on the order in which you asked the question.

To be slightly facetious, if I ask the question in a slightly Yoda way, where I use words slightly backwards, the probability of what it would predict next would be different because the order in which the words were fed into the LLM was different.

Coming back to what you said – How do I know how many pages it will feed from? – it really depends on how many similarity matches it finds in the traditional search results.”

It's funny because people might have traditionally searched in a more Yoda-ish way, as you put it, in that they've perhaps used a shorter tail keyword phrase to begin with and added words to that to try and get more defined results, certainly.

The other question that I was going to ask you was related to whether or not you felt it was a good idea to have a different web page for every single query that you're attempting to answer on your website or whether it's just as effective for AI search engines to find the answer on a page that happens to have thousands of words on it.

“It's an interesting one. We're experimenting with both sides.

Now, the way a large language model works is that, believe it or not, a large language model – all AI – is deterministic. What I mean by that is that, given the same input, you could guarantee the same output. Now, lots of people are going to be screaming at the podcast, thinking that is nonsense, because if I send in the same query (even if you did it programmatically using Python), there is an attribute where you can change the randomness of the response.

The thing is, the large language model is wrapped with a piece of software, and the bit that wraps it is the bit that creates the randomness, not the model itself.

If you had one super massive webpage, the challenge is that, unless that webpage is organised in a very, very clear way, it's very difficult for the large language model. It's just going to work out which piece of text on this page is relevant. You may have two points on that page that could be quite far apart and, depending on which model you're using, it may be tuned to support that, or it may not.

Let me put that another way. Some large language models have been trained on webpages, so they understand HTML. They understand what a main tag is, they understand what an article tag is, and things like that. Other models haven't. They've been trained on just bodies of text, or corpuses of text. They don't understand HTML. Their ability to read HTML and pull out the salient bit is almost impossible. It will either not find it, or it will find it.

Certainly, with the way large language models read web pages, they don't understand JavaScript. So, if you've got a lot of your content embedded in JavaScript with mouseovers and doing these kinds of really funky things, the LLM will just think, ‘There's nothing here for me to read because I haven't executed JavaScript. I'm not going to do that.’ So, it will move on.

To your question – Should you have one supermassive page and one not-supermassive page? – the way you really want to be is you want a piece of content that is really on point.

In the past, we used to think, ‘How do I get that People Also Ask-type piece of text?’, when Google started saying, ‘Oh, this is the answer to your question. I pulled it from this web page.’ You want to think of it in a similar way because, given this question that the person's asked a large language model, it's going to perform that semantic similarity, and it’s going to say, ‘Which body of text is the same?’

Now, where in the web page should you put that piece of text? We're experimenting with that, and it's kind of changing but, for the most part, you want to put it in a really easy place so it doesn't have to scan through lots of text, it can just pull it from the top.

To your question, should you have just one long page? It really depends on what model you're using. Some like longer pages, some like shorter pages. ChatGPT operates differently from Perplexity, for example. However, last year, a new – it’s not really a standard yet, but it stands for llms.txt, and there's another one called llms-full.txt. These are pages which are formatted in Markdown, a kind of markup, and they're a way you can start organising your content in a way that is very targeted towards how the large language model will look.

The catch, at the moment, is that there's no large language model that's actively looking at these files. But, if you included them within your prompt, your customer would get to the response very fast, rather than relying on the large language model to try and interpret what your webpage looked like and how it was organised.”

James, what's the key takeaway from the tip you shared today?

“Right now, it's like the early days of the web.

If you get your business profile in place, and you get your content very clear and have it more quantitative-based – ‘I can save you X% by following my product’ – you're more likely to stand out, rather than doing a very qualitative, lengthy, and verbose answer.”

James Hocking is co-founder at Hocking Digital, and you can find him over at HockingDigital.com.

Choose Your Own Learning Style

Webinar iconVideo

If you like to get up-close with your favourite SEO experts, these one-to-one interviews might just be for you.

Watch all of our episodes, FREE, on our dedicated SEO in 2025: Additional Insights playlist.

youtube Playlist Icon

Podcast iconPodcast

Maybe you are more of a listener than a watcher, or prefer to learn while you commute.

SEO in 2025: Additional Insights is available now via all the usual podcast platforms

Spotify Apple Podcasts Audible

Book iconSEO in 2025

Catch up on SEO tips from 106 SEO experts in the original SEO in 2025 series

Available as a video series, podcast, and a book.

SEO in 2025

Could we improve this page for you? Please tell us

Fresh Index More info about Fresh Index icon

Unique URLs crawled 194,188,717,434
Unique URLs found 618,052,412,077
Date range 07 Apr 2026 to 05 Aug 2026
Last updated 1 hour ago

Historic Index More info about Fresh Index icon

Unique URLs crawled 4,502,566,935,407
Unique URLs found 21,743,308,221,308
Date range 06 Jun 2006 to 26 Mar 2024
Last updated 03 May 2024

SOCIAL

  • LinkedIn
  • YouTube
  • Facebook
  • Bluesky
  • Twitter / X

COMPANY

  • Blog External Link
  • About
  • Terms and Conditions
  • Privacy Policy
  • GDPR
  • Contact Us

TOOLS

  • Plans & Pricing
  • Site Explorer
  • Compare Domains
  • Bulk Backlinks
  • Search Explorer
  • Developer API External Link

MAJESTIC FOR

  • Trust Flow
  • Flow Metric Scores
  • Link Context
  • Backlink Checker
  • Influencer Discovery
  • Enterprise External Link

PODCASTS & PUBLICATIONS

  • The Majestic SEO Podcast
  • SEO in 2026
  • SEO in 2025
  • SEO in 2024
  • SEO in 2023
  • SEO in 2022
  • All Podcasts
top ^