Majestic

  • Site Explorer
    • Majestic
    • Summary
    • Ref Domains
    • Backlinks
    • * New
    • * Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Author ExplorerBeta
    • Summary
    • Similar Profiles
    • Profile Backlinks
    • Attributions
  • Compare
    • Summary
    • Backlink History
    • Flow Metric History
    • Topics
    • Clique Hunter
  • Link Tools
    • My Majestic
    • Recent Activity
    • Reports
    • Campaigns
    • Verified Domains
    • OpenApps
    • API Keys
    • Keywords
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
    • Link Tools
    • Bulk Backlinks
    • Neighbourhood Checker
    • Submit URLs
    • Experimental
    • Index Merger
    • Link Profile Fight
    • Mutual Links
    • Solo Links
    • PDF Report
    • Typo Domain
    • TLD Checker New
  • Free SEO Tools
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
  • Support
    • Blog External Link
    • Support
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • Style Guide
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • About Backlinks and SEO
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Sign Up for FREE
  • Plans & Pricing
  • Login
  • Language flag icon
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文
  • Get started
  • Login
  • Plans & Pricing
  • Sign Up for FREE
    • Summary
    • Ref Domains
    • Map
    • Backlinks
    • New
    • Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Summary
      Pro
    • Backlink History
      Pro
    • Flow Metric History
      Pro
    • Topics
      Pro
    • Clique Hunter
      Pro
  • Bulk Backlinks
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
      Pro
  • Neighbourhood Checker
    Pro
    • Index Merger
      Pro
    • Link Profile Fight
      Pro
    • Mutual Links
      Pro
    • Solo Links
      Pro
    • PDF Report
      Pro
    • Typo Domain
      Pro
    • TLD Checker New
      Pro
  • Submit URLs
    • Summary
      Pro
    • Similar Profiles
      Pro
    • Profile Backlinks
      Pro
    • Attributions
      Pro
  • Custom Reports
    Pro
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • Site Updates
    • The Company
    • Style Guide
    • Terms & Conditions
    • Privacy Policy
    • GDPR
    • Contact Us
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Blog External Link
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文

Crawling and indexing management needs to be a corner stone of your SEO strategy

James McLoughlin

James McLoughlin shares that crawling and indexing management needs to be a corner stone of your SEO strategy in 2026, ensuring crawlers (both traditional search engine crawlers and LLM crawlers) can quickly and easily discover, understand and index any content you want surfaced in search engines and LLMs.

Website  
James McLoughlin 2026 Additional Insights podcast cover with logo
« Back to Additional insights
More Additional Insights YouTube Podcast Playlist Link Spotify Podcast Playlist Link Audible Podcast Playlist Link

James says: “Efficient crawling and indexing management has to be a cornerstone of your SEO strategy in 2026.

That's ensuring that crawlers – both the traditional search engine crawlers and the crawlers from the large language model tools – can quickly and easily discover, understand, and index any content that you want surfaced in the search engines or the LLMs.”

Okay. So, what does crawling and indexing management mean in practice in 2026?

“It's basically making sure that anything that you want to appear is technically crawlable. So, no disallow rules, no server-side restrictions on what crawlers can access, your links are crawlable (proper HTML links, not JavaScript buttons or any other uncrawlable elements), making sure that your important URLs are clearly prioritised (your non-relevant ones are excluded from crawling as far as possible), and that you've got your indexing signals and directives properly set up – so you're telling search engines what you want to appear in the index and what you don't want to appear.

This is absolutely essential because you can't appear in the search engine's index and you can't be part of the LLM's training data if the crawlers can't find, crawl, index, and effectively understand your content.”

It sounds like it could be a full-time job almost to do that. If you're just a single SEO working in an organisation, how much time should you spend on crawling and indexing management?

“It needs to be part of your regular workflows. It's something that you can monitor and act on when you see red flags come up. It definitely can be a full-time job, especially with big organisations that have hundreds of thousands, or even millions, of URLs, but it's something that we can at least monitor and jump on when we see a problem – mainly through regular, scheduled crawls.

If we have a baseline for how many URLs we'd expect the crawlers to be able to discover on a normal crawl of our website, and we have a baseline of how many URLs from that crawl should be indexable to the search engines, if we see any significant fluctuation in that number (so, if we all of a sudden see a big increase in the number of discovered URLs in a crawl, or we see a big change in the number of indexable URLs in a crawl), we can react to that. We can try to diagnose the issue, create tickets, ask developers to fix it, and do whatever needs to be done.

Keeping an eye on that data – keeping an eye on that crawl data, that indexing data – as long as that's part of your regular workflows (assuming that you don't already have any huge problems on the website, and you're in a good place), you can keep an eye on that and you can monitor that through regular crawling, and just make changes to that strategy if red flags pop up from your crawl data.”

What does regular workflow actually mean? Does it mean every day, every week, or a couple of hours here and there?

“I would say you need to have your weekly crawl set up. Now, if you've got an absolutely huge website, that might not be practical. If you've got a website with millions of URLs, it can sometimes take days for that crawl to run.

As long as you've got a sample crawl, or a portion crawl (which a lot of the modern cloud crawlers allow you to do), you can crawl a subsection of your website and get a flavour of what's happening on your website, from a portion.

However, you need to have some kind of weekly (and if not weekly, monthly) crawl running, just so you can see the fluctuations in your crawled and indexed URL data.”

Any particular software that you favour using at the moment for this?

“I usually use Screaming Frog and Botify, but most of the cloud crawlers – pretty much all of the cloud crawlers – will allow scheduling.”

Screaming Frog versus Botify: would you use one or the other, depending on the size of the site?

“Botify (or any of the cloud crawlers) are usually better for large websites, just because they do run cloud-based, so they're not going to affect your website. If you use Screaming Frog for an absolutely massive crawl, it's going to start slowing down your laptop significantly, and you're not going to be able to do much in the background.

Definitely, if you're running big crawls, you want to be using tools like Botify or whatever cloud crawler it might be.

The other big benefit of them is that they have a lot of these alerts built into them. So, you'll be able to see, from the headline stats from your crawl, what big changes there have been in terms of discovered URLs, indexed URLs, changes in status codes, and things like that.”

That's what tools to use and how often to use them, but how do you ensure that crawlers can quickly and easily discover, understand, and index any content that you want to be surfaced?

“You've just got to make sure that there are no technical blockers that would prevent crawlers from accessing your content. No stray disallow rules in your robots.txt file. Your links are crawlable. You've got to make sure that as much content as possible is visible in the source HTML, and that's particularly important with the advent of LLMs.

A lot of the LLMs and the answer engines have bots that aren't as sophisticated as Googlebot and can't process a lot of JavaScript content to the extent that Google and Google crawlers can. You need to make sure that any of your SEO and index-relevant content is discoverable in the source HTML, just so that it's visible for any plain HTML crawlers that are visiting your website.

In terms of indexing, you've just got to make sure that all of your indexing signals and directives are properly set up so you're only allowing indexing of content that you'd want a user to be able to find via the search engines.

On the flip side, it's not just about making sure that Google can crawl and access the stuff that you want them to crawl and index. You also need to make sure that you're preventing crawling of anything that is not useful for the search engines to crawl, and not relevant for users to land on via the search engines.

For example, I work with a publishing client who had a tonne of product reviews on their websites. On each of the websites, they had a page where you could apply filters to find the relevant reviews that you want. You could filter for the type of product, the star rating, and certain other characteristics. They had a robots.txt rule that was blocking the crawling of these filter URLs, and the links to the filters were also removed in the rendered version of the page.

However, the links to these faceted navigations, these filtering URLs, were still present in the source HTML and only removed at rendering. Also, the disallow rule in the robots.txt file had been malformed. It wasn't written correctly, so Googlebot was crawling hundreds of thousands of these low-value, low-quality URLs that just had different lists of product reviews with different filters applied.

If Googlebot can access that kind of content, you're wasting crawl budget on URLs that have absolutely no value for SEO, and that you don't need Googlebot to be crawling. It's just using and wasting time that Googlebot could be spending crawling your good quality content: your articles, your product reviews, stuff like that.

If you're not checking this stuff regularly, and if you aren't making that crawling and indexing management a focal part of your SEO strategy, you really run the risk of situations like that – where you're just wasting crawl budget on low-quality URLs, or you've got thin content or duplicate content popping up in the index. That's going to cause big problems for your overall SEO strategy further down the line.”

If you find a significant issue like the one you highlighted there, what's a significant quick fix that you could do to prevent this from happening quickly?

“With this one, because Google had discovered so many of these URLs and indexed a bunch of them before we could apply the correct robots.txt rule, we first had to noindex all of these pages to wait for them to drop out of the index, and then we could reapply the robots.txt rule correctly so Google would stop crawling them.

When you're using your crawling tools, you've got to be looking out for stuff like this – URLs that are appearing in your crawl that don't look like they have SEO value – and that will help you identify these issues. Then it's just working with the development teams or whoever it is who has control over the robots.txt file to make sure that your robots.txt files are correctly applied.

You can use tools like the robots.txt tester at TechnicalSEO.com, which you can test and apply robots.txt rules to and check that anything that you do apply is actually going to work and block what you want it to block before you put it into live. That'll give you the peace of mind that, when you do put that rule to prod in your robots.txt file, it's actually going to do what you want it to do.”

That's handling the content that you don't want to be discovered/surfaced. How do you actually define the content that you do want to be surfaced, and how do you make it more likely for it to be surfaced?

“There are a few things. As I mentioned before about the links between content, if you're using links that are reliant on rendering or JavaScript buttons, that's going to make it more difficult (or even impossible) for crawlers, depending on the crawler, to discover the content that's linked to.

Making sure that your links are proper HTML links, and that they're discovered in the source HTML, is going to make it easier for crawlers to find your content. Also, just making sure that stuff isn't buried too deep in the website. Anything that takes the crawler a long time, and it has to go down a lot of different levels to discover that content, is less likely to be discovered quickly. Making sure that there isn't too big a crawl depth to any content that you want search engine crawlers to be able to find and index.

Also, this is where crawling and indexing management and efficiency come into it. If you're wasting your crawl budget on low-value URLs, it makes it less likely that search engines will find the stuff that you want them to crawl, because they'll have spent a significant amount of time crawling low-value and low-quality thin content.

Link graph matches between mobile and desktop. I've worked on sites in the past where they'll have a bigger navigation for desktop because users obviously have a bigger screen, and they can see more links. Then, some of those links will be removed in the mobile version because they're dynamically serving two sets of code to mobile user agents and desktop user agents. However, because the crawlers are generally using a mobile user agent, they're not going to see any links that are only found in your desktop version.

Just making sure that any links to content that you want the search engines to find are included in the mobile version of your page, and not just in the desktop version of the page, is also really important. I've seen that quite a few times, where a link will be removed from the mobile version of the page, and that effectively makes that link invisible to the search engines.”

What are the main differences between crawling and indexing for traditional search engines versus LLMs?

“The big one is that a lot of the crawl activity from LLM crawlers is not rendering JavaScript. Any content that you have that relies on JavaScript rendering to be able to discover that content will either not be discovered by the LLMs or it won't be discovered as quickly.

This is why I say that anything that you want to be able to appear in AI overviews or any of the large language model search results has to be in your source HTML. They just aren't as sophisticated as Google when it comes to crawling JavaScript content. That's the big one.

This is also why you need to make sure that your crawling strategy is efficient. If you're asking the LLM crawlers to crawl stuff that just isn't relevant for their training data (so, if you're asking them to crawl sorting and filtering URLs, if they're crawling duplicate parameter pages, and things like that), it's making them work harder to discover the stuff that's actually useful for their training data.

You need to be excluding (or hiding as far as possible) any of this non-crawling-relevant stuff from not just the classic search engine crawlers, but also the LLMs crawlers.”

Finally, I'm thinking about crawling and indexing in terms of how that feeds into a content marketing strategy.

How do you take insight from what you discover with your crawling and indexing analysis and assist with determining what your content strategy should be for the coming few months?

“Crawling can help with that, because your crawler will help you identify potential gaps in your content.

If you crawled your website and there are certain topics that haven't been discovered by the crawler, that can do two things.

First of all, if that content does exist somewhere on your website, but the crawler hasn't found it, that's a big problem for your content strategy. If you have a page that covers that topic, but it's not discoverable for crawlers, then it's probably going to rank worse than if it were linked internally and discoverable – or maybe not at all.

If that topic doesn't exist, then that means, as a topic that you need to be writing about, you need to include it in your content strategy. The crawling of your website can also help you with things like that.

Tools like Screaming Frog have now allowed you to link your crawler with OpenAI, Gemini, or any of the other LLMs, so you can use the crawler to identify internal linking opportunities and potential content opportunities.

It can identify where you've got two pages that are already semantically quite similar. If you have a couple of pages that are potentially competing and about the same topic, you know that your content strategy doesn't necessarily need to include a page that covers that topic, because you've already got a couple of pages about that topic.

Actually, you should be looking to consolidate the similar content and create one page that's about that topic, which is your primary canonical page about that topic and can rank for whatever searches users are making around that topic. It’s ensuring that you've got a good oversight of your content, and you're using the crawlers to identify the entire graph of your website.

It's really important for informing your content strategy, to make sure that you're writing content about the topics that are important to users and the topics that you haven't got content about already.”

James, what's the key takeaway from the tip you shared today?

“If you aren't taking crawling efficiency as a big part of your strategy, then you're potentially going to miss out.

You're going to make the search engine crawlers and the LLM crawlers work harder to find your good content, and they're going to reward websites that are making it easier to find that content.

If you're not using your scheduled crawls and working with your development teams and your tech teams to make your website as efficient and easily crawlable as possible, then you're going to be falling behind the competition in 2026.”

James McLoughlin is a Technical SEO Specialist at EE. Find out more over at EE.co.uk.

Choose Your Own Learning Style

Webinar iconVideo

If you like to get up-close with your favourite SEO experts, these one-to-one interviews might just be for you.

Watch all of our episodes, FREE, on our dedicated SEO in 2026: Additional Insights playlist.

youtube Playlist Icon

Podcast iconPodcast

Maybe you are more of a listener than a watcher, or prefer to learn while you commute.

SEO in 2026: Additional Insights is available now via all the usual podcast platforms

Spotify Apple Podcasts Audible

Book iconSEO in 2026

Catch up on SEO tips from 117 SEO experts in the original SEO in 2026 series

Available as a video series, podcast, and a book.

SEO in 2026

Could we improve this page for you? Please tell us

Fresh Index More info about Fresh Index icon

Unique URLs crawled 194,188,717,434
Unique URLs found 618,052,412,077
Date range 07 Apr 2026 to 05 Aug 2026
Last updated 1 hour 2 minutes ago

Historic Index More info about Fresh Index icon

Unique URLs crawled 4,502,566,935,407
Unique URLs found 21,743,308,221,308
Date range 06 Jun 2006 to 26 Mar 2024
Last updated 03 May 2024

SOCIAL

  • LinkedIn
  • YouTube
  • Facebook
  • Bluesky
  • Twitter / X

COMPANY

  • Blog External Link
  • About
  • Terms and Conditions
  • Privacy Policy
  • GDPR
  • Contact Us

TOOLS

  • Plans & Pricing
  • Site Explorer
  • Compare Domains
  • Bulk Backlinks
  • Search Explorer
  • Developer API External Link

MAJESTIC FOR

  • Trust Flow
  • Flow Metric Scores
  • Link Context
  • Backlink Checker
  • Influencer Discovery
  • Enterprise External Link

PODCASTS & PUBLICATIONS

  • The Majestic SEO Podcast
  • SEO in 2026
  • SEO in 2025
  • SEO in 2024
  • SEO in 2023
  • SEO in 2022
  • All Podcasts
top ^