Majestic

  • Site Explorer
    • Majestic
    • Summary
    • Ref Domains
    • Backlinks
    • * New
    • * Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Author ExplorerBeta
    • Summary
    • Similar Profiles
    • Profile Backlinks
    • Attributions
  • Compare
    • Summary
    • Backlink History
    • Flow Metric History
    • Topics
    • Clique Hunter
  • Link Tools
    • My Majestic
    • Recent Activity
    • Reports
    • Campaigns
    • Verified Domains
    • OpenApps
    • API Keys
    • Keywords
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
    • Link Tools
    • Bulk Backlinks
    • Neighbourhood Checker
    • Submit URLs
    • Experimental
    • Index Merger
    • Link Profile Fight
    • Mutual Links
    • Solo Links
    • PDF Report
    • Typo Domain
    • TLD Checker New
  • Free SEO Tools
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
  • Support
    • Blog External Link
    • Support
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • Style Guide
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • About Backlinks and SEO
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Sign Up for FREE
  • Plans & Pricing
  • Login
  • Language flag icon
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文
  • Get started
  • Login
  • Plans & Pricing
  • Sign Up for FREE
    • Summary
    • Ref Domains
    • Map
    • Backlinks
    • New
    • Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Summary
      Pro
    • Backlink History
      Pro
    • Flow Metric History
      Pro
    • Topics
      Pro
    • Clique Hunter
      Pro
  • Bulk Backlinks
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
      Pro
  • Neighbourhood Checker
    Pro
    • Index Merger
      Pro
    • Link Profile Fight
      Pro
    • Mutual Links
      Pro
    • Solo Links
      Pro
    • PDF Report
      Pro
    • Typo Domain
      Pro
    • TLD Checker New
      Pro
  • Submit URLs
    • Summary
      Pro
    • Similar Profiles
      Pro
    • Profile Backlinks
      Pro
    • Attributions
      Pro
  • Custom Reports
    Pro
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • Site Updates
    • The Company
    • Style Guide
    • Terms & Conditions
    • Privacy Policy
    • GDPR
    • Contact Us
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Blog External Link
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文

Use Conversational SEO and Query Data Clustering via machine learning

Konrad Szymaniak

Konrad Szymaniak discusses keyword-to-topic modeling in SEO 2025, emphasizing the use of machine learning models like BERT to cluster keywords into semantic topics.

@KonradSzymaniak  
Konrad Szymaniak 2025 Additional Insights podcast cover with logo
« Back to Additional insights
More Additional Insights YouTube Podcast Playlist Link Spotify Podcast Playlist Link Audible Podcast Playlist Link

Konrad says: “My additional insight for SEO in 2025 is Keyword to Topic Modelling, using various unsupervised and supervised machine learning models.”

Superb. Okay, so a lot to unpack there. Shall we start at the beginning?

You started off with Keyword to Topic Modelling. What does that mean in practice?

“What it means is having your keyword datasets mapped out and clustered in a way that machines and AI agents would understand.

If we are talking about a supervised machine learning approach, like xbird or just fuzzy matching, or an unsupervised approach with BERTopic.

What I mean by that is, for example, in BERTopic (a model in machine learning), you can have a keyword dataset and cluster this keyword dataset into different topics. This way, your keyword dataset would be mapped out in a way that a typical human would find hard to do. So, it's essentially training your keyword dataset further.

If we have an enterprise website, we would take a product URL and a category URL and see, for those products and categories or collections, which keywords they rank for, and then we would try to identify semantic classes in that keyword dataset. That would be the general approach.

With Keyword to Topic Modelling, it's important to consider the objectives of the analysis. If we are just looking at a pure keyword dataset, a model like BERTopic would identify the topics from that keyword dataset. However, if you already have your topics or your collections and you are not looking to create new categories, you'd use another model.

That's a basic overview of this insight.”

How much human involvement does there have to be?

If you leave it all to the machines, surely the selection of topics to accompany keyword phrases isn't necessarily going to be as accurate or as relevant for that particular marketplace or business as a human might select, or am I not up to speed with what machines can do?

“To give you an example, what we would do in-house is scrape the keywords that a URL or a similar bunch of URLs would rank for. What we then do is apply Keyword to Topic Modelling (or you could phrase it as a ‘Topic Mapping Model’), like BERTopic.

We usually do this with longer queries. What I mean by longer queries is when you have long tail keywords – and keywords are queries; what I'm referring to is the same thing.

When we look at queries that users type into Google (those specific ones, which would be 4/5 or more words), we would take these keywords and we would put it into a CSV file, then we would upload it into our script, which is a model, and it basically gives us clusters or topics that are relevant to the website and what the keywords are about.

To give you an example, you could have a ‘Nike football shoes’ category, right? In a ‘Nike football shoes’ category, you would have URLs like ‘astro’, ‘grass’, and colours, like ‘blue’, ‘white’, etc. You have all of these URLs that essentially have filters on them. You would take those queries and try to explore the keywords that each URL would rank for, and we would put them into a CSV file.

Then we would put it into our model, and this model essentially generates the semantic topics that are associated with that keyword dataset.

How is this useful? We are essentially looking at which keywords belong to which topic cluster. It basically gives us a summary of the representative keywords. In plain English, this means that, if there is a keyword that we are not considering for a URL, we could insert it into a description or a meta title, if it's a high search volume keyword.

Again, this is a basic top-level overview. It’s essentially mapping keywords to their semantic clusters.”

It sounds like selecting the correct URLs to begin with is key.

How do you go about doing that? Is this something you can get in Google Search Console?

“If we're talking about first-party data, so Google Search Console, we can explore clicks, impressions, and different keywords.

If we're talking about third-party data, we can use a third-party tool to do that as well; just plug in the URL and explore the keywords.

It depends on the URL. If you're looking at one URL and just a few days, if it's an important URL, we will just use an API to get more queries, because Google Search Console is limited. It only gives you 30% of the data.”

Can you talk a little bit about other software that you use in order to do this?

“If we're talking about pure keyword data, we'll look at Semrush, simply because Semrush, for us, works on a keyword basis.

We would use Semrush to get the keywords and analyse them in our topic model, but we would also use a tool like Pi Datametrics, which would then be able to give additional metrics on top of it.

Why are we doing it? So, if it's a particular tracking solution we're looking at and, let's say, the objective is to identify new topics, we would then use Semrush to scrape the data and find the keywords.

Then, we would identify the classes/topics from this initial keyword dataset using a model like BERTopic. We would then plug in these topics and try to create new URLs or new topics based on that.

Using a tool like Pi Datametrics, we are able to see what positions are tracked over time and if there is a drop or improvement in organic search traffic based on our analysis and input.”

What tool do you use to map the intended topics to the keywords that you identify?

“So, this is a Python script in Google Colab.

It's not a publicly accessible tool, but if you do a good Google search, lots of SEO experts will provide the initial model.

What we do in-house is we train it so it works for each client. Essentially, what I mean by training is, we have a sample of the data and add additional functions or libraries to the initial script, tailoring the script for our use case.”

At what stage will you attempt to prioritise what has to be done/what URLs to focus on initially?

“When we look at prioritisation, we always try to go back to the initial objectives.

If it is an important category, like ‘Nike football boots’, then we would try to prioritise by the objective, the search volume and other SEO metrics, and the business outcome.

If we focus on a category that is in high demand – and our forecast tells us that, if we improve this category, we will be able to move the needle and essentially get more traffic and get more orders – we will prioritise those.

If we're looking at ‘Reebok trainers’ versus ‘Nike football boots’, we will prioritise based on search volume, based on the number of historical orders, and the financial figures. Essentially, we are prioritising based on the initial objective and what the business and the higher-ups care about.”

It was interesting that you used ‘get more traffic’ and ‘get more orders’ as the KPIs that you were focusing on. So, it's obviously easy to make some kind of educated guess in terms of the likely traffic that you would get from an increase in rankings, but an increase in orders is trickier.

Would you use an assumed click-through rate from someone visiting your site to making a purchase there, or would you attempt to get more granular and look at the relevance of the keyword phrase and assume a likely purchase rate based upon the type of keyword phrase that you're targeting?

“What you're referring to is third-party data.

When we engage with a client, we look at their CRM. In an enterprise environment, you have a CRM team, a conversion team, an SEO team, a paid search team, a social team, etc. In this case, we are collaborating directly with the CRM team, which has this first-party data. It's just easier than the anticipation aspect.

If it's a smaller client, and there’s a lack of data, we would look at anticipated positions, the click-through rate based on that position, and conversion rates, and we are able to anticipate and forecast the outcomes that the business cares about.

With different clients and in different scenarios, the approach is varied. With large enterprise, CRM teams provide the data, and we can work with that and combine it. With smaller clients, we’ll look at forecasting methodologies – using both historical first-party data and Search Console analytics, but also third-party data like Semrush and different third-party tools.

It's difficult to tell, but at the end of the day, it's essentially about identifying, using different decision trees, what we could rank for and when.

To elaborate on that, if we’re looking at a smaller client, we would try to get the keywords and the topics in the right order of association. We would then look at a forecasting method like, for example, if the keyword difficulty is low but the search volume is high.

It’s based on different decision trees and, based on these decisions, we would say, ‘This is the timeline we would be able to achieve this by.’ Then, looking at the rankings and our Google AIO dimensions.

Sometimes clicks is one of the KPIs, but at times we know that click-through rate has been dropped massively in the SEO industry, so a mention in the AIO is important as well.

I hope I answered your question.”

Yes, absolutely. You made me think of lots of follow-up questions, including the generation of content from the topics that you identify as well, and the trends that you're seeing from that.

But to be honest with you, that's going to be a whole other conversation, and perhaps we can have that conversation in a future episode.

“One hundred percent.”

Sounds good. Okay, so, Konrad, what's the key takeaway from the tip you shared today?

“Well, the key takeaway is looking at your keyword dataset, your topics that the model generates, and seeing what you could do with those associations of keywords and topics.

Can we add additional links to the topics (talking about internal links)? Can we create new content? Can we do something on a technical aspect that we are not doing?

In a nutshell, the key takeaway would be the objectives.

If we are increasing the sites and providing additional architecture elements, what model are we using? The objectives of the analysis, the terms of the model, and the type of keywords (long tail, short tail, etc.), that would be the key takeaway. It’s the objectives of the analysis.”

Konrad Szymaniak is Managing Enterprise SEO Consultant at Szymaniak Digital, and you can find him over at SzymaniakDigital.co.uk.

Choose Your Own Learning Style

Webinar iconVideo

If you like to get up-close with your favourite SEO experts, these one-to-one interviews might just be for you.

Watch all of our episodes, FREE, on our dedicated SEO in 2025: Additional Insights playlist.

youtube Playlist Icon

Podcast iconPodcast

Maybe you are more of a listener than a watcher, or prefer to learn while you commute.

SEO in 2025: Additional Insights is available now via all the usual podcast platforms

Spotify Apple Podcasts Audible

Book iconSEO in 2025

Catch up on SEO tips from 106 SEO experts in the original SEO in 2025 series

Available as a video series, podcast, and a book.

SEO in 2025

Could we improve this page for you? Please tell us

Fresh Index More info about Fresh Index icon

Unique URLs crawled 194,188,717,434
Unique URLs found 618,052,412,077
Date range 07 Apr 2026 to 05 Aug 2026
Last updated 1 hour 2 minutes ago

Historic Index More info about Fresh Index icon

Unique URLs crawled 4,502,566,935,407
Unique URLs found 21,743,308,221,308
Date range 06 Jun 2006 to 26 Mar 2024
Last updated 03 May 2024

SOCIAL

  • LinkedIn
  • YouTube
  • Facebook
  • Bluesky
  • Twitter / X

COMPANY

  • Blog External Link
  • About
  • Terms and Conditions
  • Privacy Policy
  • GDPR
  • Contact Us

TOOLS

  • Plans & Pricing
  • Site Explorer
  • Compare Domains
  • Bulk Backlinks
  • Search Explorer
  • Developer API External Link

MAJESTIC FOR

  • Trust Flow
  • Flow Metric Scores
  • Link Context
  • Backlink Checker
  • Influencer Discovery
  • Enterprise External Link

PODCASTS & PUBLICATIONS

  • The Majestic SEO Podcast
  • SEO in 2026
  • SEO in 2025
  • SEO in 2024
  • SEO in 2023
  • SEO in 2022
  • All Podcasts
top ^