Konrad says: “My additional insight for SEO in 2025 is Keyword to Topic Modelling, using various unsupervised and supervised machine learning models.”
Superb. Okay, so a lot to unpack there. Shall we start at the beginning?
You started off with Keyword to Topic Modelling. What does that mean in practice?
“What it means is having your keyword datasets mapped out and clustered in a way that machines and AI agents would understand.
If we are talking about a supervised machine learning approach, like xbird or just fuzzy matching, or an unsupervised approach with BERTopic.
What I mean by that is, for example, in BERTopic (a model in machine learning), you can have a keyword dataset and cluster this keyword dataset into different topics. This way, your keyword dataset would be mapped out in a way that a typical human would find hard to do. So, it's essentially training your keyword dataset further.
If we have an enterprise website, we would take a product URL and a category URL and see, for those products and categories or collections, which keywords they rank for, and then we would try to identify semantic classes in that keyword dataset. That would be the general approach.
With Keyword to Topic Modelling, it's important to consider the objectives of the analysis. If we are just looking at a pure keyword dataset, a model like BERTopic would identify the topics from that keyword dataset. However, if you already have your topics or your collections and you are not looking to create new categories, you'd use another model.
That's a basic overview of this insight.”
How much human involvement does there have to be?
If you leave it all to the machines, surely the selection of topics to accompany keyword phrases isn't necessarily going to be as accurate or as relevant for that particular marketplace or business as a human might select, or am I not up to speed with what machines can do?
“To give you an example, what we would do in-house is scrape the keywords that a URL or a similar bunch of URLs would rank for. What we then do is apply Keyword to Topic Modelling (or you could phrase it as a ‘Topic Mapping Model’), like BERTopic.
We usually do this with longer queries. What I mean by longer queries is when you have long tail keywords – and keywords are queries; what I'm referring to is the same thing.
When we look at queries that users type into Google (those specific ones, which would be 4/5 or more words), we would take these keywords and we would put it into a CSV file, then we would upload it into our script, which is a model, and it basically gives us clusters or topics that are relevant to the website and what the keywords are about.
To give you an example, you could have a ‘Nike football shoes’ category, right? In a ‘Nike football shoes’ category, you would have URLs like ‘astro’, ‘grass’, and colours, like ‘blue’, ‘white’, etc. You have all of these URLs that essentially have filters on them. You would take those queries and try to explore the keywords that each URL would rank for, and we would put them into a CSV file.
Then we would put it into our model, and this model essentially generates the semantic topics that are associated with that keyword dataset.
How is this useful? We are essentially looking at which keywords belong to which topic cluster. It basically gives us a summary of the representative keywords. In plain English, this means that, if there is a keyword that we are not considering for a URL, we could insert it into a description or a meta title, if it's a high search volume keyword.
Again, this is a basic top-level overview. It’s essentially mapping keywords to their semantic clusters.”
It sounds like selecting the correct URLs to begin with is key.
How do you go about doing that? Is this something you can get in Google Search Console?
“If we're talking about first-party data, so Google Search Console, we can explore clicks, impressions, and different keywords.
If we're talking about third-party data, we can use a third-party tool to do that as well; just plug in the URL and explore the keywords.
It depends on the URL. If you're looking at one URL and just a few days, if it's an important URL, we will just use an API to get more queries, because Google Search Console is limited. It only gives you 30% of the data.”
Can you talk a little bit about other software that you use in order to do this?
“If we're talking about pure keyword data, we'll look at Semrush, simply because Semrush, for us, works on a keyword basis.
We would use Semrush to get the keywords and analyse them in our topic model, but we would also use a tool like Pi Datametrics, which would then be able to give additional metrics on top of it.
Why are we doing it? So, if it's a particular tracking solution we're looking at and, let's say, the objective is to identify new topics, we would then use Semrush to scrape the data and find the keywords.
Then, we would identify the classes/topics from this initial keyword dataset using a model like BERTopic. We would then plug in these topics and try to create new URLs or new topics based on that.
Using a tool like Pi Datametrics, we are able to see what positions are tracked over time and if there is a drop or improvement in organic search traffic based on our analysis and input.”
What tool do you use to map the intended topics to the keywords that you identify?
“So, this is a Python script in Google Colab.
It's not a publicly accessible tool, but if you do a good Google search, lots of SEO experts will provide the initial model.
What we do in-house is we train it so it works for each client. Essentially, what I mean by training is, we have a sample of the data and add additional functions or libraries to the initial script, tailoring the script for our use case.”
At what stage will you attempt to prioritise what has to be done/what URLs to focus on initially?
“When we look at prioritisation, we always try to go back to the initial objectives.
If it is an important category, like ‘Nike football boots’, then we would try to prioritise by the objective, the search volume and other SEO metrics, and the business outcome.
If we focus on a category that is in high demand – and our forecast tells us that, if we improve this category, we will be able to move the needle and essentially get more traffic and get more orders – we will prioritise those.
If we're looking at ‘Reebok trainers’ versus ‘Nike football boots’, we will prioritise based on search volume, based on the number of historical orders, and the financial figures. Essentially, we are prioritising based on the initial objective and what the business and the higher-ups care about.”
It was interesting that you used ‘get more traffic’ and ‘get more orders’ as the KPIs that you were focusing on. So, it's obviously easy to make some kind of educated guess in terms of the likely traffic that you would get from an increase in rankings, but an increase in orders is trickier.
Would you use an assumed click-through rate from someone visiting your site to making a purchase there, or would you attempt to get more granular and look at the relevance of the keyword phrase and assume a likely purchase rate based upon the type of keyword phrase that you're targeting?
“What you're referring to is third-party data.
When we engage with a client, we look at their CRM. In an enterprise environment, you have a CRM team, a conversion team, an SEO team, a paid search team, a social team, etc. In this case, we are collaborating directly with the CRM team, which has this first-party data. It's just easier than the anticipation aspect.
If it's a smaller client, and there’s a lack of data, we would look at anticipated positions, the click-through rate based on that position, and conversion rates, and we are able to anticipate and forecast the outcomes that the business cares about.
With different clients and in different scenarios, the approach is varied. With large enterprise, CRM teams provide the data, and we can work with that and combine it. With smaller clients, we’ll look at forecasting methodologies – using both historical first-party data and Search Console analytics, but also third-party data like Semrush and different third-party tools.
It's difficult to tell, but at the end of the day, it's essentially about identifying, using different decision trees, what we could rank for and when.
To elaborate on that, if we’re looking at a smaller client, we would try to get the keywords and the topics in the right order of association. We would then look at a forecasting method like, for example, if the keyword difficulty is low but the search volume is high.
It’s based on different decision trees and, based on these decisions, we would say, ‘This is the timeline we would be able to achieve this by.’ Then, looking at the rankings and our Google AIO dimensions.
Sometimes clicks is one of the KPIs, but at times we know that click-through rate has been dropped massively in the SEO industry, so a mention in the AIO is important as well.
I hope I answered your question.”
Yes, absolutely. You made me think of lots of follow-up questions, including the generation of content from the topics that you identify as well, and the trends that you're seeing from that.
But to be honest with you, that's going to be a whole other conversation, and perhaps we can have that conversation in a future episode.
“One hundred percent.”
Sounds good. Okay, so, Konrad, what's the key takeaway from the tip you shared today?
“Well, the key takeaway is looking at your keyword dataset, your topics that the model generates, and seeing what you could do with those associations of keywords and topics.
Can we add additional links to the topics (talking about internal links)? Can we create new content? Can we do something on a technical aspect that we are not doing?
In a nutshell, the key takeaway would be the objectives.
If we are increasing the sites and providing additional architecture elements, what model are we using? The objectives of the analysis, the terms of the model, and the type of keywords (long tail, short tail, etc.), that would be the key takeaway. It’s the objectives of the analysis.”
Konrad Szymaniak is Managing Enterprise SEO Consultant at Szymaniak Digital, and you can find him over at SzymaniakDigital.co.uk.