Majestic

  • Site Explorer
    • Majestic
    • Summary
    • Ref Domains
    • Backlinks
    • * New
    • * Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Author ExplorerBeta
    • Summary
    • Similar Profiles
    • Profile Backlinks
    • Attributions
  • Compare
    • Summary
    • Backlink History
    • Flow Metric History
    • Topics
    • Clique Hunter
  • Link Tools
    • My Majestic
    • Recent Activity
    • Reports
    • Campaigns
    • Verified Domains
    • OpenApps
    • API Keys
    • Keywords
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
    • Link Tools
    • Bulk Backlinks
    • Neighbourhood Checker
    • Submit URLs
    • Experimental
    • Index Merger
    • Link Profile Fight
    • Mutual Links
    • Solo Links
    • PDF Report
    • Typo Domain
    • TLD Checker New
  • Free SEO Tools
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
  • Support
    • Blog External Link
    • Support
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • Style Guide
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • About Backlinks and SEO
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Sign Up for FREE
  • Plans & Pricing
  • Login
  • Language flag icon
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文
  • Get started
  • Login
  • Plans & Pricing
  • Sign Up for FREE
    • Summary
    • Ref Domains
    • Map
    • Backlinks
    • New
    • Lost
    • Context
    • Anchor Text
    • Pages
    • Topics
    • Link Graph
    • Related Sites
    • Advanced Tools
    • Summary
      Pro
    • Backlink History
      Pro
    • Flow Metric History
      Pro
    • Topics
      Pro
    • Clique Hunter
      Pro
  • Bulk Backlinks
    • N-grams Near Links
    • Keyword Checker
    • Search Explorer
      Pro
  • Neighbourhood Checker
    Pro
    • Index Merger
      Pro
    • Link Profile Fight
      Pro
    • Mutual Links
      Pro
    • Solo Links
      Pro
    • PDF Report
      Pro
    • Typo Domain
      Pro
    • TLD Checker New
      Pro
  • Submit URLs
    • Summary
      Pro
    • Similar Profiles
      Pro
    • Profile Backlinks
      Pro
    • Attributions
      Pro
  • Custom Reports
    Pro
    • Get started
    • Backlink Checker
    • Majestic Million
    • Browser Plugins
    • Google Sheets
    • Post Popularity
    • Social Explorer
    • Get started
    • Tools
    • Subscriptions & Billing
    • FAQs
    • Glossary
    • How To Videos
    • API Reference Guide External Link
    • Contact Us
    • Site Updates
    • The Company
    • Style Guide
    • Terms & Conditions
    • Privacy Policy
    • GDPR
    • Contact Us
    • SEO in 2026
    • The Majestic SEO Podcast
    • All Podcasts
    • What is Trust Flow?
    • Link Building Guides
  • Blog External Link
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
    • 日本語
    • Nederlands
    • Polski
    • Português
    • 中文

What Data Sources Should You Be Feeding LLMs?

Andrew Melnychuk-Oseen

Andrew Melnychuk-Oseen shares that for effective SEO in 2026, you should be aware of what data sources you're feeding LLMs.

Website  
Andrew Melnychuk-Oseen 2026 Additional Insights podcast cover with logo
« Back to Additional insights
More Additional Insights YouTube Podcast Playlist Link Spotify Podcast Playlist Link Audible Podcast Playlist Link

Andrew says: “Adopt AI agents in your workflow.

You've probably heard about them with things like Claude Code, plugging in with Ahrefs, Semrush, all your data sources – and this is kind of the product I make.

What this does is you have one tool that pulls all your data into a chat. If you want to do things like keyword research and you have a lot of data to go through, these chatbots (ChatGPT, Anthropic’s Claude, etc.) can have direct access to your data. This really reduces hallucinations and allows you to move a lot faster.

For example, if you want to do your standard SEO audit, you're going through a lot of tools, and you have to pull a lot of data from different sources. Well, if you have your AI agent plugged into these things, you can pull all that data immediately, and it'll know what to pull in. Then, you can actually have an insight into the massive amounts of data from all your different sources.

It saves you a tonne of time getting data from all your different sources, like Google Search Console, for example. There's a lot of analysis work to do. Well, the agents are very, very good at it, and they can save you a lot of time now.”

Are you suggesting that SEOs train up their own specific AI agents and they have each agent focussing on one particular task?

“When you say training, I would say, no, you don't need to do that. Training is actually getting a model and fine-tuning the weights. You don't need to do that. What you need is instructions to say how you would actually go about your work in the context of the client. Once you put that in as a set of instructions and you have it hooked up to the tools, the AI agent will go into Semrush, Google Analytics, whatever you want, and then complete your task.

Now, I've talked to my clients about this and how they operate: You still need a person behind the wheel. There's a lot of talk saying, ‘AI is going to replace people.’ It's absolutely not. You really have to look at an AI agent like a supercharged librarian. It just knows where all the information is and where it can get it.

Frankly, it can give you an average response, but I really don't recommend it. You still need your professional expertise for going over this data. You can just get your data quicker, and then you can assemble a report the way you want a lot faster to deliver to clients. You can save more than half the time doing your normal tasks if you do it this way.”

Obviously, you can use APIs, or you can use third-party tools to automatically take the data into one source.

Are you suggesting that AI agents should have full access to the individual pieces of software that you use to access the data from?

“You can do it in a number of ways. What we do is we have an agent that creates agents, right?

There's something in working with AI called ‘context poisoning’. If you're doing, let’s say, a content audit and a standard SEO intro audit, you don’t want to do those things in the same step. You will want to create two different agents that do those two different things, and that maybe have two separate tools, so they can focus. Essentially, you wouldn't do those things at the same time; you do them one step at a time. That's how you structure your agents as well.

You can build an agent that specifically has the tools that you need to get the job done. Say you're doing backlink prospecting for what a competitor is using, or you're doing competitor analysis, and maybe you just want organic search results and something like Firecrawl to scrape their site. You can keep things limited like that, and then you get better results.

That's the job. It's a new skill set, much like how literacy started. In 1450, when we got the printing press, everybody needed to learn how to read and write. Well, for the future economy, everybody needs to know how to do what I call ‘probability field shaping’ with these LLMs to get the results that you want out of them.”

You mentioned Google Search Console and a few other data sources.

What would you say are the top data sources that SEOs absolutely need to be utilising and providing their agents with?

“Oh, dude, it's Google Search Console, because that's first party, right? That's the best concrete data for the site that you're working on.

It's genuinely amazing when you hook an agent up to these things, even when you supplement it with Semrush and Ahrefs – and of course, a scraper like Firecrawl, so the agent can actually read what's on the website.

When you have a scraper, and you have Google Search Console, it's genuinely wild to see an agent work with these different data sources. It's really a magic moment.”

What initial instructions should you be giving these agents to try and ensure that they're interpreting the data in the optimal possible way for your particular business?

“Well, that's context-dependent. I was talking to clients about this, and this is where it comes down to the job. So, actually, I don't have an answer for this because it's all different. Every context is different.

You can say, at the very base, ‘Hey, go into Google Search Console and get this data.’ There isn't really an answer, because that's where the expertise comes in. You, as a professional, have to build those instructions. But once you do, they're kind of set in stone, and you can reuse them if you need to go back to them.

What I predict is that, when people are using this and professionals are using this, they're going to have a specific agent built for their client that's using their specific data. You're not going to have one agent that does it all. You're going to need them specialised to the client.”

When you talk about these agents, obviously you have conventional LLMs: everyone's using Gemini, Claude, or ChatGPT.

Are there any go-to LLMs or types of agents that you would recommend?

“When you say, ‘type of agent’, agents are really an LLM that has a simple wrapper around it that does a tool call. It's called ‘tool calling’. That allows it to reach into Google Search Console, Ahrefs, Semrush, Google Keyword Planner, or whatever it is; it can call a computer program and get the results back. Much like how you go into those services and get the data yourself, the agent can do that.

Agents are really this one general thing. They're just an agent that has tool calling, and they're going to have different tools. It's kind of remarkable. I would look at an agent as: what the web browser was for the internet, the agent is for the LLM. It allows you to browse the probability field and shape it to get your answers a lot faster.

An agent is just a thing. The LLMs you're going to adjust based on your cost and your task. This is where the skill comes in for running it. Sometimes you don't need Opus 4.6 (which is really expensive and really good) to do a task. If you just want to scrape a competitor's website and get a content audit on them, you just need a cheaper model to get that data for you. Then, once you get that data, you switch to Claude Opus, and you're like, ‘Okay, what insights are here?’ Then, the frontier models are really good at analysis.

Part of the new job is picking the right model for the right job. An agent harness wraps any model, so you can flip them out. It doesn't really matter about the model so much. Actually, the agents can pick. You can tell them to pick their model, and they're going to just switch their brain. Imagine you could just pop your brain out and switch your brain with another brain. That's what these agents can do.”

I love that piece of advice: get the agent to select the right model, because obviously there are quite a few different models out there.

“Yeah, like 150.”

Wow, yeah, and the go-to thing, initially, for people getting started is probably just to select the most expensive/best model, but that's not required.

“And you're wasting money, because we're going to a usage-based world.

Even learning how to set up a local model – it’s going to be slower, but it's not costing you money.”

What about hallucinations? Is that still a problem? It used to be talked about a lot, maybe a couple of years ago.

Will these agents/LLMs actually hallucinate based on your own data?

“So that comes from the compression of the LLM. Basically, there's this overlap in the actual tokens – in what is essentially the hyperdimensional space of the LLM. Those aren't really going away, so you have to come up with techniques to reduce them.

Pulling in live data from places like Google Search Console, Ahrefs, Semrush, etc., really reduces them. You still need to check. You still need to be a professional. You’ve got to check your data source. Did it copy?

There are techniques that agents have, like we have a chain of verification. We have some special tools that reduce hallucinations by 98%, and they're pretty straightforward, but it's just a step.

You have your agent go pull some data. Then what you do is you have another agent say, ‘Hey, can you construct questions about this data?’ Then you have these two agents that go back and forth. They generate the questions, and ask the question, then the agent will answer the question. You create these verifiers, and then the agent will verify.

That strengthens your context, and then you get significantly fewer hallucinations, but you don't want to get lazy. What our tool does is it pulls the live data in so you can look, and you can say, ‘Hey, is there a discrepancy with this data?’

You do want that human verification. You're never going to be able to get rid of that responsibility, but you can get hallucinations down lower.”

You talked about SEOs making a mistake by not selecting the right model.

Are there any other key mistakes that you see SEOs making in this process that you'd like to highlight, and they can rectify?

“The general thing I would say, from what I’ve seen with some clients who use our product, is that they will just type something in. It's more the younger guys that are doing this; the less experienced folks. They drop something in, they get an answer, they take it, and then off they go because they’ve got to get a job done. They’re in a rush. They're not being disciplined about thinking, ‘Wait, I'm going to go over this. I'm going to ask questions.’

You can sit down with an LLM and literally ask questions when it gives you back an answer. You can go to the agent and be like, ‘How do you arrive at that conclusion? Wait, no, you need to go check this. The next time you come in with this information, use the ‘ask human’ tool and check with me before you proceed.’

It's building these workflows of where to check and how to check, to get a result. However, you still need to put the work in. With all that talk about ‘Agents are going to replace people,’ that's not happening.

I sell this stuff. It should be part of my narrative to go out there and say, ‘Hey, if you buy my product, you can replace everybody.’ I'd make a tonne of money, and a lot of the frontier models are saying that, but it's not true. It is not. You're going to get some productivity boost, and it's genuinely amazing when you do, but it is not replacing people.”

What's the end result that clients are looking for here, in terms of direction? How often are they looking for reports and strategic advice? What does it look like visually?

“Basically, it kicks out a report. This is the first step. You can run an analysis a lot faster.

We have people who don't know how to use the software, and an audit takes them about four hours. They learned the software and got the result back, which was comparable – just as good as the work from four hours – in two and a half. Now, they can basically just run it again, and they get that report instantly. So, they can schedule that report consistently and immediately get the report back.

Now, still, like I said, you should check these things when they're done, but once you build that up, you just hit a button and then away you go. You simply record a set of steps to do, and you can just execute those. Those are called agent skills.”

You talked about the fact that humans aren't getting replaced here.

“No, that's just not happening.”

Are there any tasks that humans should absolutely be doing and data that humans should be analysing as opposed to letting the LLMs do it?

“The landscape's always changing, and the important thing to understand about an LLM is that its information can be up to months old. As fast as the space changes, the LLM is not going to have the latest strategies and capabilities for what it needs to do.

That's where it takes professional maintenance, instruction, and guidance to make sure that the agent's doing what you want. If Google doesn't update (as Google does), that information is not in an LLM. So, if it's operating on the old way, you're going to get old results, and they're not going to be accurate. This is the thing.

The LLMs don't continually update, but a human certainly can.”

That's a great point. You also say that AI isn't going to replace human judgement.

What does human judgement look like in practice?

“A human being, when you talk about intelligence, we learn on the fly, pretty much. To get into our long-term memory, sure, you’ve got to go to sleep for the high-level concepts to really sink in, but we can look at something and learn instantly.

An LLM cannot do that at all. There's no intelligence in these systems. They're really good new information infrastructure. They're the equivalent of an encyclopaedia, pretty much.”

Andrew, what's the key takeaway from the tip you shared today?

“I would say, agents; definitely check them out. Use them. They're going to speed you up. They can't replace you.

Definitely don't be lazy. Don't trust them. They're not going to take your job, but definitely learn to use them. It's the literacy of the future.”

Andrew Melnychuk-Oseen is Founder at SAAGA Solve. Find out more over at SAAGASolve.com.

Choose Your Own Learning Style

Webinar iconVideo

If you like to get up-close with your favourite SEO experts, these one-to-one interviews might just be for you.

Watch all of our episodes, FREE, on our dedicated SEO in 2026: Additional Insights playlist.

youtube Playlist Icon

Podcast iconPodcast

Maybe you are more of a listener than a watcher, or prefer to learn while you commute.

SEO in 2026: Additional Insights is available now via all the usual podcast platforms

Spotify Apple Podcasts Audible

Book iconSEO in 2026

Catch up on SEO tips from 117 SEO experts in the original SEO in 2026 series

Available as a video series, podcast, and a book.

SEO in 2026

Could we improve this page for you? Please tell us

Fresh Index More info about Fresh Index icon

Unique URLs crawled 194,188,717,434
Unique URLs found 618,052,412,077
Date range 07 Apr 2026 to 05 Aug 2026
Last updated 1 hour 2 minutes ago

Historic Index More info about Fresh Index icon

Unique URLs crawled 4,502,566,935,407
Unique URLs found 21,743,308,221,308
Date range 06 Jun 2006 to 26 Mar 2024
Last updated 03 May 2024

SOCIAL

  • LinkedIn
  • YouTube
  • Facebook
  • Bluesky
  • Twitter / X

COMPANY

  • Blog External Link
  • About
  • Terms and Conditions
  • Privacy Policy
  • GDPR
  • Contact Us

TOOLS

  • Plans & Pricing
  • Site Explorer
  • Compare Domains
  • Bulk Backlinks
  • Search Explorer
  • Developer API External Link

MAJESTIC FOR

  • Trust Flow
  • Flow Metric Scores
  • Link Context
  • Backlink Checker
  • Influencer Discovery
  • Enterprise External Link

PODCASTS & PUBLICATIONS

  • The Majestic SEO Podcast
  • SEO in 2026
  • SEO in 2025
  • SEO in 2024
  • SEO in 2023
  • SEO in 2022
  • All Podcasts
top ^