Episode 15 · The Search Signal
About this episode
Before an AI assistant can use a company's website in an answer, it has to reach the site and read it. A site that's technically sound for Google can still fail that step in ways nobody on the team chose. A content delivery network (CDN) default can block some AI crawlers, and a snippet tag added years ago can keep pages out of AI Overviews.
In this episode, Michael Transon makes the case that most of the technical search engine optimization (SEO) work that gets a site found in Google also gets it found in AI search, and that the differences mostly come down to which bots you let in and whether each important page makes sense on its own. Watch to learn which search index each AI platform uses, what blocking a training bot costs you, and why LLMs.txt is a default skip. Then take the episode's ordered list of checks to your team.
Michael's POV in 60 seconds
One thing
Your team has spent years improving your technical SEO so your pages show up in Google.
But does that work still matter when AI platforms answer your buyers' questions too?
So what
Google's AI Overviews and AI Mode run on Google's regular index, so the work your team already does supports your visibility there. But indexability now means being in a search index and in a model's training data, and most of the big AI companies run a separate bot for each. Block the training bot, and the next model learns about you only from what other sites say.
Now what
Keep doing the technical work already on your roadmap, and build each important page as if it's the only one an assistant will read. Make crawler access a decision someone owns, with your legal team involved before anyone restricts a training bot. The differences from traditional search are a lot smaller than the industry makes them sound.
Questions this episode answers
How do you check that the bots you want are actually getting into your site?
Confirm that Googlebot, Bingbot, OAI-SearchBot, and the ClaudeBots are allowed in robots.txt and that nothing in your firewall or content delivery network (CDN) blocks them, since robots.txt can say a bot is allowed while the CDN turns it away. Then have whoever runs your hosting or technical SEO check server logs or CDN bot reports to confirm the search bots get your pages back without errors.
If you have a legitimate reason to block AI training bots, how do you do it?
Block narrowly. Robots.txt lets you shut a specific bot out of a specific folder, so if the proprietary part of your site is a data product or a research section, block that and leave your homepage, services, product, about, and case study pages open. Keep the search bots allowed everywhere you want to be found, and if your legal team needs the block enforced, it has to happen at the server level or on your CDN.
How do you put the context an assistant needs on each commercial page?
Near the top of the page, write your company name and a plain description of what you do, in text, with a clear link back to the main page for that service. Keep your schema, and also write the facts that matter, such as prices, service areas, hours, and certifications, in the visible text, since an assistant reading the page live may never see a fact that only appears in the code. None of the important facts should be hidden in images or loaded in JavaScript after the page does.
Sound bites
Resources mentioned
Take this further
Don't miss an episode
New episodes every Thursday. Get them in your podcast app of choice, or watch on YouTube.
Full transcript
Michael 00:00 – 00:17
So blocking training data doesn't keep you completely out of the answers where the assistant runs a live search. But there's a big but. Where it does cost you is the answers that don't involve a search at all.
Michael 00:17 – 00:36
Remember, some of what these assistants say comes straight out of what the model learned in training. So when you block the training bot, you're telling the company to leave your future content out of what the next model learns. Anthropics in documentation words it almost exactly that way.
Michael 00:36 – 01:03
The next model will still learn about your company because there are other websites that mention you, and those are in its training data. But it learns about you only from what everyone else writes. And your own description of yourself is not a part of it.
Michael 01:03 – 01:30
Hey, welcome back to The Search Signal. I'm Michael Transon. I'm the founder and CEO of a search marketing agency called Victorious. And here at The Search Signal, our goal is to really help marketers understand and take action on what is happening in the world of search marketing. We have been ourselves running campaigns for nearly 13 years. We work with some of the biggest and also some of the fastest growing brands that are out there. And our goal is to bring what we learn from those experiences, plus a lot of our own
Michael 01:30 – 01:53
really interesting first party research to The Search Signal to help you with your own work. So in the last few episodes we've been mostly talking about what you publish and also what other websites say about you and how those two variables really impact a website's performance in AI search. And then today I wanna switch it a little bit. I want to talk about the technical side of AI search, which is
Michael 01:53 – 02:21
honestly, the part that decides whether these systems can even reach your website and also read it in the first place. And the question that I want to answer today is one that I am getting from a lot of marketing leaders right now. And the question is of all of the technical SEO work my team has historically done and does right now, how much of that actually counts for AI search? And what do I need to add? And is there anything that I can maybe even stop doing? So before we get into all of that
Michael 02:21 – 02:25
let's talk about what happened this week in the world of search marketing news.
Michael 02:25 – 02:48
Okay, starting with Google. Last week Google released its September 2026 spam update. If you have not seen, this is the fourth spam update that they have released this year. And it said that the rollout would probably take a couple of weeks. So the biggest ranking changes that we have seen at Victorious so far happened from Friday the 25th through Sunday.
Michael 02:48 – 03:14
And if you use rank tracking tools, it should all be recorded in there. We're only about halfway through the rollout as of today's recording. So it's still a little bit too early to say which sites are losing rankings and why they're losing rankings. And I would also be very careful with anybody saying that they've already figured out what happened and why. What I would do if I was you, just kind of hold off any changes on your website that are big until Google says that the rollout is done. And then you can just go into Search Console
Michael 03:14 – 03:41
and you can look for page that lost any traffic starting, probably, I would recommend around that start time of the 25th. And then the next thing for this week is I actually want to do a quick follow-up from episode 10. And in that episode, I talked through some of the new AI crawler categories that Cloudflare had recently set up. And I had said that the new blocking defaults would apply to newly onboarded domains. That's also true, but it's not totally complete. So
Michael 03:41 – 04:10
per Cloudflare's recent announcement, these defaults also apply to new sites that existing customers add and to every existing customer on the free plan. Now, those defaults did go live on, I believe, September 15th. And those updates, they block the training and the agent crawlers on page that are going to show ads. And then search crawlers are staying allowed. So most B2B sites that don't run ads, which is probably quite a bit of us, we won't see the change, but,
Michael 04:10 – 04:39
if your website monetizes with ads and you're on Cloudflare's free plan right now and you've never touched these settings, there's a probably a good chance some of the AI crawlers could now be blocked on your page without anyone on your team deciding that. So I would definitely look at that. And the last one this week is a little bit of a smaller update, but I think it connects really well to today's topic. So earlier this month, Google had refreshed its developer documentation for something called the Web Search Service API.
Michael 04:39 – 05:08
Which returns full Google search results to a list of their approved partners. you can't really just sign up for it, although it would be great if you could. So every one of these requests has to carry an ID tied to a partner agreement. And then sometime around the 20th, it looks like that documentation actually disappeared. And there's no notes or redirect, and Google hasn't said who the partners actually are. Now this is small, but I wanted to bring it up because every AI assistant that searches the web
Michael 05:08 – 05:33
needs an index to search and who is getting access to whose index is starting to turn into some business dealings and backroom negotiations. So I bring that up because all of this week's news is about who can read your website and on what terms they can read it. So let's dive into today's episode.
Michael 05:33 – 06:03
When a marketing leader typically comes and asks me whether their technical SEO, current technical SEO, covers them for AI search, what they are usually asking me is whether they have to go out and spend some more money on something new or need to spend more time on it. And I would just tell you that my answer is mostly gonna be no. There are definitely a few things that are going to need to change. And there are definitely a few things that I would tell you that you just don't need to do at all. And the way that we have been thinking about this at Victorious is.
Michael 06:03 – 06:29
I think pretty simple. We drop it into three buckets. Either something that we start doing, something that we need to stop doing, or something that we need to keep doing. Essentially, what does your team already do for traditional search from a technical perspective that keeps working for AI search? That's bucket one. Number two is what do you need to actually start doing for the first time that you haven't done before? And then the third is what can you just stop?
Michael 06:29 – 06:59
And just so you know up front, the stop list, in my opinion, is probably going to be pretty short. And that's honestly on purpose. And the reason is because anything from traditional technical SEO that doesn't carry over into AI search is still something that you should keep on doing for as long as your buyers use Google search to find you. And they right now, I'll tell you, they do. It's still a big channel. So nothing I say today is a reason to just stop doing any work that's keeping you visible in traditional Google
Michael 06:59 – 07:21
search results. I would say the stop list we'll talk about today is mostly about like stuff I wouldn't recommend you doing in the first place, maybe things that you see on your LinkedIn feed of sexy hacks and things that you can do to manipulate AI search results. I'm gonna say no to most of those things, and we'll talk about that when we get there. But I would also say one important point is that these platforms change how they work
Michael 07:21 – 07:33
every couple of months. It's crazy. But the basic mechanics of how a machine gets to your website and reads it are very well understood. And so that's actually where I would like to start.
Michael 07:33 – 07:53
The first technical question in SEO has always been about crawlability, meaning whether a bot can actually get to your website or your page at all. And then the second has been about indexability, meaning whether the page can get stored somewhere so that it can be found later.
Michael 07:53 – 08:07
Now, in current search environments, crawlability still means the same thing that it has always meant, but indexability now means two different things. And it's important we understand this.
Michael 08:07 – 08:36
The first thing it means is being in a search index. And that is the version that we have historically done and we all know. This is like Google's search index or Bing's search index. And then also now, the indexes of some of the AI companies that are starting to build indexes for themselves. When an AI assistant decides to go search the web in the middle of a conversation, it's looking things up in one of these. Now the second,
Michael 08:36 – 08:56
is being in the training data, which is everything that the model reads before it is released to the public. And we talked about this last episode when we got into how long it takes to see results from AI search activities. Some answers come from a live search, and then some come straight out of what a model has already learned.
Michael 08:56 – 09:18
And there's no searching being done at all. So for AI search, being indexed means we need to be in both places. You want your content to be part of what the model has actually already learned about you, and part of it, we want it to be able to see us when it looks up what it needs to look up when it's trying to search and find information for an answer. And
Michael 09:18 – 09:47
I will say the bots that you let onto your site are going to decide which of those two places you end up in. So most of the the big AI companies now run separate bots and crawlers for each different job. So OpenAI has GPTbot that collects content that might be used for its training. And then it also has OAI search bot, which is the one that decides whether you can show up in ChatGPT's search answers. And then OpenAI says these two settings
Michael 09:47 – 10:09
are totally independent of each other. And then Anthropic also, as another example, has the same kind of split. They've got ClaudeBot for training, and then they also have Claude SearchBot for search. So these companies have a separate bot that visits a page in real time when a person asks a question. That's the third type of search bot or bot that these use. And we're gonna talk about that third one in just a few minutes. But
Michael 10:09 – 10:32
the point being is some answers come from a live search, the next question needs to be which index each of these platforms is searching, because that tells you where your pages have to be findable. I will tell you this: on one of these platforms. it took us researching this episode to figure something out. So I'll get to that in a second, but let's start with Google.
Michael 10:32 – 10:54
Let's talk about AI overviews and AI mode. Those are both part of Google Search and they run on Google's regular index. And Google's own documentation says that to show up as a link in either one a page has to be indexed and eligible to show up in Google search with a snippet, and that there are no additional technical requirements. So
Michael 10:54 – 11:19
the point I want to make here for Google and for Google's AI search surfaces your technical SEO is your AI technical SEO. It's the same crawler, Googlebot, and it's also the same index. But there are two details in there that I think a lot of teams might miss. And the first is that the robots.txt setting called Google Extended, which a lot of companies use to actually opt out of Google's AI training.
Michael 11:19 – 11:43
Doesn't affect whether you show up in Google Search. And that also includes AI overviews. What it does is it controls training and grounding for the Gemini apps and also some of developer products that Google has. So if somebody on your team blocked Google Extended Thinking, thinking it would keep you out of AI overviews, I don't know why you would do that, but it didn't do that. And then the second detail is actually the opposite side of things. So
Michael 11:43 – 12:07
Google says that the old snippet controls, things like using no snippet and max snippet, also limit what it shows in AI features. So this is a maybe a unique situation for some of us, but if there were any utilization of no snippet tags that your team, put on a page or a set of page years ago for no reason your team might remember anymore those pages
Michael 12:07 – 12:35
may be excluded from AI overviews today. I had said earlier that we were learning about how these different systems are using different indexes. And one of the things we learned was new information preparing for this episode. And ChatGPT was actually one that we learned about for the first time. So for a long time, and this is still true across a lot of the industry, the working assumption was that ChatGPT runs its web searches on Bing because of the Microsoft relationship that they have.
Michael 12:35 – 12:59
Now, as of very recently, the latest research doesn't actually support that anymore. OpenAI's own documentation talks about ChatGPT pulling from what it calls "indexed and cached web content." And earlier this month, a company called Peec AI which makes AI visibility software, they had published some research that maps what looks like a set of internal ChatGPT search indexes
Michael 12:59 – 13:15
with a general web index plus separate ones for things like news and shopping and PDFs and YouTube. They also found some experiments where OpenAI was appearing to be testing its own index against outside search results to decide
Michael 13:15 – 13:43
which one it should use for a given question. So the name that I've been seeing come up for all this is something called Labrador. And also just to be clear, OpenAI has never officially confirmed it, or it's name, or how any of this works, but it also lines up with what OpenAI's head of ChatGPT had actually testified during their Google antitrust trial, which is that OpenAI had started building its own index way back in 2023. So
Michael 13:43 – 14:12
the most accurate way that I can probably put this is that ChatGPT is running on a mix of its own index and also outside providers. And the research is suggesting that OpenAI is probably moving more of it onto its own index. And OAI SearchBot is the crawler that OpenAI uses for ChatGPT search. And OpenAI says if you opt out of that one, you will not be showing up in ChatGPT's search answers.
Michael 14:12 – 14:41
So, as we continue to talk about different indexes and what different AI systems use as their indexes to determine what they find and show in answers, let's also talk about Claude, which is actually a lot simpler. Anthropic lists Brave Search as one of the vendors that it uses for Claude's web search. And you might have seen this on LinkedIn. There were some trending posts online where people have been comparing Claude citations
Michael 14:41 – 15:07
with Brave's search results and see that they mostly match. Brave is its own independent search engine and it's got its own independent index. And then lastly, Microsoft's Copilot, which is big in enterprise AI usage, runs on Bing. So, what would I do with all of that information? Four things that I would think about. So let's break it down by platform.
Michael 15:07 – 15:34
If we're trying to show up in Google, if we're trying to show up in ChatGPT, showing up in Claude, right? So if we're trying to show up in all of these systems, the question is what should we be doing from a technical perspective unique to each one of these systems? Because again we've talked about this in prior episodes. We don't want to treat AI systems and AI search as a monolith. We need to recognize that each one of these systems is unique.
Michael 15:34 – 15:56
Has a unique user set and they have different sizes of user bases. So we need to prioritize which ones matter. So let's talk about these three: Google, ChatGPT, and Claude. Very straightforward. For Google, you need to keep doing the technical SEO that you are already doing. And if you can add the Google Extended and snippet tag checks that I mentioned earlier. Now for ChatGPT.
Michael 15:56 – 16:25
First, we need to confirm that the OAI search bot is actually allowed in your robots.txt file and that there's nothing at all about your firewall or your CDN that is blocking it. I can talk about that, and I will talk about how to check that in just a minute. For Claude, if you want, you can go search a handful of your most important buyer questions on a Brave search and you can see whether your pages come up. If they don't, you know, we might say they might have a hard time showing up in Claude when it searches that index as well.
Michael 16:25 – 16:52
And then I would also say as it relates to Copilot, just straightforward, you know, these are just high-level recommendations to start today's episode. I would keep Bing's webmaster tools. I would keep that account set up with your site map clearly submitted, because Copilot again still runs on Bing. Now, a year ago, I would have told you that ChatGPT was Bing. And, these arrangements are business deals like that Google API one from the news that we talked about this week. So
Michael 16:52 – 16:59
I would just recommend we're constantly checking in which index each platform uses when we're doing our quarterly search reviews for our projects.
Michael 16:59 – 17:28
Okay, that was a lot. That was the index piece. Now let's dive into talking about what carries over. Let's talk about the continue list of technical search. This is the technical work that you're probably already doing for traditional SEO that's going to help you in AI search. So, zooming out, when we run a technical audit at Victorious, we look at six different areas. We look at crawling and indexation,
Michael 17:28 – 17:53
site architecture, internal linking, rendering and performance, structured data, and duplicate content. That's pretty much all of it. And it all carries over for the reasons that we just talked about. Google's AI features uses Google's index. So anything that keeps a page out of Google's index will keep it out of AI overview and AI mode. But I think a few of these deserve a little bit more time.
Michael 17:53 – 18:17
So let's talk about them. I want to start with site architecture and internal linking and I think this is one that I probably want to spend the most time on. And the reason why I want to do that is because they tell a machine what matters on your website and how everything connects together. For example, I was just running a chat experiment in ChatGPT a couple of weeks ago, and I asked it to recommend
Michael 18:17 – 18:33
the best companies in a random industry that I was just looking at. And then I picked one of the one of the companies that it ranked low. And I asked it, why was that company not ranked higher? And a lot of its answer came down to how clearly the site was organized.
Michael 18:33 – 18:52
It described an ideal hierarchy where you have your products and your services and then you have your capabilities under each of them, and then more specific subcapabilities if you've got them. And then moving to workflow of how you deliver those capabilities and then the evidence that you've successfully done it.
Michael 18:52 – 19:07
And then having your blog post and your blog content really supporting all of that. Now, I don't fully trust a model's explanation of its own reasoning very much because it's gonna give you a plausible answer, whether that answer is true or whether that answer is not. But
Michael 19:07 – 19:26
the structure matches how we'd probably tell you to build a site anyway. And we had actually walked through this together in a prior episode with anchor pages and supporting pages. So that's I would say that's a great recommendation for continuing to invest in that. But now one question our team couldn't answer.
Michael 19:26 – 19:44
While we were also preparing for this episode is whether an AI assistant that lands on one of your page follows the internal links on that page or goes to a new search instead. Now I haven't seen research that answers it, and this is definitely related to site architecture. So until there's evidence otherwise,
Michael 19:44 – 20:10
I would recommend from an architecture perspective, building every important page on your site as if it's the only page an assistant is going to read. And what I mean by that is we don't know if an assistant comes to one page on your website and then bounces around to the next one and the next one and the next one and gathers wide context, or for compute efficiency purposes, just visits one page to get as much information as it can and then leaves.
Michael 20:10 – 20:37
Now if you do want to know what is happening on your website related to this, the answer could be in your server logs, which record every single request that happens, including which bots ask for which pages and when. What you look for is the same bot requesting one of your pages and then right after it requesting the pages that page links to. And if server logs sound scary, that's okay. If no one on your team has ever worked with server logs before, that's okay.
Michael 20:37 – 21:04
This is just a question for whoever runs your hosting or whoever is responsible for your technical SEO. Let's also talk about rendering, meaning whether your content actually shows up in a page's source code or only after JavaScript runs. We've covered this actually twice now in two different episodes. And the short version is that most AI crawlers don't run JavaScript. So if your main content only appears after the page loads.
Michael 21:04 – 21:32
They probably cannot see it. And I would definitely go back and listen to those episodes. We talk a lot more in detail about this. And then the other thing that I would talk about is page speed. This is one that I get asked about a lot. I would say we need to keep page speed on the continue list because it counts for Google and Google's AI surfaces, like we talked about, run on Google's index. Now, whether speed on its own actually changes
Michael 21:32 – 21:59
what ChatGPT or Claude chooses to cite, I haven't seen any research that shows it one way or shows it the other way. So I would keep it at its current priority on your roadmap and not raise it just because of AI, because we haven't seen any data or research that proves that a faster website results in better AI search response outside of what we already know about Google.
Michael 21:59 – 22:27
Now, structured data is the one where I would probably change how you think about it without changing whether you do it. What I mean by that is we talked about organization schema and same as schema in our brand entity episode. I'm not gonna go back to it, but you can go listen to that episode. But what's new since then is there's been some testing on whether AI assistants
Michael 22:27 – 22:55
even read schema at all when they fetch a page live. And some of the tests I've seen have said that ChatGPT and Claude actually don't. One test had put a product price only in the schema code and nowhere in the visible text. And none of the five AI systems that they tested could find it. Now there was another test that found Gemini could read the schema when it was asked.
Michael 22:55 – 23:20
While ChatGPT and Claude couldn't find any. Now, these are small tests that are typically run by software vendors. So I wouldn't treat them as like established proof, but they do match how we tend to understand structured data, particularly related to whether we should continue doing, start doing, and stop doing technical work for AI search performance.
Michael 23:20 – 23:47
Which is that your schema gets read by search engines when they index your page. And that's how it helps you in AI search through Google and Bing's indexes. So what I would do is keep your schema and make sure any fact that matters is also written out in the visible text of your page. Your prices, your service areas, hours and certifications, all that stuff. If a fact only appears in the code, an assistant reading your page live is potentially never going to see it.
Michael 23:47 – 24:11
So that's the continue list. So let's now pivot and talk about what you should probably start doing that you're not doing right now. So the first one that I would recommend comes from how these assistants read a website, which is actually a little bit different than how Google actually reads your website. So when Google crawl your site, it's trying to map the whole site, essentially every page and how all those page connect together.
Michael 24:11 – 24:35
Then when an AI assistant fetches your site and does it live for somebody, it's usually doing something else. It's trying to answer one specific question and it's trying to do it as quickly and as cheaply as it can. And every page that it reads costs it compute and processing. So it will often go to just one page, take what it needs, and then leave. There are a couple of small log studies that I've seen that show that these live
Michael 24:35 – 25:02
fetches typically go to the home page first, and then they go to pricing and then they go to comparison pages. Now, what this means for you is that any page that an assistant lands on might be the only page that it reads. So each of your important pages needs to carry enough context about your company to make sense on its own without, this is important, taking focus away from what that specific page is about.
Michael 25:02 – 25:11
We covered the content side of this in the last episode. We talked about things like writing short descriptions of your company, for example, at the end of every blog post.
Michael 25:11 – 25:40
Now, the technical side is making sure that the context is there in a way that a machine can read. And that means your company name and also a plain description of what you do. I would recommend near the top of your commercial page, it has to be written in text. And I would also say include a clear link back to the main page for that specific service. And it also means none of the important facts in your site can be hidden in images or loaded in JavaScript after the page
Michael 25:40 – 25:41
loads.
Michael 25:41 – 25:47
That was the first thing. Now the second thing to start doing for technical AI search performance,
Michael 25:47 – 26:12
is we need to make our crawler decisions intentional and on purpose. And I want to spend some time here because there's a conversation that I keep on seeing come up with companies that have proprietary content and maybe like a little bit of a nervous legal team. And it typically goes like this: The company has decided to block AI training bots because they don't want their data to be used to train somebody else's model. I get it. And their view is,
Michael 26:12 – 26:40
it doesn't cost them anything in AI search, because the training is just training, and the search bots are the ones that decide whether you show up, it's all good to go. Now, the first time this came up for me, I also saw the point of view. And they are mostly right. OpenAI's documentation says you can block GPTbot, allow OAI searchbot, and still show up in ChatGPT search answers, technically.
Michael 26:40 – 27:03
And for Google, we just covered that Google Extended doesn't affect AI overviews. So blocking training data doesn't keep you completely out of the answers where the assistant runs a live search. But there's a big but. Where it does cost you is the answers that don't involve a search at all.
Michael 27:03 – 27:23
Remember, some of what these assistants say comes straight out of what the model learned in training. So when you block the training bot, you're telling the company to leave your future content out of what the next model learns. Anthropics in documentation words it almost exactly that way.
Michael 27:23 – 27:51
The next model will still learn about your company because there are other websites that mention you, and those are in its training data. But it learns about you only from what everyone else writes. And your own description of yourself is not a part of it. Which, after the whole episode we had last week on building a brand entity, is kind of the opposite of what you want. I would also say one very important thing, and I saw this this week, and
Michael 27:51 – 28:18
I want to talk about it in later episodes as we get more information and data. But it seems like some of these AI systems are spending less and less compute on active searches and are being more and more reliant on their training data. Now we'll see if that ends up continuing to be true. And I'd also add one more point to this. And you know, this is my own opinion. I can't prove it. But when an assistant does go and search,
Michael 28:18 – 28:39
it's typically gonna start from what it already knows, just like a human. And if it already has a very clear picture of your company, it's probably more likely to search the right sources. So, what the model learned about you might still shape what it finds and where it starts a live search, even in a live search. So, what I would do.
Michael 28:39 – 29:02
If you do have any legitimate reason to block training I would block it very narrowly. Okay. Robots.txt lets you block a specific bot from a specific folder on your website. So if the proprietary part of your site is a data product or a research section, you can block those, but leave the pages that describe your company wide open, like your homepage, your services page, your product pages, your about page, your case studies, anything that's not proprietary.
Michael 29:02 – 29:29
Those are the page that you want the next model to learn from. And keep the search bots allowed everywhere that you want to be found. I would strongly recommend you do not block those. And while we are on this topic, there's also a related assumption that I want to clear up, which is that a line in robots.txt stops a bot completely. Robots.txt is a request. the bots that respect it stay out, and the ones that don't respect it are gonna come in anyway.
Michael 29:29 – 29:57
Anthropic says that its bots honor robots.txt. OpenAI says that for a chatGPT user, or the bot called ChatGPT user, which is the bot that visits a page when a person asks ChatGPT something in real time, robots.txt rules might not apply because a person started that request. So your proprietary data might still be used in a ChatGPT answer, even if you said you didn't want them to visit that page. Now
Michael 29:57 – 30:19
what your robots.txt file says and what your server does can also be different in both directions. There's a company called HasData that sells, data in scraping tools. They tested this on a set of sites, and of the sites that banned GPTbot in their robots.txt, about four in 10 still served GPTbot the page when it asked for it.
Michael 30:19 – 30:29
And back in episode 10, we talked about the opposite problem where your robots.txt says a bot is allowed, but then your firewall or your CDN blocks it anyway. So
Michael 30:29 – 30:57
really, what I would do depends on why you're blocking. So if it's a legal requirement your legal team should know that robots.txt is just a request. And if they need the block enforced that has to be done at the server level or on your CDN. But if what you want is visibility, you need to have someone look at your server logs or your CDN bot reports and confirm that the search bots are getting your pages back successfully and not getting some sort of error message.
Michael 30:57 – 31:22
Okay, we have talked about the continue, the start doing list. Now let's finish by talking about the stop list. And I like I said, this is mostly of things that you should not start at all that you aren't doing right now. The big one that I would start with is no surprise, the llm'stxt file. So if you have not seen it,
Michael 31:22 – 31:49
this is all over the internet on a way to manipulate these results, but it is a text file that you put at the root of your website that's supposed to give AI systems a clean summary of your site and also a list of your most important pages. A lot of people are recommending it right now. And ironically, that company I said earlier has data, has a same study that found that a little under one in ten of the top websites have one. Now
Michael 31:49 – 32:14
we'll say this. Google search team has been very clear that Google Search doesn't use it. John Mueller, he's a guy who speaks very publicly for Google Search. He compared it to old keywords meta tags, which search engines had stopped paying attention to a very long time ago because site owners could put anything that they wanted inside of it. Now there is one exception.
Michael 32:14 – 32:40
Google's Chrome team added a check for LLMs.txt in the Lighthouse, which is their free site auditing tool if you've never used it, in a new section built for AI agents. So different parts of Google aren't totally saying the same thing. Mueller's take on that was was that it can make sense for developer documentation where AI coding tools read the docs, but for most other websites, it really doesn't make much sense. So my default is
Michael 32:40 – 33:05
really just don't do it. Don't waste your time unless you really have a very specific reason like developer documentation. And there are a couple of reasons for that default that go beyond whether this works. The first came from our team talking about this practically. So an LLMs.txt file is a second description of your company, and it's going to live in a file that most of your team is never going to look at.
Michael 33:05 – 33:36
So what happens? Like your marketing team launches a new service or they change how they're describing what you do on the homepage. And there's a good chance that nobody remembers to update the LLMs.txt file. Why is that bad? Well, now you've suddenly got two versions of your company out there on your site, which is the exact problem we spent in the last episode talking about trying to fix. The second is just my opinion, coming from watching search evolve for a very long time.
Michael 33:36 – 34:04
And how it relates to other things that people used to do that they thought could easily manipulate and change search results to their advantage. Every black hat SEO tactic that exists today was a white hat tactic at one point, until it wasn't. The point being is things that look harmless today as a way to influence these systems have the potential to hurt you later.
Michael 34:04 – 34:26
Once the platforms decide what they count as manipulation. And spam updates, just like the one we talked about in the news update this week that are rolling out right now, are usually when those things happen. So when you can almost in today's world do almost anything infinitely cheaply and infinitely quickly.
Michael 34:26 – 34:37
Personally, my personal recommendation, I would default to not doing something unless you can confidently explain why it's going to help, even if the reason is small.
Michael 34:37 – 35:03
Before we finish, as it goes with all of these episodes, we're talking about what we know today and what's happening in search today, but also this space is changing very, very quickly. And I want to talk about a few things that I'm watching that could potentially change parts of this. I want to keep this short because I really don't know what happens in all of them, but there's been a lot of talk on LinkedIn the last couple of weeks about,
Michael 35:03 – 35:28
I mentioned this earlier, newer models using web search less and relying more on what they already know to save on cost. This might be true, but I think efficiency also gets talked about often in the wrong way. We tend to assume that these systems want to use fewer tokens on every single step. But the bigger savings might also come from just taking fewer steps.
Michael 35:28 – 35:44
It's like handing a hard coding task to a smaller AI model. I don't know if you've used, like for example, if you use Claude, sending a coding task to Opus versus sending a coding task to Haiku, right? While
Michael 35:44 – 36:11
Opus is smarter and Haiku is less smart from an actual token utilization perspective. It can actually end up costing more in total to use Haiku because it makes more mistakes. It has to go back and redo things while a smarter model like Opus gets to the answer in fewer steps. So I wouldn't assume in that same line of thinking that the outcome is going to be less search.
Michael 36:11 – 36:37
Another really interesting thing to consider, and we are talking about internally as we look at these AI systems evolve, is that there's a real chance that live web search could be something that in the future users have to pay extra for because it's expensive to run and people might want current information. So I don't know if that will ever happen, but it would definitely change how these assistants search on our behalf.
Michael 36:37 – 36:55
And then another one to talk about is how agents are currently looking at our website and how they might look at it in the future. So right now we tell, and I've told you today, we tell people not to rely on JavaScript because most AI crawlers read your page's code and don't see what a browser draws on the screen.
Michael 36:55 – 37:25
But these systems keep getting better and better at looking at a page visually in many ways the way that a person would. If that becomes the normal way that agents are reading websites, some of today's technical advice is going to change. I don't know what that costs them to run visual analysis compared to reading the code. I haven't found a good number on it. So for for now, I would focus on the code. But there could be a real future that we live in where AI tools
Michael 37:25 – 37:49
and AI assistants and AI bots are looking at websites the same way that humans do. And then the last one I would say is: if you sell products online, is agent shopping. So my view for a while has been that when an agent buys something from you today, it mostly opens up, you know, in a browser and does what you would do. It takes over your tab, clicks through the same pages, and does the same type of checkout.
Michael 37:49 – 38:07
If that's the same workflow they always follow, your site's already built for it. And the part that actually matters here is just as it has always been, getting recommended earlier in the buying process because that's when the buyer has decided what to buy and then sends the agent to your site to go do the transaction.
Michael 38:07 – 38:34
Now, that's still very much true, but there's also a second approach that's starting to develop in agentic transactions. Google launched something called the Universal Commerce Protocol back in January, and they did this with Shopify, and they did it with a group of big retailers where the merchants were handing its product data and its checkout directly to the assistant. So the assistant didn't have to click through your site at all. OpenAI launched an in-chat checkout in ChatGPT last fall, and from what I have read,
Michael 38:34 – 39:01
it discontinued that in March and shifted to product feeds that send people to the merchant's own website. So it's still early and OpenAI already discontinued one of the biggest attempts to make agentic e-commerce a thing. But what I would do is if you're an e-commerce, is to make your product info very complete everywhere that it appears. This is what we always say. Do it on your product pages and do it in your product feeds, including your reviews, what it costs for shipping,
Michael 39:01 – 39:12
your delivery schedule and timing, if you have return policies, what they are, that's always been a good practice, even before agents show up. And it's the part that every one of these approaches do read.
Michael 39:12 – 39:37
So if you want to take this back to your team and begin optimizing your technical work for AI search in conjunction with SEO, the order I would work in is this. First off, start by confirming the search bots can actually get to your website. So that's Googlebot, Bingbot, OAI SearchBot, and the ClaudeBots in your robots.txt file and at your CDN. And ideally, if you can, confirm these in your server logs.
Michael 39:37 – 39:59
Then check that your important commercial pages are being successfully indexed in Google and in Bing, and that there's been no snippet tags that are sitting on them that have been there for maybe years, if your site's that old. Then we need to make sure that the facts that matter are written in visible text on those pages, and that each page
Michael 39:59 – 40:19
says who you are, I would recommend near the top. And that's so that if a crawler goes to just one page and does not go to other pages, it has the full context of who you are. And then I would recommend that you meet with, if you don't do it yourself, whoever owns your crawler settings and if you've got regulatory content or you've got proprietary content that you don't want to be
Michael 40:19 – 40:49
indexed in the training. Make sure that these legal teams are involved. They understand the training bot decision. And if you end up having to restrict training bots, do it on purpose and as narrowly as you possibly can. But taking a step back from this whole episode, the point of this whole episode, most of the technical work that gets you found in AI search is the same work that gets you found in Google in traditional search. And a lot of it should be, if you're not already doing it, it should be already on your roadmap.
Michael 40:49 – 41:02
And the differences are a lot smaller than the industry makes them sound. They mostly involve knowing which bots you're letting in and why and making sure every page is making sense on its own.
Michael 41:02 – 41:11
So thanks for listening to this week's episode of The Search Signal. If this was useful to you, please go ahead and subscribe so the next episode finds you. That is it for me today. I will see you next week.