Crawler access

Are crawlers actually reaching your pages?

The pages you want indexed and cited are only as good as the response crawlers get when they ask for them. Lume runs server-side to show the status codes every requester hits (told apart from your real human visitors) so you can see whether Googlebot, GPTBot, and the rest reach your pages or quietly bounce off a 403. It works the same across your CDNs, and it’s measurement, not ranking magic.

Sound familiar?

What Lume shows you

This is the /status-codes view, productized: KPIs, a trend over time, requester-type segmentation, and the tables that turn “crawlers hit some errors” into “fix these pages.” Every metric comes with something to do about it. It’s the reachability companion to agent analytics, and it earns its keep for SEO teams and publishers alike.

Status codes, split by who’s asking

Filter every status code by requester type (all, human, bot, or agent) so you can see the responses your real visitors get and the responses crawlers get side by side. When a bot is hitting a wall your browser never would, the gap is obvious. Action: open the pages returning 4xx/5xx to crawlers and fix the access rule keeping them out.

The pages bleeding crawl budget

Find the top 404-magnet URLs crawlers keep requesting (dead links, retired pages, bad redirects), ranked by how much crawler traffic they’re wasting. Action: redirect or retire them so crawlers spend their budget on the pages you actually want indexed and cited.

Cloaking and misconfig, caught

An inconsistent-status table surfaces URLs that answer differently depending on who asks: 200 to humans, 403 to crawlers being the classic tell of a WAF rule or CDN setting quietly locking bots out of a page that looks fine to you. Action: reconcile the two so the page you can see is the page crawlers can reach.

The regression is right there in the trend

A per-requester status-code trend and a crawler-status card show the response codes crawlers are hitting over time, so the day a deploy flips GPTBot or Googlebot from 200s to 403s shows up as a step change in the chart the next time you look. No static firewall rule catches a silent regression like that; a continuous server-side view does. Export the underlying data to CSV on Starter and up.

Verify it’s really Googlebot first

A 403 on a fake Googlebot and a 403 on the real one mean completely different things: one is a scraper you may want to keep out, the other is a regression costing you indexing. That’s why crawler access rests on verified identity: Lume confirms who a requester really is with forward-confirmed reverse DNS, published IP ranges, and Web Bot Auth signatures across 7 major operators before its status codes count toward the picture. Verified isn’t the same as safe or well-behaved. It just means the name on the request is real, which is the floor every access decision needs.

FAQ

Does Lume alert me when a crawler starts getting blocked?

No. Lume doesn’t send push, email, or Slack notifications, and it won’t warn you the moment a status code changes. What it does is make the change visible: the day a crawler flips from 200s to 403s appears as a step in your status-code trend, so it’s waiting for you in the dashboard the next time you look. It’s a continuous server-side record, not an inbox.

Should I block AI crawlers?

That’s your call. Lume measures and identifies; it does not block. It shows you which crawlers reach which pages and what they get back, so you can decide what to allow, then act through robots.txt, your CDN, or your WAF. The blocking is always yours to do.

Will this get me cited in AI answers?

No, and we won’t pretend otherwise. This page is about reachability (whether crawlers can physically fetch the pages you want indexed and cited), not about whether you get ranked or quoted in an AI answer. Reaching the page is a precondition for being cited, not a guarantee of it. Lume shows you the reachability half honestly and leaves the ranking claims to others.

How do you know it’s really Googlebot and not a 403 on a fake?

A status code only means something once you trust the requester behind it. Lume verifies crawler identity with forward-confirmed reverse DNS, published IP ranges, and Web Bot Auth signatures across 7 major operators, so a “Googlebot” arriving from an IP that isn’t Google’s is flagged rather than counted. The access story you read is built on requesters that are who they claim to be.

Find out what crawlers really get on your pages.

Free plan, one ingest token, server-side across your CDNs. See the status codes bots hit, and catch the regression the day it starts.

Start for free →