Blog
Notes from the non-human web
Crawl budget: what it is and how to protect it
Crawl budget is the number of URLs bots crawl on your site in a set window. Learn what wastes it, how logs expose the leaks, and why AI crawlers now spend it too.
How to detect bot traffic in your logs
A server-log-first method to see and classify bot traffic yourself: the signals that identify a bot, how to separate good crawlers from scrapers, and why verification matters.
How to get cited in Google AI Overviews and ChatGPT
A practical guide to earning citations in Google AI Overviews and ChatGPT, then confirming they worked by measuring the AI-referral traffic they send back.
Generative engine optimization (GEO): the complete guide
A complete guide to generative engine optimization: what GEO is, how it differs from SEO and AEO, the citation factors that matter, and the crawler-access and measurement layers most guides skip.
llms.txt explained: what it is and how to create one
llms.txt is a proposed Markdown file that points AI models to your best content. Here is how to create one, plus the 2025-2026 evidence on whether crawlers use it.
Reverse DNS lookup: how bot verification actually works
Reverse DNS lookup maps an IP to a hostname, but verifying a crawler needs forward-confirmed reverse DNS. Here is how FCrDNS proves a bot is real.
How to optimize your site for AI search
AI search optimization starts with a technical question most guides skip: can AI crawlers actually reach and read your content? Here is the checklist.
How to verify Googlebot (and catch spoofed bots)
A user-agent string is not proof. Learn the layered way to verify Googlebot with reverse DNS, published IP ranges, and Web Bot Auth signatures.
Web crawler vs. web scraper: what's the difference?
A web crawler discovers pages by following links; a web scraper extracts specific data from them. Here is how to tell them apart in your server logs.
What is a web crawler? A modern taxonomy for 2026
A web crawler is an automated program that fetches web pages and follows links. Here is how crawlers work, the five types you now see, and how to tell them apart in your own logs.
What is GPTBot? A field guide to AI crawlers
GPTBot is OpenAI's training crawler. Learn what it does, its user-agent token, and how it differs from the search and user-fetch bots you should not block.
What is Web Bot Auth? A plain-English guide to signed bot traffic
Web Bot Auth is an emerging IETF method that lets bots and AI agents cryptographically sign their HTTP requests so sites can verify them. Here is how the handshake works and why it beats user-agent strings.
How to see ChatGPT and Perplexity referral traffic
AI referral traffic is the humans arriving from ChatGPT, Perplexity, and Gemini. Here are the exact referrer hostnames, GA4 setup, and the gotchas that hide it.
How to spot server throttling and inconsistent status codes blocking your crawlers
Intermittent 403, 429, 503 responses to Googlebot and AI crawlers quietly drop pages from search and AI answers. Here is how to detect, diagnose, and fix it.
Log file analysis for SEO and AI search: a practical guide
Log file analysis reveals exactly what search and AI crawlers do on your site. Here's how to read your logs to protect crawl budget and AI visibility.
301 redirects: how to use them effectively and handle edge cases
A practical guide to 301 redirects: when to use them over 302/307/308, how link equity passes, and how to handle chains, migrations, and edge cases.
See the non-human half of your traffic.
Lume shows you every crawler, scraper and AI agent on your site, and verifies which ones are real. Set up in minutes.
Start for free →