Field notes
Answer engines, crawlers and browser agents: a field guide to AI visitors
Not every AI visitor wants the same thing. A practical taxonomy of the agents on your site, what each one is after, and how to treat it, from robots.txt to forms.
Most websites still put every AI visitor in one bucket labelled "bots", then make one decision about it: allow or block. That's like treating a search engine, a scraper and a customer as the same person because none of them filled in your newsletter form.
AI visitors come in distinct types. They're sent for different reasons, behave differently on the page and deserve different treatment. Block the wrong one and you vanish from AI answers. Welcome the wrong one and you pay for traffic that gives nothing back.
Here is the taxonomy we use in Ethogram, with what each type wants and how to think about it.
Answer engines
Who: assistants fetching a page because a person just asked them something. ChatGPT, Claude, Perplexity, Gemini and others, usually through user-triggered fetchers such as ChatGPT-User, Claude-User and Perplexity-User.
What they want: a fast, readable answer. The price, the spec, the opening hours, the returns window, the API parameter. They typically read one or a handful of pages and leave.
How they behave: mostly simple HTTP fetches without JavaScript. Visits are short and bursty, and they follow the questions people are asking: a product in the news can get read many times in an hour.
How to treat them: these are the closest thing to a referral you'll get from AI. Let them in. Make sure the facts people ask about are in the HTML, not rendered after load. Keep important pages at stable URLs. And keep facts consistent across pages, because an assistant quoting your old price from a forgotten landing page is worse than not being quoted at all.
One nuance: several providers document that their user-triggered fetchers don't always follow robots.txt the way crawlers do, on the reasoning that a person asked for that specific page. Check each provider's documentation before relying on robots.txt to control them.
Search crawlers
Who: crawlers that build the indexes behind AI search, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot. Google's AI features rely on the regular Googlebot index.
What they want: coverage and freshness. They visit many pages, revisit the ones that change and use that index to decide what to cite when someone asks a question later.
How they behave: much like classic search crawlers. Systematic, polite when well built, guided by your sitemap and robots.txt.
How to treat them: if you want to appear in AI search results and be cited as a source, allow them. Everything that helps classic SEO helps here: a sitemap, canonical URLs, server-rendered content, clear titles and descriptions. Watch their crawl rate on large sites, but blocking them is usually a decision to disappear from a growing share of search.
Training crawlers
Who: crawlers that collect content to train models, such as GPTBot, ClaudeBot, CCBot (Common Crawl, whose dataset is widely used for training), Bytespider and Meta-ExternalAgent. Some companies also offer robots.txt tokens that control training use without a separate crawler, such as Google-Extended and Applebot-Extended.
What they want: text, at volume. They're not answering anyone right now, and they don't send visitors back.
How they behave: broad, deep crawls. They can be heavy on bandwidth and are the main reason some sites started blocking "AI bots" wholesale.
How to treat them: this is a business decision, not a technical one. Some publishers block training to protect their content; many companies welcome it because they want models to know their product, their docs and their name. The useful thing to know is that the major providers separate training from search: blocking GPTBot doesn't remove you from ChatGPT search results, and blocking Google-Extended doesn't affect Google Search. You can decide each one on its merits, in robots.txt.
Browser agents
Who: agents that drive a real browser to complete a task for a person, like ChatGPT agent, Perplexity's Comet or Claude in Chrome, and the wider family of computer-use agents.
What they want: to finish the job. Find the product, pick the size, fill in the address, pay, book, sign up.
How they behave: like a very fast, very literal person. They run your JavaScript, click buttons, type into fields and read the page through its structure and accessibility tree as much as through what it looks like. When something doesn't respond, they try again, often several times, and then give the task back to their person.
How to treat them: these visits carry the most intent of any AI traffic, because someone delegated a real task. Don't block them by default. Make your flows usable without a mouse: real buttons and links, labelled form fields, native or properly accessible controls, error messages in text. Some browser agents now sign their requests, which lets you tell a legitimate agent from a scraper pretending to be one.
Personal agents
A growing branch of browser agents deserves its own name. Personal agents are always on, have their own cloud computer and browser, keep their own logins and work toward goals for one person over days or weeks. They might check your pricing every Monday, watch a product until it's back in stock, or run an outbound routine on a schedule.
They behave like loyal, very busy customers. They come back, they sign in, and they remember what failed last time. Ethogram groups them with browser agents in traffic reports and gives each one its own profile, including when in the day it visits.
The taxonomy at a glance
| Type | Sent by | Wants | Runs JavaScript | Sensible default |
|---|---|---|---|---|
| Answer engines | A person's question, right now | One fact, fast | Rarely | Allow; put facts in HTML |
| Search crawlers | An AI search index | Coverage and freshness | Rarely | Allow; sitemap and canonicals |
| Training crawlers | A model training pipeline | Text at volume | Rarely | Your call, per crawler |
| Browser agents | A person's task | To complete it | Yes | Allow; make flows accessible |
| Personal agents | One person, ongoing | Goals over time | Yes | Allow; expect return visits |
Name it before you decide
The common thread: you can't set a good policy for a visitor you can't identify. User agent strings are a start, but they're easy to fake, and browser agents often look like an ordinary browser. Reliable classification combines several signals: the declared user agent, published IP ranges, request signatures where they exist, and how the visitor actually behaves on the page.
That's what Ethogram does for every visit. Each agent is named, grouped by type and recorded step by step, so your robots.txt, your rate limits and your roadmap can be based on who actually visits, not on one bucket called "bots".
- agent types
- robots.txt
- crawlers