Field notes

Answer engines, crawlers and browser agents: a field guide to AI visitors

Not every AI visitor wants the same thing. A practical taxonomy of the agents on your site, what each one is after, and how to treat it, from robots.txt to forms.

5 min readBy Ethogram

Most websites still put every AI visitor in one bucket labelled "bots", then make one decision about it: allow or block. That's like treating a search engine, a scraper and a customer as the same person because none of them filled in your newsletter form.

AI visitors come in distinct types. They're sent for different reasons, behave differently on the page and deserve different treatment. Block the wrong one and you vanish from AI answers. Welcome the wrong one and you pay for traffic that gives nothing back.

Here is the taxonomy we use in Ethogram, with what each type wants and how to think about it.

Answer engines

Who: assistants fetching a page because a person just asked them something. ChatGPT, Claude, Perplexity, Gemini and others, usually through user-triggered fetchers such as ChatGPT-User, Claude-User and Perplexity-User.

What they want: a fast, readable answer. The price, the spec, the opening hours, the returns window, the API parameter. They typically read one or a handful of pages and leave.

How they behave: mostly simple HTTP fetches without JavaScript. Visits are short and bursty, and they follow the questions people are asking: a product in the news can get read many times in an hour.

How to treat them: these are the closest thing to a referral you'll get from AI. Let them in. Make sure the facts people ask about are in the HTML, not rendered after load. Keep important pages at stable URLs. And keep facts consistent across pages, because an assistant quoting your old price from a forgotten landing page is worse than not being quoted at all.

One nuance: several providers document that their user-triggered fetchers don't always follow robots.txt the way crawlers do, on the reasoning that a person asked for that specific page. Check each provider's documentation before relying on robots.txt to control them.

Search crawlers

Who: crawlers that build the indexes behind AI search, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot. Google's AI features rely on the regular Googlebot index.

What they want: coverage and freshness. They visit many pages, revisit the ones that change and use that index to decide what to cite when someone asks a question later.

How they behave: much like classic search crawlers. Systematic, polite when well built, guided by your sitemap and robots.txt.

How to treat them: if you want to appear in AI search results and be cited as a source, allow them. Everything that helps classic SEO helps here: a sitemap, canonical URLs, server-rendered content, clear titles and descriptions. Watch their crawl rate on large sites, but blocking them is usually a decision to disappear from a growing share of search.

Training crawlers

Who: crawlers that collect content to train models, such as GPTBot, ClaudeBot, CCBot (Common Crawl, whose dataset is widely used for training), Bytespider and Meta-ExternalAgent. Some companies also offer robots.txt tokens that control training use without a separate crawler, such as Google-Extended and Applebot-Extended.

What they want: text, at volume. They're not answering anyone right now, and they don't send visitors back.

How they behave: broad, deep crawls. They can be heavy on bandwidth and are the main reason some sites started blocking "AI bots" wholesale.

How to treat them: this is a business decision, not a technical one. Some publishers block training to protect their content; many companies welcome it because they want models to know their product, their docs and their name. The useful thing to know is that the major providers separate training from search: blocking GPTBot doesn't remove you from ChatGPT search results, and blocking Google-Extended doesn't affect Google Search. You can decide each one on its merits, in robots.txt.

Browser agents

Who: agents that drive a real browser to complete a task for a person, like ChatGPT agent, Perplexity's Comet or Claude in Chrome, and the wider family of computer-use agents.

What they want: to finish the job. Find the product, pick the size, fill in the address, pay, book, sign up.

How they behave: like a very fast, very literal person. They run your JavaScript, click buttons, type into fields and read the page through its structure and accessibility tree as much as through what it looks like. When something doesn't respond, they try again, often several times, and then give the task back to their person.

How to treat them: these visits carry the most intent of any AI traffic, because someone delegated a real task. Don't block them by default. Make your flows usable without a mouse: real buttons and links, labelled form fields, native or properly accessible controls, error messages in text. Some browser agents now sign their requests, which lets you tell a legitimate agent from a scraper pretending to be one.

Personal agents

A growing branch of browser agents deserves its own name. Personal agents are always on, have their own cloud computer and browser, keep their own logins and work toward goals for one person over days or weeks. They might check your pricing every Monday, watch a product until it's back in stock, or run an outbound routine on a schedule.

They behave like loyal, very busy customers. They come back, they sign in, and they remember what failed last time. Ethogram groups them with browser agents in traffic reports and gives each one its own profile, including when in the day it visits.

The taxonomy at a glance

Type Sent by Wants Runs JavaScript Sensible default
Answer engines A person's question, right now One fact, fast Rarely Allow; put facts in HTML
Search crawlers An AI search index Coverage and freshness Rarely Allow; sitemap and canonicals
Training crawlers A model training pipeline Text at volume Rarely Your call, per crawler
Browser agents A person's task To complete it Yes Allow; make flows accessible
Personal agents One person, ongoing Goals over time Yes Allow; expect return visits

Name it before you decide

The common thread: you can't set a good policy for a visitor you can't identify. User agent strings are a start, but they're easy to fake, and browser agents often look like an ordinary browser. Reliable classification combines several signals: the declared user agent, published IP ranges, request signatures where they exist, and how the visitor actually behaves on the page.

That's what Ethogram does for every visit. Each agent is named, grouped by type and recorded step by step, so your robots.txt, your rate limits and your roadmap can be based on who actually visits, not on one bucket called "bots".

  • agent types
  • robots.txt
  • crawlers