# Track AI assistant traffic

> Install the tracking file on a Next.js site, check that it works, and read which AI assistants fetch your pages.

Source: https://docs.use-mark.com/docs/track-ai-assistant-traffic

AI assistants like ChatGPT, Claude, and Perplexity fetch web pages, but they
don't run the site tag, so the tag never sees them. A small file on your site
records each fetch in your PostHog project. **Performance** then shows the
hits in its **AI assistant traffic** section.

A hit means an assistant fetched a page. It doesn't mean the assistant cited
or recommended you.

## Before you start

* The site runs Next.js 16 or later. Other site builders aren't supported yet.
* PostHog is connected in **Connections** > **PostHog**. Without it, the
  site shows "Connect PostHog" instead of the file.

## Install the file

1. Open **Settings** > **Websites**.
2. Pick the website from the menu on its name if you have more than one.
3. Under **AI assistant tracking**, select **Copy** on the file.
4. Save it as `proxy.ts` at the root of the site's project and deploy. If the
   site already has a `proxy.ts`, the comment at the top of the file says how
   to merge the two.

The file already holds your PostHog project's public token, the same one the
site tag uses. Copy a fresh file from Mark rather than editing it by hand.

## Check the install

Select **Check AI assistant tracking**. Mark loads the site's home page as
**Mark-Probe**, then waits up to 30 seconds for that visit to reach PostHog.
Mark-Probe is Mark's own check. It never counts as an assistant.

| Result                                                 | What it means                                                                                                                                        |
| ------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Receiving**                                          | An AI assistant fetched a page in the last 7 days.                                                                                                   |
| **Installed**                                          | Mark's check reached PostHog. No assistant has visited in the last 7 days yet; on a small site that can take days.                                   |
| **Not installed**                                      | The site answered, but nothing reached PostHog. Deploy the file, or check again in a minute if you just did.                                         |
| **Blocked by the site's firewall**                     | The site refused Mark's check, so Mark can't tell. [Let Mark-Probe through](#let-assistants-and-mark-probe-through-your-firewall), then check again. |
| **Site didn't answer** or **Site couldn't be reached** | Mark couldn't load the home page. Check that the site is up.                                                                                         |
| **Check failed**                                       | Mark couldn't read your PostHog project. Check again in a few minutes.                                                                               |

Each website keeps its own last result under the button, with the time Mark
checked. If the page says **The check didn't finish**, Mark has no result for
that run; select the button again.

## Let assistants and Mark-Probe through your firewall

A firewall in front of your site runs before the tracking file. When it
blocks a request, the file never sees it and no hit is recorded. Mark's check
only tests **Mark-Probe**, so a firewall can block real assistants while the
check still says **Installed**. If **Performance** shows no hits for weeks,
check the firewall first.

Let these user agents through: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot,
Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, and
Mark-Probe. Rules that match on a user agent let in anyone who sends that
name, so allow only these.

### Vercel Firewall

1. Open the site's project in Vercel and select **Firewall**.
2. Select **Configure**, then **Add New** > **Rule**.
3. Under **If**, pick **User Agent**, choose a contains operator, and enter
   `Mark-Probe`. Add one condition per assistant name, joined with **OR**.
4. Under **Then**, pick **Bypass**.
5. Select **Save Rule**, then **Review Changes** and **Publish**.
6. Drag the rule above any rule that denies or challenges traffic.

A **Bypass** rule gets a request past your other custom rules and past
**Bot Protection**. Under **Bot Management**, keep **AI Bots Ruleset** on
**Allow**. **Attack Mode** lets known bots through, but it may still
challenge Mark-Probe, so check again after you turn it off.

### Cloudflare

1. In **Security Settings**, open **Configure AI bot policies**. Set
   **Search** and **Agent** to &#x2A;*Allow (do not block)**. Agent covers
   ChatGPT-User, Claude-User, and Perplexity-User, which fetch a page when a
   person asks about it. **Training** covers GPTBot and ClaudeBot; allow it if
   you want to count those too.
2. If you use **AI Crawl Control**, open **Crawlers** and set each assistant
   to **Allow**.
3. On the Free plan, turn off **Bot fight mode** under **Security Settings** >
   **Bot traffic**. Cloudflare doesn't let a rule skip it.
4. On Pro and above, add a skip rule instead. Open **Security rules** >
   **Create rule** > **Custom rules**, enter the expression
   `http.user_agent contains "Mark-Probe"`, pick **Skip**, and select
   **All Super Bot Fight Mode rules**, **All managed rules**, and
   **All remaining custom rules**. Add the assistant names with `or`. Place
   the rule first.

### Arcjet

Arcjet's rules live in your site's code, not a dashboard.

* In `detectBot`, add `"CATEGORY:AI"` to `allow`. It covers all eight
  assistants.
* Arcjet has no allow entry for Mark-Probe. Mark-Probe only loads your home
  page, so skip `protect` when the path is `/` and the user agent contains
  `Mark-Probe`. Or run the tracking file before Arcjet in your proxy, which
  records the hit even when Arcjet then denies the request.

After you change a firewall, select **Check AI assistant tracking** again.

## Read AI assistant traffic

Open **Performance** and pick a range. The **AI assistant traffic** section
shows:

* **Hits**: every assistant fetch in the range, split into **Verified** and
  **Unverified**.
* Hits, verified, and unverified per assistant, split by kind:
  * **Crawler** collects pages for training or an index (GPTBot, ClaudeBot,
    PerplexityBot).
  * **Search** builds the assistant's own search results (OAI-SearchBot,
    Claude-SearchBot).
  * **User fetch** loads a page because a person asked the assistant about it
    (ChatGPT-User, Claude-User, Perplexity-User).
* **Top pages**: the 10 pages fetched most.
* A "Data through" line with the newest hit.

With no hits in the range, the section says so and links to the Website
connection.

Select a path in **Top pages** to see that one page. Each assistant gets a
row with its hits on the page in the range and the date of its newest fetch,
or "None in range". Pick a 30-day range to see which assistants haven't
fetched the page in 30 days. **All pages** goes back, and the range and
filters stay as you set them. The one-page view shows hits only, not
verified and unverified.

A hit counts when the request's user agent names an assistant. Anyone can send
a request that claims to be GPTBot, so Mark also checks where it came from.

## Verified and unverified hits

OpenAI, Anthropic, and Perplexity each publish the IP addresses their bots
use. A hit is **verified** when it came from an address on the list for its
assistant. Mark reloads the lists once a day.

A hit is **unverified** when Mark can't match it to a list. That happens when:

* Someone sent a fake assistant user agent.
* The request had no IP address.
* The vendor's list couldn't be loaded. The section says so. Mark checks the
  same hits again once the list loads.
* The site's PostHog project has **Discard client IP data** turned on. Every
  hit is then unverified.

Anthropic publishes one list for all its bots. A verified ClaudeBot,
Claude-SearchBot, or Claude-User hit came from Anthropic, but the list can't
show which of the three sent it. OpenAI and Perplexity publish a list per bot.

Verified doesn't mean the assistant cited or recommended your site. It only
means the fetch really came from that company.
