AI and Search Crawler Tracking
Monitor when AI models, user-triggered answer engines, and search crawlers fetch and index your pages.
Agent guidance
- Crawler reporting runs on your server or edge, not in the browser. Do not add the tracking snippet expecting it to see crawlers: crawlers do not run JavaScript.
- The reporter uses a Server API Token with the
write:crawler_hitsscope. Keep it in a server environment variable namedSOURCETRACK_API_KEY. Never paste it into a chat, a prompt, a front-end file or any client-side code, and never commit it. - The public site key from the tracking snippet is a different value and cannot report crawler hits.
What this is
Standard client-side web tracking pixels only execute when a human visitor opens a page in a full browser environment with JavaScript enabled. Automated LLM training bots, retrieval fetchers (such as ChatGPT browsing or Claude search), and search engine indexers do not execute client-side scripts when fetching your HTML.
So crawler reporting needs code on your server or at your edge. Paste one of the snippets below. It recognises known AI and search crawlers (including GPTBot, ClaudeBot, PerplexityBot, Applebot, Googlebot and Bingbot) and reports each fetch. Each snippet is a complete file: nothing to install, and the crawler list updates itself from SourceTrack.
Architecture and privacy guarantees:
- Off the request path: The JavaScript snippets report after your response is sent (Express
finish, orwaitUntil()on Next.js and edge runtimes). The PHP snippet runs in a shutdown function after the page is generated, with a 3-second network cap. - Fail-open: Any network or internal error degrades silently to a no-op so your application never crashes.
- Zero PII: The payload carries only the crawler name, a verification label, the path (no query string), the status code and a timestamp. No IP address or User-Agent is ever sent. If you pass a trusted client IP, it is checked locally against the crawler's published IP ranges and then discarded.
Where this works
Crawler visibility needs access to your server or edge. A hosted store or site builder that runs none of your code cannot report crawler fetches, and the tracking snippet cannot either.
| Platform | Crawler reporting | How |
|---|---|---|
| Node.js / Express | Available | Express snippet. |
| Next.js (Vercel or self-hosted) | Available | Next.js snippet. The status code is not available because the proxy runs before the page renders. |
| Cloudflare Workers / Pages, Hono | Available | Request/Response snippet. |
| PHP (Laravel, plain PHP) | Available | PHP snippet. |
| WordPress | Available | PHP snippet as a must-use plugin, or the SourceTrack crawler plugin. Pages served from a page cache or CDN are not seen. |
| Webflow, Squarespace, Framer and other hosted builders | Only behind Cloudflare | Only if your domain is proxied through Cloudflare: deploy the Request/Response snippet as a Cloudflare Worker. Webflow supports this officially; check your builder before relying on it. |
| Shopify (not Hydrogen) | Not available | Not available. Shopify does not run your server code and does not support a Cloudflare proxy in front of the store. |
| Wix | Not available | Not available. Wix does not run your server code and does not support proxied Cloudflare records. |
Step 1: Generate a Server API Token
Crawler reporting requires a dedicated Server API Token with the write:crawler_hits scope.
- Go to the app at Settings, Advanced, Server API Tokens (https://app.sourcetrack.ai/settings).
- Click Create Token.
- Give your token a descriptive name (for example
Production Crawler Middleware) and check the write:crawler_hits permission box. - Copy the generated token (
st_live_...) and save it in your server's environment variables asSOURCETRACK_API_KEY.
Scope requirement: The token must carry the write:crawler_hits scope. A token created with standard event permissions (write:events) will be rejected with a 403 Forbidden status on crawler ingestion routes.
Step 2: Paste the snippet for your platform
Save the file next to your code and wire it in as shown. It reads the token from SOURCETRACK_API_KEY. Without a trusted client IP, hits are labelled user-agent only; each snippet shows where to pass one.
Express / Node.js
Save as sourcetrack-crawler.js:
// sourcetrack-crawler.js: SourceTrack AI crawler reporting for Express (ES module).
// ── SourceTrack AI crawler reporter ─────────────────────────────────────────
// Self-contained: nothing to install. Needs a SourceTrack API key with the
// write:crawler_hits scope.
// Sent to SourceTrack: the crawler's name, a verification verdict, the path
// (query string removed), the status code and a timestamp. Nothing else.
// Never sent: IP address, User-Agent, query string, cookies, headers, body.
// The crawler list and published IP ranges come from SourceTrack
// (cached for 6 hours), so new crawlers are picked up without re-pasting.
// IP verification runs here, locally, only with an IP you pass in.
// Every error is swallowed: a failure never affects your response.
const ST_API_BASE = 'https://api.srctk.com'
const ST_TTL_MS = 6 * 60 * 60 * 1000
const ST_RETRY_MS = 5 * 60 * 1000
const ST_TIMEOUT_MS = 5000
const ST_MAX_PENDING = 20
let stRegistry = null
let stFetchedAt = 0
let stRetryAt = 0
let stInFlight = null
let stPending = 0
function stSignal () {
return typeof AbortSignal !== 'undefined' && typeof AbortSignal.timeout === 'function'
? AbortSignal.timeout(ST_TIMEOUT_MS)
: undefined
}
function stSafePath (value) {
const raw = typeof value === 'string' && value.length > 0 ? value : '/'
const cut = raw.split('?')[0].split('#')[0]
return (cut || '/').slice(0, 512)
}
function stParseIp (ip) {
const raw = String(ip || '').trim().toLowerCase()
if (!raw) return null
let acc = BigInt(0)
if (raw.includes(':')) {
if (raw.includes('.')) return null
const halves = raw.split('::')
if (halves.length > 2) return null
const head = halves[0] ? halves[0].split(':') : []
const tail = halves.length === 2 && halves[1] ? halves[1].split(':') : []
if (halves.length === 1 && head.length !== 8) return null
const fill = 8 - head.length - tail.length
if (fill < 0) return null
const groups = head.concat(new Array(halves.length === 2 ? fill : 0).fill('0'), tail)
if (groups.length !== 8) return null
for (const g of groups) {
if (!/^[0-9a-f]{1,4}$/.test(g)) return null
acc = (acc << BigInt(16)) | BigInt(parseInt(g, 16))
}
return { value: acc, bits: 128 }
}
const parts = raw.split('.')
if (parts.length !== 4) return null
for (const p of parts) {
if (!/^\d{1,3}$/.test(p) || Number(p) > 255) return null
acc = (acc << BigInt(8)) | BigInt(Number(p))
}
return { value: acc, bits: 32 }
}
function stIpInCidr (ip, cidr) {
try {
const parts = String(cidr || '').trim().split('/')
const prefix = Number(parts[1])
const target = stParseIp(ip)
const base = stParseIp(parts[0])
if (!target || !base || target.bits !== base.bits) return false
if (!Number.isInteger(prefix) || prefix < 0 || prefix > target.bits) return false
const host = BigInt(target.bits - prefix)
return (target.value >> host) === (base.value >> host)
} catch (_) {
return false
}
}
async function stFetchRegistry (base, apiKey, doFetch) {
try {
const res = await doFetch(base + '/api/server/crawler-ranges', {
method: 'GET',
headers: { Authorization: 'Bearer ' + apiKey },
signal: stSignal()
})
const body = res && res.ok ? await res.json() : null
const data = body && body.data ? body.data : {}
const bots = []
for (const b of Array.isArray(data.registry) ? data.registry : []) {
if (b && typeof b.token === 'string' && b.token && typeof b.name === 'string' && b.name) {
bots.push({ token: b.token, needle: b.token.toLowerCase(), name: b.name, rangesPublished: b.ranges_published === true })
}
}
if (bots.length === 0) {
stRetryAt = Date.now() + ST_RETRY_MS
return stRegistry
}
const ranges = new Map()
const served = data.bots && typeof data.bots === 'object' ? data.bots : {}
for (const token of Object.keys(served)) {
if (Array.isArray(served[token]) && served[token].length > 0) ranges.set(token, served[token])
}
stRegistry = { bots, ranges }
// No ranges served (the list is stale on our side): keep the crawler list, ask again soon.
stFetchedAt = ranges.size > 0 ? Date.now() : Date.now() - ST_TTL_MS + ST_RETRY_MS
return stRegistry
} catch (_) {
stRetryAt = Date.now() + ST_RETRY_MS
return stRegistry
}
}
function stLoadRegistry (base, apiKey, doFetch) {
const now = Date.now()
if (stRegistry && now - stFetchedAt < ST_TTL_MS) return Promise.resolve(stRegistry)
if (now < stRetryAt) return Promise.resolve(stRegistry)
if (stInFlight) return stInFlight
const pending = stFetchRegistry(base, apiKey, doFetch)
stInFlight = pending
pending.then(() => { if (stInFlight === pending) stInFlight = null })
return pending
}
function stClassify (registry, userAgent, clientIp) {
const ua = String(userAgent || '').toLowerCase()
if (!ua) return null
const bot = registry.bots.find((b) => ua.includes(b.needle))
if (!bot) return null
if (!bot.rangesPublished) return { name: bot.name, verification: 'ua_only_no_ranges_published' }
const cidrs = registry.ranges.get(bot.token)
if (!clientIp || !stParseIp(clientIp) || !cidrs) return { name: bot.name, verification: 'ua_only' }
const matched = cidrs.some((cidr) => stIpInCidr(clientIp, cidr))
return { name: bot.name, verification: matched ? 'ip_verified' : 'ip_mismatch' }
}
async function stReport (options, input) {
try {
const apiKey = options && options.apiKey
if (!apiKey) return
const doFetch = options.fetch || fetch
const base = String(options.apiBase || ST_API_BASE).replace(/\/+$/, '')
const registry = await stLoadRegistry(base, apiKey, doFetch)
if (!registry) return
const hit = stClassify(registry, input.userAgent, input.clientIp)
if (!hit || stPending >= ST_MAX_PENDING) return
stPending++
try {
await doFetch(base + '/api/server/crawler-hit', {
method: 'POST',
headers: { 'Content-Type': 'application/json', Authorization: 'Bearer ' + apiKey },
body: JSON.stringify({
bot_name: hit.name,
verification: hit.verification,
path: stSafePath(input.path),
status_code: Number.isInteger(input.statusCode) ? input.statusCode : 0,
collection_source: 'edge_middleware',
timestamp: new Date().toISOString()
}),
signal: stSignal()
})
} finally {
stPending--
}
} catch (_) {
// Fail-open: reporting never affects the site.
}
}
function stClientIp (getClientIp, request) {
try {
return typeof getClientIp === 'function' ? getClientIp(request) : null
} catch (_) {
return null
}
}
// Express middleware. next() is called at once; the report runs after the
// response has been sent, so the status code is the real one.
export function sourcetrackCrawler (options) {
const opts = options || {}
return function sourcetrackCrawlerMiddleware (req, res, next) {
try {
res.once('finish', function () {
setImmediate(function () {
stReport(opts, {
userAgent: req.headers['user-agent'],
path: req.originalUrl || req.url,
statusCode: res.statusCode,
clientIp: stClientIp(opts.getClientIp, req)
})
})
})
} catch (_) {
// Fail-open.
}
next()
}
}
Register it in your app:
import express from 'express'
import { sourcetrackCrawler } from './sourcetrack-crawler.js'
const app = express()
// Register before your routes.
app.use(sourcetrackCrawler({
apiKey: process.env.SOURCETRACK_API_KEY
// Optional IP verification. Only if req.ip is the real visitor address: no proxy
// in front of Express, or app.set('trust proxy', ...) set for your load balancer.
// getClientIp: (req) => req.ip
}))
Next.js (proxy.js / middleware.js)
Save as proxy.js:
// proxy.js (Next.js 16+) or middleware.js (Next.js 15 and older; rename the
// exported function from proxy to middleware). Put it at the project root, next
// to app/ or pages/. SourceTrack AI crawler reporting, self-contained.
// ── SourceTrack AI crawler reporter ─────────────────────────────────────────
// Self-contained: nothing to install. Needs a SourceTrack API key with the
// write:crawler_hits scope.
// Sent to SourceTrack: the crawler's name, a verification verdict, the path
// (query string removed), the status code and a timestamp. Nothing else.
// Never sent: IP address, User-Agent, query string, cookies, headers, body.
// The crawler list and published IP ranges come from SourceTrack
// (cached for 6 hours), so new crawlers are picked up without re-pasting.
// IP verification runs here, locally, only with an IP you pass in.
// Every error is swallowed: a failure never affects your response.
const ST_API_BASE = 'https://api.srctk.com'
const ST_TTL_MS = 6 * 60 * 60 * 1000
const ST_RETRY_MS = 5 * 60 * 1000
const ST_TIMEOUT_MS = 5000
const ST_MAX_PENDING = 20
let stRegistry = null
let stFetchedAt = 0
let stRetryAt = 0
let stInFlight = null
let stPending = 0
function stSignal () {
return typeof AbortSignal !== 'undefined' && typeof AbortSignal.timeout === 'function'
? AbortSignal.timeout(ST_TIMEOUT_MS)
: undefined
}
function stSafePath (value) {
const raw = typeof value === 'string' && value.length > 0 ? value : '/'
const cut = raw.split('?')[0].split('#')[0]
return (cut || '/').slice(0, 512)
}
function stParseIp (ip) {
const raw = String(ip || '').trim().toLowerCase()
if (!raw) return null
let acc = BigInt(0)
if (raw.includes(':')) {
if (raw.includes('.')) return null
const halves = raw.split('::')
if (halves.length > 2) return null
const head = halves[0] ? halves[0].split(':') : []
const tail = halves.length === 2 && halves[1] ? halves[1].split(':') : []
if (halves.length === 1 && head.length !== 8) return null
const fill = 8 - head.length - tail.length
if (fill < 0) return null
const groups = head.concat(new Array(halves.length === 2 ? fill : 0).fill('0'), tail)
if (groups.length !== 8) return null
for (const g of groups) {
if (!/^[0-9a-f]{1,4}$/.test(g)) return null
acc = (acc << BigInt(16)) | BigInt(parseInt(g, 16))
}
return { value: acc, bits: 128 }
}
const parts = raw.split('.')
if (parts.length !== 4) return null
for (const p of parts) {
if (!/^\d{1,3}$/.test(p) || Number(p) > 255) return null
acc = (acc << BigInt(8)) | BigInt(Number(p))
}
return { value: acc, bits: 32 }
}
function stIpInCidr (ip, cidr) {
try {
const parts = String(cidr || '').trim().split('/')
const prefix = Number(parts[1])
const target = stParseIp(ip)
const base = stParseIp(parts[0])
if (!target || !base || target.bits !== base.bits) return false
if (!Number.isInteger(prefix) || prefix < 0 || prefix > target.bits) return false
const host = BigInt(target.bits - prefix)
return (target.value >> host) === (base.value >> host)
} catch (_) {
return false
}
}
async function stFetchRegistry (base, apiKey, doFetch) {
try {
const res = await doFetch(base + '/api/server/crawler-ranges', {
method: 'GET',
headers: { Authorization: 'Bearer ' + apiKey },
signal: stSignal()
})
const body = res && res.ok ? await res.json() : null
const data = body && body.data ? body.data : {}
const bots = []
for (const b of Array.isArray(data.registry) ? data.registry : []) {
if (b && typeof b.token === 'string' && b.token && typeof b.name === 'string' && b.name) {
bots.push({ token: b.token, needle: b.token.toLowerCase(), name: b.name, rangesPublished: b.ranges_published === true })
}
}
if (bots.length === 0) {
stRetryAt = Date.now() + ST_RETRY_MS
return stRegistry
}
const ranges = new Map()
const served = data.bots && typeof data.bots === 'object' ? data.bots : {}
for (const token of Object.keys(served)) {
if (Array.isArray(served[token]) && served[token].length > 0) ranges.set(token, served[token])
}
stRegistry = { bots, ranges }
// No ranges served (the list is stale on our side): keep the crawler list, ask again soon.
stFetchedAt = ranges.size > 0 ? Date.now() : Date.now() - ST_TTL_MS + ST_RETRY_MS
return stRegistry
} catch (_) {
stRetryAt = Date.now() + ST_RETRY_MS
return stRegistry
}
}
function stLoadRegistry (base, apiKey, doFetch) {
const now = Date.now()
if (stRegistry && now - stFetchedAt < ST_TTL_MS) return Promise.resolve(stRegistry)
if (now < stRetryAt) return Promise.resolve(stRegistry)
if (stInFlight) return stInFlight
const pending = stFetchRegistry(base, apiKey, doFetch)
stInFlight = pending
pending.then(() => { if (stInFlight === pending) stInFlight = null })
return pending
}
function stClassify (registry, userAgent, clientIp) {
const ua = String(userAgent || '').toLowerCase()
if (!ua) return null
const bot = registry.bots.find((b) => ua.includes(b.needle))
if (!bot) return null
if (!bot.rangesPublished) return { name: bot.name, verification: 'ua_only_no_ranges_published' }
const cidrs = registry.ranges.get(bot.token)
if (!clientIp || !stParseIp(clientIp) || !cidrs) return { name: bot.name, verification: 'ua_only' }
const matched = cidrs.some((cidr) => stIpInCidr(clientIp, cidr))
return { name: bot.name, verification: matched ? 'ip_verified' : 'ip_mismatch' }
}
async function stReport (options, input) {
try {
const apiKey = options && options.apiKey
if (!apiKey) return
const doFetch = options.fetch || fetch
const base = String(options.apiBase || ST_API_BASE).replace(/\/+$/, '')
const registry = await stLoadRegistry(base, apiKey, doFetch)
if (!registry) return
const hit = stClassify(registry, input.userAgent, input.clientIp)
if (!hit || stPending >= ST_MAX_PENDING) return
stPending++
try {
await doFetch(base + '/api/server/crawler-hit', {
method: 'POST',
headers: { 'Content-Type': 'application/json', Authorization: 'Bearer ' + apiKey },
body: JSON.stringify({
bot_name: hit.name,
verification: hit.verification,
path: stSafePath(input.path),
status_code: Number.isInteger(input.statusCode) ? input.statusCode : 0,
collection_source: 'edge_middleware',
timestamp: new Date().toISOString()
}),
signal: stSignal()
})
} finally {
stPending--
}
} catch (_) {
// Fail-open: reporting never affects the site.
}
}
function stClientIp (getClientIp, request) {
try {
return typeof getClientIp === 'function' ? getClientIp(request) : null
} catch (_) {
return null
}
}
const ST_OPTIONS = {
apiKey: process.env.SOURCETRACK_API_KEY
// Optional IP verification. Only on a host that overwrites x-real-ip with the
// visitor's address (Vercel documents that it does). Leave it out elsewhere.
// getClientIp: (request) => request.headers.get('x-real-ip')
}
// Runs before the page renders, so the status code is not known and is sent as 0
// ("not reported"). The report runs in waitUntil, after the response.
export function proxy (request, event) {
try {
const work = stReport(ST_OPTIONS, {
userAgent: request.headers.get('user-agent'),
path: request.nextUrl ? request.nextUrl.pathname : new URL(request.url).pathname,
statusCode: 0,
clientIp: stClientIp(ST_OPTIONS.getClientIp, request)
})
if (event && typeof event.waitUntil === 'function') event.waitUntil(work)
} catch (_) {
// Fail-open: returning nothing continues to your route.
}
}
// robots.txt, sitemap.xml and llms.txt are NOT excluded: crawlers fetch them most.
export const config = {
matcher: ['/((?!api|_next/static|_next/image|favicon.ico).*)']
}
Cloudflare Workers / Pages, Hono and other Request/Response runtimes
Save as sourcetrack-crawler.js:
// sourcetrack-crawler.js: SourceTrack AI crawler reporting for fetch-style
// runtimes (Cloudflare Workers and Pages, Hono, Vercel Edge, Deno, Bun).
// ── SourceTrack AI crawler reporter ─────────────────────────────────────────
// Self-contained: nothing to install. Needs a SourceTrack API key with the
// write:crawler_hits scope.
// Sent to SourceTrack: the crawler's name, a verification verdict, the path
// (query string removed), the status code and a timestamp. Nothing else.
// Never sent: IP address, User-Agent, query string, cookies, headers, body.
// The crawler list and published IP ranges come from SourceTrack
// (cached for 6 hours), so new crawlers are picked up without re-pasting.
// IP verification runs here, locally, only with an IP you pass in.
// Every error is swallowed: a failure never affects your response.
const ST_API_BASE = 'https://api.srctk.com'
const ST_TTL_MS = 6 * 60 * 60 * 1000
const ST_RETRY_MS = 5 * 60 * 1000
const ST_TIMEOUT_MS = 5000
const ST_MAX_PENDING = 20
let stRegistry = null
let stFetchedAt = 0
let stRetryAt = 0
let stInFlight = null
let stPending = 0
function stSignal () {
return typeof AbortSignal !== 'undefined' && typeof AbortSignal.timeout === 'function'
? AbortSignal.timeout(ST_TIMEOUT_MS)
: undefined
}
function stSafePath (value) {
const raw = typeof value === 'string' && value.length > 0 ? value : '/'
const cut = raw.split('?')[0].split('#')[0]
return (cut || '/').slice(0, 512)
}
function stParseIp (ip) {
const raw = String(ip || '').trim().toLowerCase()
if (!raw) return null
let acc = BigInt(0)
if (raw.includes(':')) {
if (raw.includes('.')) return null
const halves = raw.split('::')
if (halves.length > 2) return null
const head = halves[0] ? halves[0].split(':') : []
const tail = halves.length === 2 && halves[1] ? halves[1].split(':') : []
if (halves.length === 1 && head.length !== 8) return null
const fill = 8 - head.length - tail.length
if (fill < 0) return null
const groups = head.concat(new Array(halves.length === 2 ? fill : 0).fill('0'), tail)
if (groups.length !== 8) return null
for (const g of groups) {
if (!/^[0-9a-f]{1,4}$/.test(g)) return null
acc = (acc << BigInt(16)) | BigInt(parseInt(g, 16))
}
return { value: acc, bits: 128 }
}
const parts = raw.split('.')
if (parts.length !== 4) return null
for (const p of parts) {
if (!/^\d{1,3}$/.test(p) || Number(p) > 255) return null
acc = (acc << BigInt(8)) | BigInt(Number(p))
}
return { value: acc, bits: 32 }
}
function stIpInCidr (ip, cidr) {
try {
const parts = String(cidr || '').trim().split('/')
const prefix = Number(parts[1])
const target = stParseIp(ip)
const base = stParseIp(parts[0])
if (!target || !base || target.bits !== base.bits) return false
if (!Number.isInteger(prefix) || prefix < 0 || prefix > target.bits) return false
const host = BigInt(target.bits - prefix)
return (target.value >> host) === (base.value >> host)
} catch (_) {
return false
}
}
async function stFetchRegistry (base, apiKey, doFetch) {
try {
const res = await doFetch(base + '/api/server/crawler-ranges', {
method: 'GET',
headers: { Authorization: 'Bearer ' + apiKey },
signal: stSignal()
})
const body = res && res.ok ? await res.json() : null
const data = body && body.data ? body.data : {}
const bots = []
for (const b of Array.isArray(data.registry) ? data.registry : []) {
if (b && typeof b.token === 'string' && b.token && typeof b.name === 'string' && b.name) {
bots.push({ token: b.token, needle: b.token.toLowerCase(), name: b.name, rangesPublished: b.ranges_published === true })
}
}
if (bots.length === 0) {
stRetryAt = Date.now() + ST_RETRY_MS
return stRegistry
}
const ranges = new Map()
const served = data.bots && typeof data.bots === 'object' ? data.bots : {}
for (const token of Object.keys(served)) {
if (Array.isArray(served[token]) && served[token].length > 0) ranges.set(token, served[token])
}
stRegistry = { bots, ranges }
// No ranges served (the list is stale on our side): keep the crawler list, ask again soon.
stFetchedAt = ranges.size > 0 ? Date.now() : Date.now() - ST_TTL_MS + ST_RETRY_MS
return stRegistry
} catch (_) {
stRetryAt = Date.now() + ST_RETRY_MS
return stRegistry
}
}
function stLoadRegistry (base, apiKey, doFetch) {
const now = Date.now()
if (stRegistry && now - stFetchedAt < ST_TTL_MS) return Promise.resolve(stRegistry)
if (now < stRetryAt) return Promise.resolve(stRegistry)
if (stInFlight) return stInFlight
const pending = stFetchRegistry(base, apiKey, doFetch)
stInFlight = pending
pending.then(() => { if (stInFlight === pending) stInFlight = null })
return pending
}
function stClassify (registry, userAgent, clientIp) {
const ua = String(userAgent || '').toLowerCase()
if (!ua) return null
const bot = registry.bots.find((b) => ua.includes(b.needle))
if (!bot) return null
if (!bot.rangesPublished) return { name: bot.name, verification: 'ua_only_no_ranges_published' }
const cidrs = registry.ranges.get(bot.token)
if (!clientIp || !stParseIp(clientIp) || !cidrs) return { name: bot.name, verification: 'ua_only' }
const matched = cidrs.some((cidr) => stIpInCidr(clientIp, cidr))
return { name: bot.name, verification: matched ? 'ip_verified' : 'ip_mismatch' }
}
async function stReport (options, input) {
try {
const apiKey = options && options.apiKey
if (!apiKey) return
const doFetch = options.fetch || fetch
const base = String(options.apiBase || ST_API_BASE).replace(/\/+$/, '')
const registry = await stLoadRegistry(base, apiKey, doFetch)
if (!registry) return
const hit = stClassify(registry, input.userAgent, input.clientIp)
if (!hit || stPending >= ST_MAX_PENDING) return
stPending++
try {
await doFetch(base + '/api/server/crawler-hit', {
method: 'POST',
headers: { 'Content-Type': 'application/json', Authorization: 'Bearer ' + apiKey },
body: JSON.stringify({
bot_name: hit.name,
verification: hit.verification,
path: stSafePath(input.path),
status_code: Number.isInteger(input.statusCode) ? input.statusCode : 0,
collection_source: 'edge_middleware',
timestamp: new Date().toISOString()
}),
signal: stSignal()
})
} finally {
stPending--
}
} catch (_) {
// Fail-open: reporting never affects the site.
}
}
function stClientIp (getClientIp, request) {
try {
return typeof getClientIp === 'function' ? getClientIp(request) : null
} catch (_) {
return null
}
}
// Call after you have the response and hand the returned promise to your
// runtime's waitUntil. It never rejects.
// options.apiKey SourceTrack API key (write:crawler_hits)
// options.status the response status code
// options.clientIp optional; only an address your platform sets itself,
// e.g. Cloudflare's CF-Connecting-IP header
export function reportCrawlerRequest (request, options) {
const opts = options || {}
try {
let path = '/'
try { path = new URL(request.url).pathname } catch (_) {}
return stReport(opts, {
userAgent: request.headers.get('user-agent'),
path,
statusCode: Number(opts.status),
clientIp: opts.clientIp || null
})
} catch (_) {
return Promise.resolve()
}
}
Cloudflare Worker in front of any site:
// worker.js: a Cloudflare Worker on your site's route (the DNS record must be
// proxied, "orange cloud"). Store the key as a secret: wrangler secret put SOURCETRACK_API_KEY
import { reportCrawlerRequest } from './sourcetrack-crawler.js'
export default {
async fetch (request, env, ctx) {
const response = await fetch(request) // passes the request on to your site
ctx.waitUntil(reportCrawlerRequest(request, {
apiKey: env.SOURCETRACK_API_KEY,
status: response.status,
clientIp: request.headers.get('CF-Connecting-IP') // set by Cloudflare
}))
return response
}
}
Cloudflare Pages:
// functions/_middleware.js on Cloudflare Pages
import { reportCrawlerRequest } from '../sourcetrack-crawler.js'
export async function onRequest (context) {
const response = await context.next()
context.waitUntil(reportCrawlerRequest(context.request, {
apiKey: context.env.SOURCETRACK_API_KEY,
status: response.status,
clientIp: context.request.headers.get('CF-Connecting-IP')
}))
return response
}
Hono on Cloudflare Workers:
// Hono on Cloudflare Workers
import { reportCrawlerRequest } from './sourcetrack-crawler.js'
app.use('*', async (c, next) => {
await next()
c.executionCtx.waitUntil(reportCrawlerRequest(c.req.raw, {
apiKey: c.env.SOURCETRACK_API_KEY,
status: c.res.status,
clientIp: c.req.header('CF-Connecting-IP')
}))
})
PHP (Laravel, plain PHP, WordPress must-use plugin)
Save as sourcetrack-crawler.php:
<?php
/**
* sourcetrack-crawler.php: SourceTrack AI crawler reporting for PHP.
* Self-contained, nothing to install. Needs a SourceTrack API key with the
* write:crawler_hits scope, in the SOURCETRACK_API_KEY environment variable or
* define('SOURCETRACK_API_KEY', '...') (on WordPress, in wp-config.php).
*
* Install: require it at the top of your front controller (index.php), or on
* WordPress save it as wp-content/mu-plugins/sourcetrack-crawler.php.
*
* Sent: the crawler's name, a verification verdict, the path (query string
* removed), the status code and a timestamp. Never sent: IP address,
* User-Agent, query string, cookies. The crawler list comes from SourceTrack
* and is cached in the system temp directory for 6 hours.
*
* Optional IP verification: define('SOURCETRACK_CLIENT_IP_KEY', 'REMOTE_ADDR')
* when nothing sits in front of PHP, or 'HTTP_CF_CONNECTING_IP' behind
* Cloudflare. Without it, hits are reported as 'ua_only'.
*
* Runs in a shutdown function after your page is generated. Every error is
* swallowed. Only requests that reach PHP are seen: files your web server
* serves directly and pages served from a CDN or page cache are not.
*/
if (!function_exists('sourcetrack_crawler_report')) {
function sourcetrack_crawler_config()
{
$key = getenv('SOURCETRACK_API_KEY');
if (!is_string($key) || $key === '') {
$key = defined('SOURCETRACK_API_KEY') ? (string) constant('SOURCETRACK_API_KEY') : '';
}
$base = getenv('SOURCETRACK_API_BASE');
return array(
'api_key' => trim($key),
'api_base' => rtrim(is_string($base) && $base !== '' ? $base : 'https://api.srctk.com', '/'),
'ip_key' => defined('SOURCETRACK_CLIENT_IP_KEY') ? (string) constant('SOURCETRACK_CLIENT_IP_KEY') : '',
);
}
function sourcetrack_crawler_http($method, $url, $api_key, $body = null)
{
$headers = array('Authorization: Bearer ' . $api_key);
if ($body !== null) {
$headers[] = 'Content-Type: application/json';
}
if (function_exists('curl_init')) {
$ch = curl_init($url);
curl_setopt_array($ch, array(
CURLOPT_CUSTOMREQUEST => $method,
CURLOPT_HTTPHEADER => $headers,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_CONNECTTIMEOUT => 2,
CURLOPT_TIMEOUT => 3,
CURLOPT_FOLLOWLOCATION => false,
));
if ($body !== null) {
curl_setopt($ch, CURLOPT_POSTFIELDS, $body);
}
$out = curl_exec($ch);
$code = (int) curl_getinfo($ch, CURLINFO_HTTP_CODE);
return array($code, is_string($out) ? $out : '');
}
$context = stream_context_create(array('http' => array(
'method' => $method,
'header' => implode("\r\n", $headers),
'content' => $body === null ? '' : $body,
'timeout' => 3,
'ignore_errors' => true,
'follow_location' => 0,
)));
$out = @file_get_contents($url, false, $context);
$code = 0;
if (isset($http_response_header[0]) && preg_match('#\s(\d{3})(\s|$)#', $http_response_header[0], $m)) {
$code = (int) $m[1];
}
return array($code, is_string($out) ? $out : '');
}
function sourcetrack_crawler_cache_file($cfg, $name)
{
return rtrim(sys_get_temp_dir(), '/\\') . '/sourcetrack-crawler-' . md5($cfg['api_base']) . '-' . $name . '.json';
}
function sourcetrack_crawler_write($file, $data)
{
$tmp = $file . '.' . getmypid() . '.tmp';
if (@file_put_contents($tmp, json_encode($data)) !== false) {
@rename($tmp, $file);
}
}
// The crawler list (small, read on every request). The ranges are written to a
// second file that is only read when a crawler matched and an IP is configured.
function sourcetrack_crawler_registry($cfg)
{
$file = sourcetrack_crawler_cache_file($cfg, 'registry');
$raw = @file_get_contents($file);
$cached = is_string($raw) && $raw !== '' ? json_decode($raw, true) : null;
$now = time();
if (is_array($cached) && isset($cached['refresh_at']) && $now < (int) $cached['refresh_at']) {
return $cached;
}
list($code, $body) = sourcetrack_crawler_http('GET', $cfg['api_base'] . '/api/server/crawler-ranges', $cfg['api_key']);
$json = $code === 200 ? json_decode($body, true) : null;
$bots = array();
if (is_array($json) && isset($json['data']['registry']) && is_array($json['data']['registry'])) {
foreach ($json['data']['registry'] as $b) {
if (is_array($b) && isset($b['token'], $b['name']) && is_string($b['token']) && is_string($b['name']) && $b['token'] !== '' && $b['name'] !== '') {
$bots[] = array('token' => $b['token'], 'name' => $b['name'], 'ranges_published' => !empty($b['ranges_published']));
}
}
}
if (count($bots) === 0) {
// Failed: keep what we had and try again in 5 minutes, not on every request.
$fresh = is_array($cached) && isset($cached['bots']) ? $cached : array('bots' => array());
$fresh['refresh_at'] = $now + 300;
sourcetrack_crawler_write($file, $fresh);
return $fresh;
}
$ranges = isset($json['data']['bots']) && is_array($json['data']['bots']) ? $json['data']['bots'] : array();
sourcetrack_crawler_write(sourcetrack_crawler_cache_file($cfg, 'ranges'), $ranges);
$fresh = array('bots' => $bots, 'refresh_at' => $now + (count($ranges) > 0 ? 21600 : 300));
sourcetrack_crawler_write($file, $fresh);
return $fresh;
}
function sourcetrack_crawler_ip_in_cidr($ip, $cidr)
{
$parts = explode('/', trim($cidr), 2);
if (count($parts) !== 2 || !ctype_digit($parts[1])) {
return false;
}
$ip_bin = @inet_pton($ip);
$net_bin = @inet_pton($parts[0]);
if ($ip_bin === false || $net_bin === false || strlen($ip_bin) !== strlen($net_bin)) {
return false;
}
$prefix = (int) $parts[1];
if ($prefix > strlen($ip_bin) * 8) {
return false;
}
$bytes = intdiv($prefix, 8);
$bits = $prefix % 8;
if (substr($ip_bin, 0, $bytes) !== substr($net_bin, 0, $bytes)) {
return false;
}
if ($bits === 0) {
return true;
}
$mask = (0xFF << (8 - $bits)) & 0xFF;
return (ord($ip_bin[$bytes]) & $mask) === (ord($net_bin[$bytes]) & $mask);
}
function sourcetrack_crawler_report()
{
try {
$cfg = sourcetrack_crawler_config();
$ua = isset($_SERVER['HTTP_USER_AGENT']) ? strtolower((string) $_SERVER['HTTP_USER_AGENT']) : '';
if ($cfg['api_key'] === '' || $ua === '') {
return;
}
$registry = sourcetrack_crawler_registry($cfg);
$bot = null;
foreach ($registry['bots'] as $b) {
if (strpos($ua, strtolower($b['token'])) !== false) {
$bot = $b;
break;
}
}
if ($bot === null) {
return;
}
$verification = 'ua_only_no_ranges_published';
if ($bot['ranges_published']) {
$verification = 'ua_only';
$ip = ($cfg['ip_key'] !== '' && isset($_SERVER[$cfg['ip_key']])) ? trim((string) $_SERVER[$cfg['ip_key']]) : '';
if ($ip !== '' && @inet_pton($ip) !== false) {
$raw = @file_get_contents(sourcetrack_crawler_cache_file($cfg, 'ranges'));
$ranges = is_string($raw) ? json_decode($raw, true) : null;
$cidrs = is_array($ranges) && isset($ranges[$bot['token']]) && is_array($ranges[$bot['token']]) ? $ranges[$bot['token']] : array();
if (count($cidrs) > 0) {
$verification = 'ip_mismatch';
foreach ($cidrs as $cidr) {
if (sourcetrack_crawler_ip_in_cidr($ip, (string) $cidr)) {
$verification = 'ip_verified';
break;
}
}
}
}
}
$uri = isset($_SERVER['REQUEST_URI']) ? (string) $_SERVER['REQUEST_URI'] : '/';
$path = explode('#', explode('?', $uri, 2)[0], 2)[0];
$path = substr($path === '' ? '/' : $path, 0, 512);
$status = function_exists('http_response_code') ? http_response_code() : 0;
$payload = array(
'bot_name' => $bot['name'],
'verification' => $verification,
'path' => $path,
'status_code' => is_int($status) ? $status : 0,
'collection_source' => 'edge_middleware',
'timestamp' => gmdate('Y-m-d\TH:i:s\Z'),
);
sourcetrack_crawler_http('POST', $cfg['api_base'] . '/api/server/crawler-hit', $cfg['api_key'], json_encode($payload));
} catch (\Throwable $e) {
// Fail-open: never affect the page.
}
}
register_shutdown_function('sourcetrack_crawler_report');
}
WordPress
Both options run on your WordPress server. Pages served from a page-cache plugin or a CDN never reach WordPress, so those crawler fetches are not seen.
- Must-use plugin (recommended): save the PHP snippet above as
wp-content/mu-plugins/sourcetrack-crawler.phpand adddefine('SOURCETRACK_API_KEY', 'st_live_...');towp-config.php. It reports the real status code. - SourceTrack AI Crawler Visibility plugin: a regular plugin with a settings screen (Settings, SourceTrack AI Crawlers) where you paste the token; the token is stored encrypted. It labels every hit user-agent only and does not report status codes. It is not in the WordPress plugin directory yet: email support@sourcetrack.ai for a copy.
Verification and testing
You can check your installation by sending a simulated crawler request to your application using curl:
# Test with a simulated GPTBot User-Agent
curl -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)" https://yourdomain.com/
The snippet recognises the crawler name, reports the fetch after your response is sent, and the hit appears on the AI Visibility page. A test like this is labelled user-agent only or IP mismatch, because it does not come from the crawler's published IP range.
Related documentation
Looking to connect AI assistants over MCP? See the AI assistants (MCP) guide: https://www.sourcetrack.ai/docs/mcp. For server-side conversion ingestion, see the Offline Conversions API reference: https://app.sourcetrack.ai/developers/offline-conversions.