When a crawler finds a shared link
On 17 September 2026 a Bing Webmaster Tools site scan of briefgate.dev flagged
two pages for a missing H1. Both were demo client portals on
p.briefgate.dev, each with a magic token in the query string
(p.briefgate.dev/<slug>?t=…). Bing reached them by following the redirect
from the public demo link on the marketing site — nobody had linked a real
client's portal from anywhere crawlable, but the demo one, by design, is.
Nothing leaked: the intakes behind those URLs are demo data, confirmed against the database. Nobody was locked out either, because the token isn't single-use — reasons below. But a portal URL has no business sitting in a search engine's crawl log at all, and the fix is worth writing up, because every tool that emails people an access link has the same problem.
Four kinds of automated opener
A link mailed to a human doesn't only get opened by that human:
- Search crawlers (Googlebot, Bingbot, site-scan tools). Increasingly
render JavaScript, but usually not on the first pass — a scan does a fast
synchronous
GETof raw HTML first, and only some pages get queued for a later, rendered pass. The first fetch alone logs the URL, token included. - Chat and link-preview fetchers — Slack, iMessage, WhatsApp, Discord,
Teams. A plain
GETbuilds a preview card the instant a link is pasted, whether anyone clicks it or not. No JavaScript runs. - Mail security scanners — Microsoft Defender Safe Links, Proofpoint,
Mimecast. Rewrite links in incoming mail and fetch the original at
delivery time, sometimes again at click time — the classic killer of
single-use tokens, since the scanner's
GETcan spend the link before the recipient opens the email. - Corporate web proxies and gateways — outbound filters that fetch a URL as part of scanning traffic, same shape as a mail scanner, different trigger.
Only search crawlers meaningfully execute JavaScript, and not on every
request. The other three are GET-only — no script runs. Any protection that
depends on JavaScript running does nothing for three out of four.
"It is unlisted" is not a control
A link nobody linked to feels private. It isn't: a preview bot fetches it the
moment it's pasted into a channel, a mail scanner fetches every link in every
message regardless of who can see it, and a page that loads its own images or
analytics sends the current URL as Referer on those requests. Anywhere a
fetch gets logged — a crawler's history, a proxy's access log, a
webmaster-tools dashboard — the fact that a link was never linked from an
indexed page doesn't stop it being crawled once, and once is enough to be
recorded forever. "Unlisted" describes how a human finds a link; none of the
four openers above find links that way.
What actually happened, mechanically
A portal URL is <host>/<slug>?t=<token>. The page is client-rendered
(Nuxt): the initial HTML is close to an empty shell, and the page's own
JavaScript reads the t query parameter on mount, exchanges it for a
session, then calls history.replaceState() to strip the token back out of
the address bar. Bing's fast first-pass fetch never runs that script — it
reads the raw, near-empty HTML, which is why the scan reported a missing H1
rather than a rendered page.
The exchange itself also isn't something a GET can trigger. GET /portal/:slug requires an existing session cookie; it isn't how a token gets
redeemed. Redemption happens at POST /portal/:slug/redeem, called by the
page's own script, never by a browser just following a link. A GET-only
fetcher — the common case for three of the four opener types — cannot spend
the token, mint a session, or trigger the client.viewed webhook, no matter
how often it requests the URL.
Even where a token is redeemed, it isn't consumed: redeemMagicToken
checks it against a stored hash, constant-time, and accepts it until it
expires on its own schedule (portal_link_ttl_days, 30 days by default,
counted from the intake's most recent send), not on first use. That's what
kept this incident from locking anyone out — a deliberate trade, not an
accident that happened to help.
Why noindex alone wasn't the answer
These pages already carried a noindex meta tag, and it wasn't enough,
because a crawler has to fetch the page to read it — the fetch, query string
included, has already happened and been logged by the time the crawler learns
it shouldn't index what it just read. The fix that addresses the fetch itself
is robots.txt on the portal host:
User-agent: *
Allow: /app/login
Disallow: /app.briefgate.dev and p.briefgate.dev share one build, so one file covers
both. The dashboard sits behind login, so only its login page is worth
indexing; everything else, portals above all, is disallowed. Allow: /app/login wins over Disallow: / by longest-match, keeping login crawlable
while every token-bearing path stays untouched. A compliant crawler now never
issues the GET that would have logged the URL — stronger than asking it not
to index what it already fetched. robots.txt only binds crawlers that honor
it, which is why it's one layer, not the whole answer.
The concrete measures
- Disallow the token-bearing host in
robots.txt— stops the fetch, not just the index entry. noindexas defense in depth, useful against crawlers that ignorerobots.txt, useless against the fetch itself.- Hash tokens at rest. Portal tokens, like API keys, are 256 bits of randomness stored as a SHA-256 digest; the plaintext exists only in the link and briefly in memory during redemption.
- Compare in constant time, burning the same cost on an unknown slug as a known one, so timing can't be used to guess a token or enumerate slugs.
- Extend the link's life on every reminder, cap it between sends. An
intake has one stable link, so there is nothing to rotate — but each
reminder pushes
portal_link_ttl_daysout from that send, and if the sends stop, so does the extension. An old copy sitting in an inbox or a crawler's log stays live only as long as the intake is still being chased, and goes dead once the last send falls out of the window. A session already in progress is a separate cookie and is unaffected either way. - Keep the session short-lived and scoped — a separate
httpOnlycookie limited to portal paths, good for 14 days, tied to one intake. - Exchange over
POST, neverGET— nothing that just fetches the URL can spend anything. - Rate-limit the exchange endpoint, per IP and per slug.
- Strip the token from the address bar right after use, so nothing the
page loads afterward carries it in a
Refererheader or browser history.
Why the token isn't single-use, on purpose
Consuming a token on first read looks like the obvious fix, but it creates a
worse failure: a mail scanner's GET — the opener most likely to fetch
before the human does — would burn the link before the recipient opens their
inbox, and they'd see "this link has expired" through no fault of their own.
BriefGate's token stays valid until it expires on its own schedule, an
expiry the next reminder simply pushes back rather than replacing the token
itself. Security comes from the token being unguessable, hashed at rest, and
rate-limited on the redeem path, not from single use — the same property
that makes the system resilient to a scanner's premature fetch is what made
a crawler's fetch harmless here.
If you mail people access links
- Put the token-bearing host under
robots.txt, not justnoindex. - Assume every link gets fetched by a preview bot and a mail scanner before the recipient.
- Make the state-changing exchange a script-triggered
POST, never a bareGET. - Hash tokens at rest and compare them in constant time.
- Cap the link's lifetime and extend it only from your own sends, rather than trusting one link to stay valid indefinitely.
- Choose deliberately between single-use and scanner-safe — replay and premature expiry trade off against each other.
- Strip the token from the address bar as soon as it's used.
See Security for BriefGate's handling of client data, Chase for reminders and the portal link's lifetime end to end, and Portal for what a client sees on the other end of the link.