When a crawler finds a shared link

Last updated: · By Radim Sekera

On 17 September 2026 a Bing Webmaster Tools site scan of briefgate.dev flagged two pages for a missing H1. Both were demo client portals on p.briefgate.dev, each with a magic token in the query string (p.briefgate.dev/<slug>?t=…). Bing reached them by following the redirect from the public demo link on the marketing site — nobody had linked a real client's portal from anywhere crawlable, but the demo one, by design, is.

Nothing leaked: the intakes behind those URLs are demo data, confirmed against the database. Nobody was locked out either, because the token isn't single-use — reasons below. But a portal URL has no business sitting in a search engine's crawl log at all, and the fix is worth writing up, because every tool that emails people an access link has the same problem.

Four kinds of automated opener

A link mailed to a human doesn't only get opened by that human:

Only search crawlers meaningfully execute JavaScript, and not on every request. The other three are GET-only — no script runs. Any protection that depends on JavaScript running does nothing for three out of four.

"It is unlisted" is not a control

A link nobody linked to feels private. It isn't: a preview bot fetches it the moment it's pasted into a channel, a mail scanner fetches every link in every message regardless of who can see it, and a page that loads its own images or analytics sends the current URL as Referer on those requests. Anywhere a fetch gets logged — a crawler's history, a proxy's access log, a webmaster-tools dashboard — the fact that a link was never linked from an indexed page doesn't stop it being crawled once, and once is enough to be recorded forever. "Unlisted" describes how a human finds a link; none of the four openers above find links that way.

What actually happened, mechanically

A portal URL is <host>/<slug>?t=<token>. The page is client-rendered (Nuxt): the initial HTML is close to an empty shell, and the page's own JavaScript reads the t query parameter on mount, exchanges it for a session, then calls history.replaceState() to strip the token back out of the address bar. Bing's fast first-pass fetch never runs that script — it reads the raw, near-empty HTML, which is why the scan reported a missing H1 rather than a rendered page.

The exchange itself also isn't something a GET can trigger. GET /portal/:slug requires an existing session cookie; it isn't how a token gets redeemed. Redemption happens at POST /portal/:slug/redeem, called by the page's own script, never by a browser just following a link. A GET-only fetcher — the common case for three of the four opener types — cannot spend the token, mint a session, or trigger the client.viewed webhook, no matter how often it requests the URL.

Even where a token is redeemed, it isn't consumed: redeemMagicToken checks it against a stored hash, constant-time, and accepts it until it expires on its own schedule (portal_link_ttl_days, 30 days by default, counted from the intake's most recent send), not on first use. That's what kept this incident from locking anyone out — a deliberate trade, not an accident that happened to help.

Why noindex alone wasn't the answer

These pages already carried a noindex meta tag, and it wasn't enough, because a crawler has to fetch the page to read it — the fetch, query string included, has already happened and been logged by the time the crawler learns it shouldn't index what it just read. The fix that addresses the fetch itself is robots.txt on the portal host:

User-agent: *
Allow: /app/login
Disallow: /

app.briefgate.dev and p.briefgate.dev share one build, so one file covers both. The dashboard sits behind login, so only its login page is worth indexing; everything else, portals above all, is disallowed. Allow: /app/login wins over Disallow: / by longest-match, keeping login crawlable while every token-bearing path stays untouched. A compliant crawler now never issues the GET that would have logged the URL — stronger than asking it not to index what it already fetched. robots.txt only binds crawlers that honor it, which is why it's one layer, not the whole answer.

The concrete measures

Why the token isn't single-use, on purpose

Consuming a token on first read looks like the obvious fix, but it creates a worse failure: a mail scanner's GET — the opener most likely to fetch before the human does — would burn the link before the recipient opens their inbox, and they'd see "this link has expired" through no fault of their own. BriefGate's token stays valid until it expires on its own schedule, an expiry the next reminder simply pushes back rather than replacing the token itself. Security comes from the token being unguessable, hashed at rest, and rate-limited on the redeem path, not from single use — the same property that makes the system resilient to a scanner's premature fetch is what made a crawler's fetch harmless here.

  1. Put the token-bearing host under robots.txt, not just noindex.
  2. Assume every link gets fetched by a preview bot and a mail scanner before the recipient.
  3. Make the state-changing exchange a script-triggered POST, never a bare GET.
  4. Hash tokens at rest and compare them in constant time.
  5. Cap the link's lifetime and extend it only from your own sends, rather than trusting one link to stay valid indefinitely.
  6. Choose deliberately between single-use and scanner-safe — replay and premature expiry trade off against each other.
  7. Strip the token from the address bar as soon as it's used.

See Security for BriefGate's handling of client data, Chase for reminders and the portal link's lifetime end to end, and Portal for what a client sees on the other end of the link.