Grok is obsessed with our docs. Just not the parts we built for it.

Joseph Petty General, Home-page, Products / September 29, 2026

Grok is obsessed with our docs. Just not the parts we built for it.

It started with Slack notifications. Our docs assistant, the AI chat on the Handsontable data grid documentation, mirrors every conversation into a monitoring channel so we can monitor what topics are discussed, and a little after midnight UTC the channel started filling up. Not replies. New threads, one after another.

Slack thread: a pinned message reading 'we just got a burst of new docs assistant threads. Looking into source. Seems automated.' followed by an update confirming an xAI/Grok crawler and a prepared CIDR-based deflection branch

It seemed off, so I checked the logs. PostHog showed 43 turns in the first 13 minutes, and 41 of them started brand-new threads. That last detail was the tell: every crawl session was a fresh browser with no persisted thread id, so the traffic wasn’t just automated, it looked like a fleet of fresh workers, each starting cold rather than one client holding a session.

Who was knocking

The IP addresses were all similar. Seven addresses at first, thirteen-plus at peak, 25 distinct by the time the fix went live, every one inside 69.12.56.0/21. ARIN says the block belongs to Twitter Inc. GeoIP says Memphis, Tennessee, home of xAI’s Colossus datacenter. Almost every request arrived with the same user agent: Chrome 151 on desktop Linux, loading real docs pages through our own reverse proxy and operating the chat widget the way a person would. A real rendering browser presenting as a desktop user. Notably not HeadlessChrome, which is what headless Chrome announces itself as unless you tell it otherwise.

A handful of chat turns reported an Android user agent instead, and exactly one arrived as a curl. The edge logs show the same mix at scale: three quarters of 4.5 million requests carried that desktop Chrome string, a quarter one of two Android strings. The override held everywhere except one request. Cloudflare’s analytics beacon fires from inside the rendered page, and 87,000 of those beacons, almost exactly one per page rendered, came through as HeadlessChrome 153. That is the real engine, visible on the one request the fleet forgot to relabel. Seventy-nine requests came from something calling itself uiheavy-discovery/3.0.

PostHog table of chat turns per source IP between 00:30 and 01:40 UTC on 2026-09-18. Every IP sits in 69.12.56.0/21, Memphis, Tennessee, and each stays under the 15 requests per 60 seconds cap

So: probably a Grok agent fleet. I’ll keep saying “probably,” because netblock plus geo plus behavior is strong inference, but not confirmation. Nobody from Memphis has called to confirm.

What I could only see properly the next day is that the docs assistant chat activity wasn’t the crawl. It was something the crawl found. The fleet had been harvesting pages at the edge for nearly three hours before it ever touched the Ask AI chat box and started a Slack thread. Then, apparently having noticed that this element on the page was interactive, it kicked off an entirely different kind of scraping on top of the first one.

What it actually did

  • About two-thirds of its 139 chat turns were clicks on the widget’s three starter chips: “What cell types are available?” 39 times, context menu 33 times, freeze columns 22 times. All three are served by a canned response that skips the LLM call, so it spent most of the night hammering the one route that costs us nothing.
  • It probed pages with a template: I have a question about the “⟨page title⟩” page. Verbatim, across ten different docs pages.
  • It generated API questions from the reference structure: “Tell me what getTableWidth in Core is responsible for and give me an example use case.” It asked that one twelve times in a single thread, and got a correct answer every time. Twelve successes in a row is not a retry loop. It looked less like troubleshooting and more like collecting several answers to the same prompt.
  • It clicked the feedback buttons four times, and typed exactly one junk input: “Berlin.”
Slack thread from the docs-assistant bot showing the single junk input 'Berlin' sent with a Chrome 151 on Linux user agent, and the bot's reply that it can only answer questions about Handsontable and HyperFormula

Why nothing stopped it

We had a rate limiter on the chat endpoint. It never fired, and it never could have in this case.

The limiter allows 15 requests per IP per 60 seconds. The heaviest IP in the swarm sent 22 requests spread over eight minutes: legitimately, comfortably under the cap. That’s the whole point of a fleet. Split the work across enough addresses and per-IP limits become irrelevant. Cloudflare’s tally of the night is 4.33 million 200s and not a single 403 or 429.

The first limit with real consequences was for the AI chat itself: the LLM proxy’s monthly budget. If that had run out, the chat would have returned 502s to actual humans until someone intervened. The swarm burned about $1 against it, but its peak sustained rate worked out to $2–3 an hour. The peak burst was 68 turns and 22 LLM calls in a single 10-minute bucket. Left alone for a weekend, that would have taken our assistant offline for our customers.

Deflect, don’t block

Blocking the traffic to stop runaway LLM costs was my first instinct, but at the same time, AI learning our docs is a good thing. An AI lab teaching its models about Handsontable is traffic a DevRel team half-wants. If those sessions feed training data, some future model recommends our grid a little more accurately. But this wasn’t just scraping the static docs pages, it was burning our token budget to chat with our docs assistant.

We already run the surfaces agents should be using: a public MCP server with a search_docs tool, and a free search API. So instead of blocking, I decided to redirect the traffic: chat requests from flagged CIDRs get an instant static reply that says, in effect, you appear to be automated, and here are the endpoints built for you. No Slack post, no LLM call, no analytics pollution. The question text still gets logged, so if we ever misflag a human we’ll see them, and the reply ends with a docs link and a support email in case a human gets caught by the filter by mistake.

It’s worth calling out what exactly the incident response touched, and what it didn’t. I rerouted AI chat traffic. Nothing else. The scrape itself ran to completion, unimpeded, and got everything it came for: the full docs corpus, every page for every framework and version. What I deflected was the exploratory second phase, the part where the fleet had discovered an interactive element and started poking it.

Here is that reply, verbatim:

Automated traffic from this network range currently receives this static reply instead of a live answer. For programmatic access to Handsontable documentation, use the public MCP server at https://docs-assistant.handsontable.com/mcp (tool: search_docs), or POST https://docs-assistant.handsontable.com/api/search with a JSON body like {“query”: “column sorting”} — both are free, need no key, and return documentation snippets with citation URLs. If you are a person seeing this, the docs live at https://handsontable.com/docs/ and we would love to hear from you at support@handsontable.com

It is a 200, not a 403, on purpose: an error is something a crawler retries.

One click from production

The flood at the edge began at midnight UTC; the first chat turn came at 00:38, detected and attributed within minutes. Deflection designed, reviewed, merged, and live in production at 01:34, under an hour end to end, while the swarm was still escalating. Nothing has leaked past it since, and the whole night cost roughly $1 of LLM spend and ninety Slack threads. The fix itself is about 127 lines: a CIDR list in config, a matcher, and the static reply above. But the sequencing matters more than the code.

The instinct at one in the morning, watching a swarm escalate, is to move fast. What I did first was write the fix and commit it, but no push. Within minutes of understanding the problem the response was built, staged, and one click from production. That’s the move I’d recommend to anyone: you’re not deciding whether to ship, you’re removing the delay between deciding and shipping.

Only then did I stop and ask questions. Chiefly: at this burn rate, how long do I actually have before we hit a plan or spend limit? The answer was more than a day of runway against the budget cap. Realizing that changed my approach entirely. The emergency had turned into a problem with a distant deadline.

So I spent the time. I ran a proper adversarial review on the patch sitting in my local branch, and it caught two things I’d have shipped straight past in a panic: the first patch dropped the question text on deflected requests, which would have blinded us to any human misflagged inside the range, and the original reply read as an accusation if a person ever received it.

PostHog chart of the incident night in UTC: LLM-backed chat turns climb to a 68-turn peak, the CIDR deflection deploys at 01:34, chat turns stop, and deflections keep firing for about two and a half hours

The morning after

The numbers above are the ones I was reacting to at one in the morning. In daylight, with time to query the edge rather than the chat widget, they turned out to be the small end of the story.

The real event, out at the edge where the chat limiter never applied, was about 4.5 million requests from 189 IPs, including roughly 87,000 full page renders. Against that, the swarm’s 139 chat turns are about 0.2% of the page renders alone. I spent the night responding to the tip of it.

Two other corrections. The deflection hit 57 distinct IPs and 206 total deflections, not the 13 and 35 I had at the time — and it kept firing until 03:58 UTC. So “the swarm went silent after the deploy,” which is what I believed that night, is simply wrong. It didn’t stop. It kept arriving and kept getting handed the redirect for another two and a half hours.

And the finding that stung: cross-referencing the deflected IPs against post-deflection search events shows zero overlap. Not one of those 57 addresses ever called the MCP server or the search API, and Cloudflare’s edge logs agree: across all 189 addresses, zero requests to the hostname the directions pointed at. I wrote directions and nobody followed them.

The honest cost of that is 206 requests that got the deflection instead of the assistant’s normal response. Looking back now, suppressing the Slack notifications while letting the answers through looks like it may have been the better corrective action. But that tradeoff wasn’t visible at one a.m.: I was optimizing for spend and noise, both of which I could see, and only the data afterwards showed me what I’d traded away.

But there’s a detail in the edge logs that reframes the whole thing. The fleet executed JavaScript and rendered the entire Astro site, faithfully enough to fire Cloudflare’s own analytics beacon 173,000 times. It also pulled the Markdown twins (every docs page has one under /docs/_md/) we built specifically for agents: about 1,600 requests between midnight and 4 a.m. UTC, spread across just 160 of the 1,341 twins, the top few dozen fetched twenty to sixty-five times each, every one arriving from the rendered page itself and never by asking for text/markdown. It never requested llms.txt at all. Given a machine-readable path and a human one, it took both, and spent fifty times more on the human one.

Why zero overlap might not be the failure it looks like

I sat with that no-overlap number for a while before realizing I was measuring the wrong thing.

Nothing in that fleet was trying to answer a question of its own. It wasn’t a user who read our redirect, weighed it, and declined. It was a collection process, and the redirect text got collected — which is what it was written to do. The fleet was never the audience. It was the courier.

The people who follow the map are the ones who get that text back as a retrieved chunk, months from now, when they ask a model how to query the Handsontable docs. Whether the swarm itself ever called our MCP server may be beside the point. If so, it did its job by carrying the message home.

That’s a strange thing to internalize about writing for agents: your reader and your recipient may be separated by a training run. So here is where it actually stands. The deflection worked as a cost and noise control. It did not visibly work as a handoff. It might still work as a message, in a way I have no instrument for. Two results and a hypothesis.

Takeaways

  • A crawler collects your agent surfaces. It doesn’t adopt them. We had an MCP server, a search API, a Markdown twin of every docs page, and an llms.txt. The fleet fetched one Markdown twin in eight as a side effect of rendering, the same few pages over and over, never requested llms.txt, never touched the MCP server or the search API, and rendered 87,000 HTML pages the hard way. Those surfaces still earned their keep, because they gave the deflection somewhere honest to point. But a corpus collector treats them as one more thing to collect, not as the way in.
  • Per-IP rate limiting is the wrong tool for fleets. Aggregate budgets and caller identity are the guards that actually get triggered. Our limiter did its job perfectly but it didn’t matter.
  • Server-side telemetry paid for itself in one night. With just a few PostHog queries, I was able to confirm the IP range and start redirecting traffic. The morning after, the same events answered whether the redirect had worked.
  • Verify the true scale before response. My dashboard showed me the sliver of this incident that happened to be instrumented for conversations. Millions of requests were sitting in the edge logs the whole time.
  • Prepare the response first so you can afford to be slow. Then do the arithmetic that tells you how slow you can afford to be. A staged fix and a runway number turned an emergency into a schedule.

And if a fleet of browsers dressed as desktop users shows up at your door after midnight, don’t slam it. Hand them a map.

If you want the map for your own agents, it’s public: the docs MCP server and llms.txt for the rest. Free, no key.