Building Robotswise: Balancing AI Crawler Governance and Google SEO
Over the past two years, the fundamental contract between website owners and web crawlers has fractured.
For decades, the deal was straightforward: search engines like Google and Bing crawl your public pages, index them, and in return, send organic search traffic back to your domain. Today, a new wave of autonomous bots—OpenAI’s GPTBot, Anthropic’s ClaudeBot, ByteDance’s Bytespider, and Common Crawl’s CCBot—harvest terabytes of written work, documentation, and open-source tutorials to train foundational models.
For many publishers, independent builders, and creators, this feels like an unfair exchange: our work is ingested into proprietary weights, often with zero attribution and zero downstream traffic.
To address this, I built and launched Robotswise (AI Robots.txt Generator). Here is the engineering breakdown of how it works and the architectural decisions behind it.
1. The Core Dilemma: Nuance vs. Scorched Earth
When creators first discover AI crawlers consuming their bandwidth and training on their data, the knee-jerk reaction is often:
User-agent: *
Disallow: /
This “scorched earth” approach is disastrous for organic discovery. It abruptly de-indexes your domain from Google Search, Bing, and DuckDuckGo, cutting off the lifeblood of your project.
Even when attempting to configure a more selective robots.txt, non-trivial questions emerge:
- If I block
GPTBot, does that also break ChatGPT’s ability to cite my product when users queryOAI-SearchBotin search mode? - Is
Google-Extendedresponsible for Google search ranking? (No—it solely controls Gemini training and Grounding, whileGooglebothandles traditional search indexing). - How do I keep sensitive administrative endpoints (
/admin/,/private/) private without accidentally advertising their paths to malicious actors?
A manual text file was no longer sufficient. Site owners needed an intuitive policy compiler that separates AI Training Crawlers from AI Search Engines and Standard Web Indexers.
2. Architectural Principles
Zero-Knowledge, 100% Client-Side Processing
When crafting a robots.txt file, developers frequently input internal directory paths (/api/internal/, /staging/) and their canonical sitemap locations. Sending these paths to a remote server creates an unnecessary security liability.
In Robotswise, no data ever leaves the user’s browser. The state machine, token validation, and string interpolation run entirely within the client bundle.
Multi-tiered Policy Presets
To minimize cognitive fatigue, the tool provides three one-click baseline policies:
- Block All AI Crawlers: Blocks every known frontier model scraper (
GPTBot,ClaudeBot,Google-Extended,Meta-ExternalAgent,CCBot,Bytespider, etc.). - Allow AI Search, Block Training: Keeps search citation bots (
OAI-SearchBot,PerplexityBot) open so your brand appears in AI answers, while shutting the door on bulk training bots. - Open Access: Retains standard discovery protocols.
3. The Technical Stack
Frontend: Next.js + React 19 + Tailwind CSS v4
Component Architecture: Radix UI primitives + Lucide Icons
Edge Runtime: Vinext + Cloudflare Pages (Serverless static edge)
Protocol Extensions: WebMCP (In-browser Model Context Protocol)
In-Browser WebMCP Integration
One unique experimental feature I baked into Robotswise is WebMCP (Web Model Context Protocol). If a user visits the tool with an AI Agent or browser assistant that supports the protocol, Robotswise registers a client-side tool dynamically:
const context = (document as WebMcpDocument).modelContext;
if (context?.registerTool) {
context.registerTool({
name: 'get_robots_txt',
title: 'Generate robots.txt',
description: 'Return the robots.txt currently configured in this generator.',
inputSchema: { type: 'object', properties: {} },
annotations: { readOnlyHint: true, untrustedContentHint: false },
execute: () => ({ filename: 'robots.txt', content: robotTextRef.current }),
});
}
This allows agents navigating the web to interactively invoke the generator, retrieve the generated directives, and write them directly into a project repository without manual copy-pasting.
4. Key Takeaways & What’s Next
- Crawler tokens are a moving target: Bot identifiers change quickly as labs fork their scrapers. Maintaining an open, versioned list of tokens is essential.
robots.txtis an advisory signal, not access control: A critical reminder highlighted in the UI is thatrobots.txtcommunicates access preferences to cooperative bots. True confidential endpoints must always be defended by authentication, WAF rules, andnoindexheaders.
The tool is live, free, and bilingual (English and Chinese): 👉 Try it here: https://robot-generate.toolyard.cc/