Useful resources
Scanners, registries, indexes and standards-body pages that tell you something checkable about the agentic web. Each entry says what it does and which articles it covers.
Documentation
- Permit Cloudflare AI Crawl Control bot reference
Names the AI crawlers Cloudflare recognises, with the operator, user-agent token and stated purpose of each.
- Identify Cloudflare Web Bot Auth documentation
Documents how Cloudflare verifies HTTP message signatures from bots at the edge and how a bot operator registers keys.
- Read llmstxt.org
The llms.txt proposal's own site: the file format, its rationale and a directory of sites publishing one.
Indexes
- Discover ardvark
Crawls the web for ai-catalog.json documents, verifies them against the ARD spec and indexes the agents, MCP servers and skills they declare.
- Discover Neuronto ARD publisher list
Lists the domains it has found publishing an ARD manifest, each checked by fetching the manifest and testing endpoint reachability.
- Pay The Agent Almanac
Aggregates public numbers on agent services, payments and infrastructure across several commerce and agent protocols.
- Pay x402 Foundation members
Lists the member organisations of the x402 Foundation.
Registries
- Discover AGNTCY Directory specification
Specification and API reference for the AGNTCY Agent Directory, a distributed registry of agent records.
- Act MCP Registry
The Model Context Protocol project's own registry of published MCP servers, with its API and server list in the open.
Scanners
- Read Agent Ready
Scores a given site against the Vercel Agent Readability Spec and llmstxt.org and lists the fixes it found missing.
- Read Lumar Agentic Readiness Scanner
Checks which agent-facing technologies a site publishes — robots.txt, sitemaps, MCP, schema and Markdown twins — and whether each is implemented correctly.
Standards bodies
- Permit IETF aipref working group
The IETF working group page for AI preferences: charter, drafts and meeting materials.
- Identify IETF webbotauth working group
The IETF working group page for Web Bot Auth: charter, adopted drafts and meeting materials.
- Identify IETF wimse working group
The IETF working group page for workload identity in multi-service environments: charter, drafts and meeting materials.
Tools
- Permit CrawlerCheck
Tests whether a site's robots.txt allows named search and AI crawlers to fetch a given URL.
Trackers
- Act Agents Welcome — Agent Protocol Atlas
Groups agent-facing protocols by layer with links to each specification and to sample code.
- Act Chrome Platform Status: WebMCP
Tracks the WebMCP feature's implementation status in Chrome, including its origin trial and standards positions.