robots.txt
PulseAugur coverage of robots.txt — every cluster mentioning robots.txt across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
New bot directive file standard emerges beyond llms.txt
The success of Anna's Archive's llms.txt suggests a growing need for more nuanced bot directives than robots.txt offers. It's plausible that other organizations will adopt or create similar convention-based files to guide AI crawlers for specific purposes, potentially leading to a new de facto standard for AI-specific web access control.
Websites increasingly block AI crawlers via IP ranges, not just robots.txt
Evidence shows users are actively exploring and recommending blocking Google's AI search scans via IP ranges, rather than solely relying on robots.txt. This indicates a shift in strategy as websites become wary of AI crawlers' impact and the perceived inadequacy of robots.txt for controlling AI-specific access.
Google to deprecate robots.txt for AI crawlers due to complexity
Given the documented issues with Google's crawler documentation and the increasing complexity of AI content access needs, it's plausible Google may eventually move away from relying solely on robots.txt for its AI crawlers. They might introduce a more sophisticated, AI-specific directive system or API to manage access, especially as they shift to an AI-first search model.
-
Cloudflare enables opt-out of AI training while preserving search indexing
Cloudflare is introducing a new feature that allows website owners to differentiate between content indexing for search engines and content usage for AI training. Starting September 15, 2026, websites can opt out of AI …
-
Guide: Control AI Crawlers with robots.txt, Nginx, and WAF
This article provides a technical guide on how website administrators can control AI crawlers using robots.txt, Nginx configurations, and web application firewall (WAF) rules. It aims to help site owners manage which pa…
-
New `llms.txt` standard guides AI agents on website access
A new technical guide explains the purpose and implementation of `llms.txt`, a file designed to direct AI agents on how to access and interact with a website. Similar to `robots.txt` for search engines, `llms.txt` provi…
-
AI bots like GPTBot may miss Google-ranked content due to robots.txt
Web crawlers like GPTBot and ClaudeBot may not be able to access content that ranks in Google search results. This is because a website's robots.txt file can prevent these AI bots from indexing specific pages, even if t…
-
Robots.txt insufficient for blocking AI bots; server-level controls needed
While robots.txt can request AI bots to avoid crawling a website, it cannot physically prevent them. More robust protection can be achieved through server, CDN, or WAF limitations, though these require greater technical…
-
New method uses 'canary tokens' to identify AI web scrapers feeding LLMs
Researchers have developed a novel method to identify which large language models (LLMs) are trained on data scraped from specific websites. The technique involves hosting dynamic websites that serve unique "canary toke…
-
AI agents vulnerable to code execution via `llms.txt` supply-chain attacks
Researchers have demonstrated a significant vulnerability in AI agents used by Fortune 500 companies, allowing them to execute arbitrary code through a supply-chain attack. By manipulating the `llms.txt` guidance files,…
-
Author names eight AI crawlers in robots.txt
An author has identified and named eight specific AI crawlers within their robots.txt file. This action aims to provide transparency and control over which AI systems are permitted to access and process content from the…
-
Cloudflare launches Bot Preference Sync for AI bot management
Cloudflare has introduced Bot Preference Sync, a new feature designed to harmonize robots.txt directives with specific AI bot configurations. This tool allows website owners to manage how search engine crawlers, AI agen…
-
AI training data and copyright spark heated debate over robots.txt
A debate is intensifying over the use of copyrighted material for AI training, with a particular focus on the role of robots.txt files. There is a growing desire for AI developers to respect these files, which indicate …
-
New `llms.txt` convention aims to guide AI assistants and agents
A new convention called `llms.txt` is proposed as a way for websites to communicate directly with AI assistants and agents. This single Markdown file, placed at the site's root, would serve as a guide for AI models, sim…
-
New ShieldFont disrupts AI scrapers by altering web text for bots
A new font called ShieldFont has been developed to combat unauthorized AI data scraping. Created by designers Isaque Seneda and Gabriel Abrucio, ShieldFont subtly alters words on webpages through ligatures, making the t…
-
Apache server admin fights AI scrapers with mod_qos
A server administrator has implemented mod_qos on an Apache2 server to mitigate excessive requests from AI scrapers that were ignoring robots.txt. These scrapers were overwhelming the Git hosting service with inefficien…
-
AI crawlers struggle to access 27% of Show HN products
A recent scan of 383 products launched on Show HN revealed significant accessibility issues for AI crawlers. Approximately 27% of these products were unreadable to automated systems, often due to client-side rendering w…
-
Cloudflare questions robots.txt adherence amid bot dominance
Cloudflare's general counsel is questioning whether websites should adhere to the directives in their robots.txt files, especially as bots increasingly dominate internet traffic. This consideration arises as the interne…
-
Website SEO audit guide focuses on machine readability
A technical guide outlines 20 checks for website optimization, focusing on machine readability for search engine crawlers. The audit is divided into four groups: access, delivery, structure, and content, with specific i…
-
Claude shared links exposed via Google Search due to public sharing
Reports from July 2026 indicated that publicly shared Claude chats and Artifacts were appearing in Google Search results, not due to a security breach, but because users had made the content public. This highlights a co…
-
Anthropic's Claude chats exposed via Google and Bing search results
Conversations shared via Anthropic's Claude chatbot were inadvertently exposed through Google and Bing search results due to a missing "noindex" tag on shared pages. This allowed search engines to index links that users…
-
Robots.txt, sitemap.xml, and llms.txt: Understanding AI's new web curation file
Website owners in 2026 will need to manage three distinct plain-text files: robots.txt, sitemap.xml, and the new llms.txt. Robots.txt controls which web paths bots can access, sitemap.xml lists all indexable pages for s…
-
Pagecord blocks AI training bots by default, offers custom robots.txt
Pagecord has implemented a default setting to block AI training bots from accessing its content, while still allowing AI search engines. The platform also offers premium customers the ability to customize their robots.t…