Should You Let AI Bots Read Your Website?
Understanding the difference between search discovery, AI training and assistants accessing content for their users.
For many businesses, allowing useful discovery while making a separate decision about AI training is a sensible starting point. The right policy depends on what your website publishes, how it earns its income and what visitors need to do. “Allow AI” and “block AI” are usually too broad to describe that policy.
An assistant checking your opening hours for a potential customer is doing a different job from a crawler collecting material for model training. A search service building an index is different again. All three might request the same page, but the business relationship behind each request can be very different.
Cloudflare’s 2026 announcements make these distinctions easier to manage. They also expose an important limit: telling a crawler what you prefer and technically preventing a request are separate actions. Your website needs a clear position on both.
What changed in 2026?
On 1 July 2026, Cloudflare introduced more granular controls for Search, Agent and Training traffic, including on its Free tier. The announcement moved beyond a single broad AI-bot setting towards choices based on what the automation does. [1]
On 21 August, it announced Bot Preference Sync, which reflects a site’s bot-policy choices in its public robots.txt response. Cloudflare said the feature would prepend managed instructions to an existing file and maintain the relevant bot list, reducing the need to keep two configurations aligned manually. [2]
A 15 September follow-up refined the distinction further. Under Disallow AI Training, Cloudflare publishes the relevant preference while allowing its designated Accountable mixed-use crawlers to continue serving search; other training crawlers are blocked. The stricter Training Block options also affect mixed-use crawlers, potentially including Googlebot, Bingbot and Applebot. [3]
This article reflects primary documentation checked on 6 October 2026. The September update matters when interpreting the earlier launch material. Before changing a live website, check the settings currently applied to that domain and the current documentation, rather than relying on an old screenshot or the wording of a previous toggle.
Search, training and a user’s assistant are different uses
Search discovery involves collecting or indexing information so a service can find relevant material later. Your business may want its services, products and helpful explanations represented. However, being available for discovery does not guarantee a citation, a prominent position or a visit.
Model training concerns using material to develop or update a model’s capabilities. This is different from retrieving a particular page to answer a question now. Agreeing to one use should not be assumed to express a considered business preference about the other.
User-requested access happens when an assistant fetches information while helping someone complete a task. A prospective customer might ask it to compare delivery policies or read a product specification. Blocking that request can affect the person waiting for the answer, even though they have not opened your website themselves.
These are useful categories, but individual services implement them differently. One operator may separate the uses into different bot names. Another may use the same crawler for several purposes and offer additional controls over downstream use. That is why the provider’s documentation matters as much as the label in a dashboard.
What the main providers actually document
| Provider | Documented distinction | Practical consequence |
|---|---|---|
| OpenAI | OAI-SearchBot supports search; GPTBot concerns potential model-training material; ChatGPT-User handles certain user-triggered requests. [4] | Search and training preferences can differ. OpenAI says robots.txt rules may not apply to user-initiated ChatGPT-User actions. [4] |
| Anthropic | ClaudeBot collects potential training material; Claude-SearchBot supports search; Claude-User retrieves content at a user’s request. [5] | Anthropic documents robots.txt controls for all three. Blocking its user or search bots may reduce those forms of visibility. [5] |
| Googlebot crawls for Search. Google-Extended is a robots.txt control token without a separate HTTP user-agent string. [6] | Google-Extended is not a name you should expect as a distinct crawler in request logs. Its documented choices do not affect ordinary Search inclusion or ranking. [6] | |
| Microsoft Bing | Bing documents NOARCHIVE and NOCACHE controls with different effects on AI answers and training. [16] | They are not interchangeable. Review the intended effect and any existing tags before making a change. |
OpenAI explicitly documents that a site can allow OAI-SearchBot while disallowing GPTBot. It also says an OAI-SearchBot opt-out prevents inclusion in ChatGPT search answers, although navigational links can remain. [4] A search opt-out therefore has a different consequence from a training opt-out.
Anthropic’s documentation describes its training opt-out in terms of future material. [5] More generally, a changed crawler policy should not be presented as evidence that previously collected copies have been removed or that a trained model has been altered. Those are separate questions for the operator and any applicable removal process.
Avoid maintaining a copied list of bot names without a review date. Services add agents, revise their documentation and change infrastructure. A managed list can reduce that maintenance burden, but someone should still own the policy and check that the settings continue to match it.
Google: AI search participation is also a separate choice
Google’s current Search Console help documents a Search generative AI control, rolled out worldwide by 31 August 2026. It covers AI Overviews, AI Mode and generative AI features in Discover. Excluding a site removes its links and content from those features; Google says that choice is not a ranking or inclusion signal for other parts of Search. [7]
The same help page says this control does not govern AI training. It points publishers to Google-Extended for that purpose. It also explains that a property may inherit a parent property’s choice. [7] Review the effective setting, rather than assuming the page you opened represents a wholly independent policy.
Google-Extended’s crawler documentation additionally covers certain uses of content for training and grounding in Gemini products. [6] Grounding means supplying information to help a model answer a request. It is another reason not to describe every use of website content by an AI system as training.
For a business owner, the useful question is which outcome you want to change. Preventing model training, excluding generative search features and removing a page from search altogether are different objectives. Ask your website team to name the objective before selecting a control.
Preferences and enforcement work at different layers
The Robots Exclusion Protocol, standardised in RFC 9309, defines instructions crawlers are requested to honour. It explicitly states that these rules are not access authorisation. [8] Publishing an instruction can guide a cooperating crawler; the text file itself does not stand between a request and your server’s response.
| Control | What it is for | Boundary to understand |
|---|---|---|
| robots.txt | Communicating crawl instructions for matching crawlers and paths. [8] | It is public and depends on the crawler following the protocol. |
| Content-use signals | Expressing preferences about uses such as search, AI input and training. [11] | Recognition and implementation must be checked for each operator. |
| Search inclusion controls | Managing indexing or participation in supported search features. [7][10] | Their scope is provider-specific; they are not access passwords. |
| Network or server rules | Allowing, blocking or limiting matching requests. | Effectiveness depends on identification, rule order and the traffic passing through that control. |
| Application permissions | Restricting protected content to authorised users. | Private data needs controls on every relevant delivery path, including downloads. |
Google warns that a URL blocked in robots.txt may still appear in search if discovered through other links. [9] Preventing a crawl is therefore different from preventing an index entry. This distinction is especially important when someone proposes a global robots.txt change to solve a visibility problem.
Where a public page should stay out of Google’s index, a supported noindex tag or response header may be appropriate. Google must be able to retrieve that instruction: blocking the page in robots.txt can prevent it from seeing the noindex rule. [10] For genuinely private information, use authentication and permissions instead of depending on indexing preferences.
What about Content Signals and robots.txt extensions?
Cloudflare documents Content Signals for search, ai-input and ai-train. Its definitions distinguish conventional search results from AI-generated summaries, and describe ai-input as supplying content for activities such as retrieval and grounding. [11] These names should not be treated as exact synonyms for every dashboard category or provider’s policy.
Read the definition as well as the value. A site owner who permits “search” may imagine a list of links, while another setting may describe collecting information for generated answers. Similar-looking labels can cover different behaviours. Record which specification or provider meaning your policy relies on.
A signal also needs an operator that understands it. Publishing a new field does not establish universal adoption, technical enforcement or agreement about its interpretation. Check the provider’s supported controls and verify the actual response your website serves after any automatic additions.
What Bot Preference Sync simplifies—and what remains your job
Bot Preference Sync helps align the public instructions with the category choices configured in Cloudflare. Its launch announcement describes preserving existing Disallow material while prepending managed rules and updating the relevant bot list. [2] That can reduce one source of configuration drift between the website and its traffic controls.
It does not choose the business policy for you. Decide whether the site benefits from search discovery, whether user-directed access is useful and how you want training handled. Then check that other firewall rules or application settings do not contradict the intended result.
Cloudflare’s September article uses Accountable for operators with qualifying capabilities or time-bound commitments. It specifically says Bing’s robots.txt training opt-out is targeted for early 2027; until then, Disallow AI Training does not automatically communicate that choice to Bing. [3] An operator’s promised capability and its implemented capability should be recorded separately.
Bing’s documented NOARCHIVE choice excludes content from its AI answers and future foundation-model training, while preserving ordinary search results. NOCACHE has a narrower effect, allowing limited material, and Bing says NOCACHE takes precedence when both are present. [16] That deserves a deliberate implementation review, not an assumption that any tag containing “no” has the same effect.
Treat the final deployed configuration as the thing to inspect. Cloudflare’s July and August announcements explain how the product evolved; the September update changes how earlier labels should be understood. Save the date, setting and intended outcome whenever you change the policy.
A bot name is not the same as a verified identity
A request can carry a familiar user-agent name without that name alone proving who sent it. Cloudflare’s verification documentation describes stronger identification methods, including published IP information, reverse DNS and cryptographic signatures, alongside behaviour requirements. [12] Use the appropriate verification method when making consequential access decisions.
This matters in both directions. A permissive rule based only on a claimed name can admit unwanted traffic. An overbroad block can affect a useful service. Ask how the system identifies the operator, what happens when identification is uncertain and where false positives are reviewed.
Also check the route to your content. Cloudflare explains that proxied DNS records send web traffic through its network, while DNS-only records point clients to the origin directly. [13] A control at one network layer can only act on the traffic that reaches it. Include relevant subdomains, media hosts and alternative public endpoints in the review.
Even a successful request does not reveal every subsequent use of the content. Logs can show that a request occurred and how the site responded. They cannot, by themselves, prove whether the material later entered a model-training dataset. Separate observed traffic from provider commitments and downstream transparency reports.
Choose a policy that fits the business
| Website model | A useful starting position | What to assess |
|---|---|---|
| Local service business | Support discovery of public services and useful assistant access; decide training separately. | Whether customers can find accurate information and complete enquiries. |
| Online shop | Make public product and policy information accessible through chosen channels. | Retrieval accuracy, unwanted scraping, load and the integrity of checkout permissions. |
| Publisher or paid research site | Separate discoverable previews from the content that funds the business. | Subscription value, advertising, referrals, summary preferences and licensing options. |
| Private portal or internal resource | Require proper access controls for restricted information. | Authorised users, document delivery and whether connected assistants should have any access. |
These are proposed starting points, not default settings for every website in each category. A small business may publish valuable original research. A publisher may want assistants to retrieve a useful public explainer. Work at the level of the content and the business purpose wherever the available controls support that distinction.
For an online shop, reading a public product page must remain separate from acting on an account or completing a purchase. A crawler allowlist is not customer authorisation. Keep transaction permissions within the application and define which actions require the customer’s identity or approval.
When a platform offers only a domain-wide switch, recognise the trade-off. A business policy can be more nuanced than a particular control allows. Document that limitation, consider whether provider-specific or path-specific controls are available, and avoid presenting a broad toggle as a precise content licence.
Content ownership needs a publishing policy too
Start with an inventory of what you publish: original articles, commissioned photography, supplier descriptions, embedded media, user contributions and licensed resources. Keep records of the permissions and agreements attached to those materials. The appropriate choices may differ across that inventory.
Then describe the business concern clearly. Is it unwanted model training, lengthy summaries that replace a visit, repeated extraction of a catalogue, server load or disclosure of private information? These concerns point to different responses. A single “protect our content” requirement is difficult to implement or verify.
Crawler settings should sit alongside your publishing and licensing decisions. They do not resolve an ownership dispute or establish what a third-party licence permits. If a commercial licence or legal claim depends on a particular use of content, examine that agreement and the applicable rules separately from the technical configuration.
A useful arrangement might publish an open summary while reserving a detailed report for subscribers, or make product specifications public while protecting customer-specific pricing. Design that distinction into the website itself. Relying on bots to recognise an unwritten commercial expectation is a weak foundation for the business.
A practical review for WordPress and Elementor websites
Begin with the response served at your public /robots.txt address. WordPress can generate robots.txt dynamically and exposes a filter through which its output can change. [14] The file visible on the internet may therefore differ from what someone expects after looking only in the hosting file manager.
Identify every layer involved: WordPress, SEO or security plugins, hosting rules, caching and any network service that modifies the response. For an Elementor site, changing page content or layout is a different operation from changing these crawler controls. Record which person or service owns each setting.
| Check | Evidence to keep | Question it answers |
|---|---|---|
| Public instructions | The served robots.txt response, date and relevant bot groups. | Is the intended preference actually being published? |
| Page and provider controls | Relevant meta tags, headers and effective platform settings. | Do indexing, summaries and training choices match the brief? |
| Traffic enforcement | Verified request identity, matching rule and observed response. | Is the intended traffic allowed or blocked for the expected reason? |
| Important user journeys | Results for public pages, enquiries, account access and checkout where relevant. | Did the change interfere with the website’s useful functions? |
| Ownership and rollback | Previous settings, change owner, review date and reversal steps. | Can the team explain, maintain or undo the decision? |
Keep the existing instructions unless there is a reason to change them. Review how bot-specific groups interact with broader rules, and preserve required sitemap information. Ask for a merged configuration that fits the actual site, rather than replacing everything with a copied “block AI” file.
A manual request using a changed user-agent can help inspect a response, but it is not proof that a real verified crawler will receive the same treatment. Combine suitable testing tools with actual request evidence and the provider’s reporting. Test the intended behaviour at the layer where the control operates.
Measure useful outcomes, not just blocked requests
Cloudflare’s AI Crawl Control analytics distinguish request activity, successful responses, bandwidth and referral information. Its documentation marks some referral metrics as available on paid plans. [15] Check the reporting included with your account before promising a particular dashboard or measurement.
A crawler request is not a customer visit. A successful fetch is not proof of an AI citation. A citation is not an enquiry or sale. Keep those stages separate so a busy bot chart does not become a misleading business-success chart.
Before changing access, record a baseline for the pages affected: search visibility, relevant referrals, completed enquiries or orders, server demand and reported problems. Make a defined change and keep a note of its date. Review the result over a period appropriate to the site’s traffic and the provider’s processing time.
Interpret the comparison cautiously. A fall in visits can coincide with seasonality, campaigns, site changes or a search update. A rise in blocked requests may simply show that a rule is matching more traffic. Use the evidence to investigate the effect of a change rather than assuming one metric establishes its cause.
What to watch in 2027
Announced direction: Cloudflare’s September update sets a goal of offering more control over how much content appears in AI summaries by early 2027. It is a provider goal, not an already-delivered capability. [3] Keep roadmap statements separate from controls your team can demonstrate today.
Our forecast: purpose-based policies will become a more normal part of website maintenance. Businesses will increasingly ask which services may discover, retrieve, summarise, train on or transact with their content. Useful tools will need to make those differences understandable without pretending that every operator implements them identically.
Our forecast: the evidence around those choices will matter more. Expect greater interest in verified bot identity, records of policy changes and reports that distinguish access from downstream use. The useful question will be whether a control can deliver the intended outcome for this website and this provider.
We cannot responsibly predict a universal best setting, a guaranteed increase in AI referrals or complete prevention of unauthorised reuse. Build a review process that can adapt as provider capabilities, your audience and the business model change.
Questions website owners are asking
Can I block AI training but still appear in search?
Often, yes, where the operator supports separate controls. The implementation differs: some providers use distinct bots, while others use a shared crawler with additional preferences. Confirm the relevant provider’s method and check that a broader firewall block does not defeat the search access you intended.
Will allowing AI bots improve my rankings or sales?
Access can make discovery or retrieval possible; it does not establish a ranking or sales benefit. Keep the content accurate and useful, measure relevant referrals and business outcomes, and evaluate each channel on the evidence your own website produces.
Does robots.txt protect private documents?
No. It publishes instructions to crawlers and is itself publicly accessible. Use authentication, authorisation and protected file delivery for private material. Do not treat a disallowed path as confidential simply because it appears in that file.
Why would a user-requested assistant ignore a robots.txt rule?
Provider policies differ from one another and from automatic crawling rules. OpenAI says robots.txt may not apply to user-initiated ChatGPT-User requests; Anthropic documents robots.txt controls for Claude-User. [4][5] Check the operator’s policy, then use technical access controls where access must actually be restricted.
Does disallowing training remove content already collected?
Do not assume so. A future-access preference is not evidence that old copies have been deleted or a model has been retrained. Ask the relevant provider about its removal procedures and what they cover; preserve records of the request and response.
Do I need Cloudflare to set a crawler policy?
No. You can publish supported crawler instructions and use controls offered by your hosting, application and search providers. Cloudflare is one way to combine managed preferences, traffic controls and reporting. Choose the arrangement your team can implement and maintain correctly.
Can I make bots pay to read my website?
Cloudflare documents Pay Per Crawl as a private beta within AI Crawl Control. [17] That is an access and participation model, not a guarantee of demand or income. Verify eligibility and terms, and do not budget on the assumption that every crawler will pay.
What should a small business do first?
Inventory the public and private content, check the current settings and identify which discovery channels matter. Write a short policy covering search, training and assistants. Have the implementation reviewed, preserve a rollback and set a date to examine the results.
Put the business decision before the bot setting
A useful bot policy should be understandable without opening a firewall dashboard. It should explain which content you want discovered, which uses you support, what you want restricted and who maintains the controls. Your website team can then translate that policy into supported settings and check the result.
For Addweb, the practical starting point is a conversation about the website’s purpose. A service business, online store and subscription publisher may all use WordPress, but they can have very different reasons for allowing a crawler to read a page.