A publisher’s crawler policy can affect whether its pages enter a particular AI retrieval pipeline. Google’s documented integration of Parallel Web Search with Gemini Enterprise Agent Platform makes that relationship explicit: the choice concerns a named grounding provider and a named crawler, rather than AI visibility in the abstract.
In its Parallel grounding guide, Google states that the integration does not use pages for grounding when they have disallowed ShapBot. Access is available through a Google Cloud Marketplace subscription or a Parallel API key.
What the ShapBot rule establishes
Parallel’s crawler documentation describes ShapBot as discovering and indexing websites for its web APIs. It recommends allowing the crawler in robots.txt and permitting connections from its designated IP ranges to improve visibility in its search results.
For a publisher, the immediate implication is eligibility within this retrieval route. If the documented exclusion applies to a page, that page cannot serve as grounding evidence through this integration. If access is allowed, the page still has to be discovered, retrieved for a relevant question and used in an answer.
That sequence is why “allowing ShapBot guarantees citations” would be an unsupported conclusion. It is also why “blocking ShapBot removes a site from Gemini” is too broad. The Google statement concerns grounding with Parallel on this platform, not every information source available to every Gemini application.
The application can narrow the source pool too
Google documents include and exclude domain lists of up to 200 entries each, country-based geotargeting and basic or advanced search modes. Advanced mode seeks more thorough results at higher latency.
These controls add a second decision layer. A publisher may permit crawler access while an application’s source configuration still limits whether the domain can participate in a particular request. Visibility analysis therefore needs to account for both publisher policy and the retrieval settings of the workflow being tested.
Country targeting is another reason to avoid treating one observed answer as universal. An application serving a particular market may retrieve a different set of useful pages from an application answering the same broad question elsewhere. Tests should preserve the location setting and report it alongside the results.
Crawler access is a distribution decision
The publishing question is which routes to machine use support the organization’s goals. A public reference page, a subscription archive and an original dataset may have different purposes. Teams can review those assets deliberately instead of treating every crawler rule as a single sitewide choice.
The relevant outcome is not simply how many bots arrive. Useful evidence includes whether a page becomes available to the intended retrieval service, whether its claims are represented accurately and whether attribution helps users identify the original publisher. Traffic is a further outcome that should be measured separately.
This article does not establish a new robots standard or recommend an automatic access change. It identifies the documented relationship between an existing publisher control and a specific grounding integration.
How to evaluate the effect
A practical test would use a stable set of questions and known relevant pages. Record the application’s provider, domain filters, location and search mode. Inspect the evidence returned and the final answer, distinguishing a retrieved page from a visible citation and from a subsequent visit.
When a page does not appear, investigate the retrieval conditions before attributing the absence to model preference. A blocked crawler, a source exclusion, a poor match to the question and an answer that omits attribution are different explanations. Only some of them can be resolved by changing content.
A comparison should also retain unchanged pages and questions where possible. That helps prevent normal variation in retrieval from being mistaken for the effect of a crawler-policy change.
The documentation date is not a launch date
At verification, Google’s page displays “Last updated 2026-10-01 UTC.” It confirms a documented capability but does not establish its original release date.
For NetContentSEO readers, the useful development is the explicit connection between crawler access and eligibility for a particular grounding path. AI visibility depends partly on how applications obtain evidence. Publisher strategy can address that route directly, while remaining clear that eligibility, citation and referral traffic are separate results.