xAI has documented model-specific API rate limits across five spending tiers, making it easier for developers to plan the throughput of Grok-powered applications and AI agents. Its official Rate Limits documentation specifies two principal ceilings for language models: requests per second (RPS) and tokens per minute (TPM). For grok-4.7, published limits range from 150 RPS and 50 million TPM at Tier 0 to 500 RPS and 100 million TPM at Tier 4.
The documentation was highlighted in an October 5, 2026 editorial alert. It describes capacity limits rather than a newly announced model launch. Crucially, an RPS allowance is not a guarantee that the same number of simultaneous web searches can complete successfully: token consumption, request duration, tool behavior and other operational constraints matter.
Grok 4.7 API rate limits by tier
| Tier | Cumulative spend threshold | Grok 4.7 RPS | Grok 4.7 TPM |
|---|---|---|---|
| Tier 0 | $0 | 150 | 50,000,000 |
| Tier 1 | $50 | 172 | 53,000,000 |
| Tier 2 | $250 | 208 | 60,000,000 |
| Tier 3 | $1,000 | 312 | 74,000,000 |
| Tier 4 | $5,000 | 500 | 100,000,000 |
Source: xAI Rate Limits. The documentation also mentions an Enterprise option available on request. Published values are platform documentation, not a guarantee of a particular team's effective throughput; personalized limits can be viewed in the xAI Console.
How xAI determines a team's tier
xAI states that tier qualification is based on cumulative API spend since January 1, 2026. Qualifying spend comes from prepaid credit purchases or successfully paid invoices. Tiers unlock automatically at the stated thresholds and, according to the documentation, do not downgrade after qualification.
That is different from a recurring monthly subscription or a monthly spend requirement. Developers should confirm their actual team tier and model limits in the xAI Console before making capacity commitments.
RPS and TPM measure different bottlenecks
RPS limits how many API requests can be initiated per second. TPM constrains the token volume processed within a minute. A workload can hit either ceiling first.
For example, a large batch of brief requests may be limited by RPS, while fewer requests carrying long prompts, extensive tool context or large generated outputs may consume the TPM allowance first. Developers should monitor both dimensions, rather than assuming a high RPS value means unrestricted throughput.
The current documentation also explains that the per-second limit is derived from a per-minute request budget, preventing the full minute's allowance from being exhausted in a single burst. It is therefore important to pace requests instead of treating the advertised RPS as an invitation to send an unlimited instantaneous spike.
Cached and reasoning tokens still matter
According to the supplied documentation alert, cached tokens and reasoning tokens count toward the TPM calculation. This is important because lower billing costs for cached input should not be confused with an exemption from throughput accounting. Similarly, reasoning-intensive requests may consume more tokens than the visible answer alone suggests.
Teams should use the API's actual usage information and the current documentation to measure consumption, including any hidden or returned reasoning-related usage categories that apply to the chosen model and endpoint.
What happens when a limit is exceeded?
xAI documents 429 Too Many Requests responses when a rate limit is exceeded. A robust client should handle rate-limit responses explicitly, reduce request pressure and retry with appropriate backoff rather than hammering the endpoint continuously.
Production agent systems benefit from centralized request scheduling, bounded concurrency, queueing and usage telemetry. These safeguards are particularly useful when many users or automated tasks share one API team allocation.
Why this matters for agents using Web Search and X Search
Grok can be used with search tools for retrieving current information, and agents may chain multiple searches, analysis steps and tool calls in a single user task. A workflow that appears to require only one answer can therefore involve multiple operations and substantial token usage.
However, the language-model RPS figure must not be interpreted as an independently verified number of parallel Web Search or X Search tool calls. Search-tool availability, tool-specific restrictions, endpoint behavior and task duration can impose additional constraints. The documentation provides model-level API limits; it does not establish a universal maximum of 150 or 500 concurrent searches.
Illustrative capacity planning
Suppose an agent averages 100,000 TPM-accounted tokens per completed task, including relevant context and processing. A 50-million-TPM allowance corresponds mathematically to 500 such task-token budgets per minute before considering other limits. This is only a hypothetical token-budget calculation, not a measured completion rate or guaranteed capacity. The actual result depends on request scheduling, token distribution, tool latency and model behavior.
For reliable estimates, teams should benchmark representative tasks and track peak RPS, rolling TPM, latency, tool errors and successful completions.
Implications for SEO monitoring and AI visibility platforms
SEO and GEO applications often run recurring checks across many domains, queries and markets. When those checks depend on a model with web-search capabilities, operational design can influence how many observations are collected and how quickly a report is delivered.
Teams building AI visibility monitors should distinguish three layers:
- Model throughput: RPS and TPM allowances for the selected Grok model.
- Tool execution: availability and behavior of web-search, X-search or other integrated tools.
- Research quality: reproducibility, source verification, sampling strategy and completeness of the resulting observations.
More API capacity can support a larger workload, but it does not automatically improve citation accuracy, discoverability or the quality of AI search findings.
Practical checklist for developers
- Check the team's active tier and model-specific limits in the xAI Console.
- Implement request pacing and a shared concurrency budget.
- Track TPM consumption, including cached and reasoning-related usage as documented.
- Handle 429 responses with bounded retries and backoff.
- Benchmark complete agent workflows, not just isolated model calls.
- Record tool errors and incomplete research tasks separately from successful responses.
- Recheck official limits because API documentation and available models may change.
FAQ
What are Grok 4.7's Tier 0 rate limits?
The published limits are 150 requests per second and 50 million tokens per minute.
What are the Tier 4 limits?
The documentation lists 500 RPS and 100 million TPM for Grok 4.7.
How many spending tiers does xAI list?
Five standard tiers, from Tier 0 through Tier 4, plus an Enterprise option on request.
Do cached tokens count toward TPM?
The supplied official-documentation summary says cached and reasoning tokens are included in TPM accounting. Developers should verify usage accounting for their particular endpoint.
Does 500 RPS mean 500 simultaneous web searches?
No. RPS measures API request rate, not guaranteed concurrent search-tool execution or completed searches.
What error appears after exceeding a limit?
xAI documents HTTP 429 Too Many Requests.
Does this announcement change Grok's search rankings?
No. The documentation describes API operational limits, not changes to public search ranking or citation selection.
NetContentSEO analysis
For agent-based SEO research, scale depends on more than choosing a capable model. xAI's published RPS and TPM ceilings make one part of the capacity equation visible, while search-tool behavior, token budgets, concurrency and verification determine practical throughput. Teams should measure complete workflows under realistic conditions before promising large-scale parallel search coverage.
Primary sources: xAI Rate Limits; xAI Grok 4.7 Developer Guide.