Perplexity Retires Sonar Chat Completions: Web Search Must Now Be Explicitly Triggered—and Citation Data Moves Inside the Agent Output

Perplexity Retires Sonar Chat Completions: Web Search Must Now Be Explicitly Triggered—and Citation Data Moves Inside the Agent Output
Sponsored

Perplexity’s Sonar Chat Completions endpoint reaches its retirement date on September 27, 2026, forcing integrations built around POST /chat/completions to move to the company’s Agent API. The migration is more than an endpoint rename. It changes when web retrieval happens, where source data appears in the response and how strictly Perplexity validates requests.

Perplexity announced the transition in August, giving developers 45 days to migrate. A Perplexity staff response in the company’s API forum subsequently confirmed that the Sonar chat-completions endpoint would retire on September 27 and that workloads need to move to the Agent API responses architecture. Perplexity’s own developer repository now describes the Agent API as its primary API and labels /chat/completions deprecated. The original Perplexity announcement and official migration guide are the canonical references for teams completing the change. citeturn0search6turn0search9turn0search4

For AI-search developers, the most consequential difference is architectural. Sonar made grounded web search feel like an intrinsic property of the chat-completions call. Agent API treats search as a capability in a broader agent runtime. Developers now have to reason explicitly about whether the request should search, which tools it may use and where the resulting evidence appears.

Sonar’s chat contract gives way to an agent contract

Perplexity’s official migration skill maps raw HTTP integrations from POST https://api.perplexity.ai/chat/completions to POST https://api.perplexity.ai/v1/agent. OpenAI-compatible SDK integrations instead point their base URL at https://api.perplexity.ai/v1 and use the Responses-style interface, which resolves through /v1/responses. The Perplexity SDK likewise moves from chat.completions.create() to responses.create(). citeturn0search0

That difference reflects the larger product change. Perplexity says the Agent API retains grounded web search while adding multi-step research, code execution, built-in tools and access to multiple models through a single API. The company’s API cookbook describes it as the primary platform for hosted agents with native web access, citations, code execution, subagents and durable long-running work. citeturn0search4turn0search6

Developers migrating from Sonar therefore should not assume that preserving the same prompt preserves the same execution behavior. The request now enters a system designed to orchestrate capabilities rather than a search-first chat endpoint.

Web search is no longer automatic

This is the migration detail most likely to affect AI-search behavior. Perplexity’s official migration skill states that a bare Agent API model request does not automatically search the web. Developers must add the web_search tool or use a preset that bundles search. Merely making the tool available also does not guarantee that the model will call it. citeturn0search0

For applications where current web evidence is optional, letting the agent decide can save cost and latency. For citation-critical workflows, Perplexity recommends forcing web search with an appropriate tool_choice or using a preset. Even then, its migration guidance warns that an occasional run can return no sources, so integrations need to handle an empty search result cleanly. citeturn0search0

This is a meaningful change in the visibility pipeline. Under the new architecture, a page cannot become evidence simply because the developer called an API historically associated with search. Search itself first has to be invoked.

Retrieval now has an explicit activation gate

AI-search visibility already involves several gates. A system must be able to access a source, retrieve it for the query, select it as useful evidence and potentially expose it as a citation. Agent API inserts another developer-controlled decision before those stages: whether a web search should occur at all.

That means two requests with similar natural-language prompts can have fundamentally different source behavior depending on the tools and presets supplied by the application. One can answer from model knowledge without web grounding; another can trigger retrieval and create a source-backed output.

For publishers measuring AI visibility, this is another reason not to treat a model or product name as a single discovery surface. Application configuration can determine whether the open web is even eligible to influence the response.

Citations have moved out of the top-level response

The response contract also changes. Perplexity’s migration documentation warns that Agent API responses do not expose the old top-level citations or search_results fields. Developers need to inspect the response’s output[] array and read the search_results output item. citeturn0search0

The documentation specifically cautions against relying on message annotations for source extraction because those annotations can be empty. An integration that successfully migrates the request but keeps parsing citations using the Sonar response shape can therefore appear to “lose” its sources even when web search actually ran. citeturn0search0

This is not merely a serialization detail for applications that display source attribution, calculate citation analytics or store evidence for audits. Citation extraction is part of the migration.

Agent output distinguishes generated text from evidence-producing events

The move into output[] is consistent with an agent system that can produce more than one kind of result. Instead of treating the entire response as one assistant message plus some metadata, the API can represent different output items generated during execution.

That makes source provenance more structurally explicit. Search results are artifacts of a tool operation inside the agent run rather than universal metadata attached to every answer.

For developers, the practical lesson is to model grounded output as a sequence of execution artifacts. Generated text is one artifact; retrieved sources are another. Code execution, function calls and other tools introduce still more.

Old Sonar parameters can now fail loudly

Perplexity identifies leftover Sonar fields as the number-one migration failure. The Agent API rejects unknown request fields with an HTTP 400 error instead of silently ignoring them. The company’s migration skill instructs developers to compare every request body against the supported parameter mapping and remove fields that no longer have a direct equivalent. citeturn0search0

This strict validation is useful once an integration is correct because it prevents unsupported configuration from masquerading as active behavior. During migration, however, it means a mechanical endpoint substitution is insufficient.

Search filters, model identifiers, token controls, function calling and response parsing all need to be checked against the Agent API contract rather than carried forward by assumption.

Status handling changes as well

Another migration trap sits outside search and citations. Perplexity says failed or cancelled Agent API runs can return HTTP 200 while carrying a response status of failed or cancelled and a populated error field. Applications therefore need to inspect the response status rather than treating an HTTP 200 as proof that the agent completed successfully. citeturn0search0

Streaming consumers also need to recognize multiple terminal events. Perplexity documents completed, failed, incomplete and cancelled outcomes, plus transport-level errors. Code written around a chat stream that waits only for a conventional completion event can hang or mishandle unsuccessful agent runs. citeturn0search0

The broader lesson is that a long-running agent is operationally different from a synchronous completion request. Its execution state becomes part of the API contract.

The migration can change citation analytics without changing publisher content

For SEO and GEO teams, the migration creates an important measurement caveat. A sudden drop in citations inside a Perplexity-powered application after September 27 does not automatically imply that publishers lost retrieval visibility.

The application may no longer be triggering web search. It may be searching successfully but failing to parse the new search_results output item. A preset or model choice may have changed. Or the agent may simply decide that a web tool is unnecessary when the integration leaves tool use optional.

Before interpreting a citation trend as a content-performance change, teams should verify the application’s retrieval configuration and response parser.

Source visibility now depends more visibly on orchestration

NetContentSEO has previously examined the first-stage retrieval gate in Perplexity’s Q2D-Web retrieval benchmark. The Agent API migration adds a layer above that gate. Before a retriever can decide whether a page enters the evidence pool, the agent or application must decide to invoke retrieval in the first place.

The distinction is useful when diagnosing visibility. “Not cited” can now describe several different failures: no search was initiated, the source was not retrieved, the retrieved source was not selected for the answer, or the source was present in the agent output but the application did not parse or display it.

Those are different technical problems and require different fixes.

Agent API broadens what developers can optimize

The tradeoff for the more explicit architecture is substantially more control. Perplexity’s Agent API combines web search with multi-step research, code execution, built-in tools and multiple model choices. The official cookbook also describes subagents and durable long-running workflows as part of the platform. citeturn0search4turn0search6

A developer can therefore design workflows in which search is one stage rather than the entire product. An agent can retrieve evidence, analyze it, run code, call other tools and continue working across multiple steps.

That flexibility makes the retrieval configuration more important, not less. Once search becomes one tool among several, developers have to decide when fresh external evidence is required and when an ungrounded model response is acceptable.

Model choice and search choice are becoming separate decisions

Sonar’s product identity tightly associated model access with web search. The Agent API makes those concerns more separable. Perplexity can offer multiple models through the same orchestration layer while web retrieval is controlled through tools or presets. citeturn0search0turn0search4

That architecture mirrors a broader shift in agent platforms: the model is the reasoning component, while search is a capability the runtime can invoke. From a visibility perspective, being supported by the platform does not necessarily mean every request reaches the web.

Developers building research or citation-sensitive products should therefore document both decisions: which model handled the task and whether web retrieval was required, optional or absent.

September 27 is a contract boundary, not just a deprecation date

Perplexity’s retirement of Sonar Chat Completions marks the end of a simpler integration assumption: send a chat request to a search-oriented model and receive grounded text with citation metadata in familiar top-level fields.

The Agent API asks developers to be more explicit. Search must be enabled through a tool or preset and can be forced when grounding is essential. Source data must be read from the search_results item inside output[]. Unsupported Sonar-era fields can produce HTTP 400 errors. Agent status and streaming terminals require new handling. citeturn0search0

For application developers, that means September 27 is a migration deadline. For AI-discovery teams, it is also a reminder that citation visibility is increasingly determined by orchestration logic that sits upstream of retrieval itself.

A publisher can make a page perfectly accessible and highly relevant, yet it cannot be retrieved by a Perplexity-powered application if that application never asks its agent to search. In the Agent API era, the first visibility decision may happen before the search engine sees the query at all.

0%