AI search engines are increasingly being used for tasks that look less like casual question answering and more like purchasing research: compare products, filter options by budget, check features and recommend the best choice.
That makes factual freshness a much bigger problem than an occasional awkward sentence.
In a hands-on comparison published September 9, Semrush tested nine AI search products across common queries, follow-up questions and source verification. The review covered Google AI Mode, ChatGPT, Perplexity, Microsoft Copilot, Brave Leo, Duck.ai, Claude, Consensus and DeepSeek.
The most consequential finding was not that one engine universally beat the others. It was that several could produce polished, useful product recommendations while still getting important commercial details wrong.
Google AI Mode sometimes combined a product’s lowest advertised price with features available only on more expensive plans. Brave Leo returned some outdated prices. DeepSeek produced stale or inaccurate pricing multiple times, including after follow-up questions.
For users, the test is a reminder that an AI-generated comparison table can look much more precise than the underlying evidence warrants.
Semrush used the same accounting-software problem where comparison made sense
Semrush author Zach Paruch says he personally tested each engine across multiple queries.
For a more direct comparison among the general-purpose tools, he used an accounting-software search with the same budget and feature requirements. He then asked follow-up questions, opened citations and checked important details against the original companies’ websites.
Where possible, the tests began on free versions while signed out, approximating what a new user could access without immediately paying or creating an account.
This makes the article more informative than a feature checklist, but it is important not to overstate the methodology.
It is an editorial hands-on review, not a controlled accuracy benchmark involving thousands of standardized prompts. Semrush does not publish error rates, confidence intervals or a randomized test corpus from which a universal winner could be calculated.
Google AI Mode mixed entry-level prices with higher-tier features
Google AI Mode performed well initially in the accounting-software task.
Paruch describes its answers as clear, concise and easy to scan. AI Mode was also able to work with multiple requirements and present options in a format useful for product discovery.
The problem emerged during verification.
After asking more detailed questions and checking the recommendations against software companies’ websites, Semrush found cases where AI Mode presented the lowest price for a product alongside features that were actually available only on more expensive plans.
This is a particularly important class of AI error because every individual component can appear plausible.
The product exists. The low price exists. The premium feature exists. The mistake is in combining those facts into a package the vendor does not actually sell at that price.
Composite errors can be harder to detect than obvious hallucinations
When an AI invents a nonexistent company or produces a nonsensical number, users have a chance of noticing something is wrong.
A cross-plan error is subtler.
A user searching for accounting software under a fixed monthly budget can see a familiar brand, a real price and a real feature, then reasonably conclude that the plan satisfies the requirement.
Only a visit to the vendor’s pricing page reveals that the price and capability belong to different tiers.
That means factual verification in AI shopping and software research needs to test relationships between facts, not merely whether each fact exists somewhere online.
Brave Leo’s checked sources leaned heavily on third parties
Brave Leo produced concise recommendations and required relatively little setup, but Semrush found its product information less reliable in this test.
When Paruch opened the sources behind the accounting-software recommendations, most were third-party review sites and software marketplaces. He says he did not see official vendor websites among the sources he inspected.
Checking the quoted prices against the software companies’ current pricing pages exposed a couple of outdated figures.
This does not prove that third-party sources are inherently unreliable or that Leo always avoids primary sources.
It illustrates a source-selection problem specific to rapidly changing commercial facts. A review site can be excellent for user experience, comparative analysis and product strengths while still contain an old subscription price.
For current pricing, the vendor’s own pricing page is normally the stronger primary source.
DeepSeek continued returning stale prices after follow-ups
DeepSeek showed a similar problem more persistently.
Semrush found that its shortlist was focused and relevant to the accounting-software requirements, but current pricing was inaccurate multiple times.
The author asked for current prices, followed up about specific plans and challenged unclear information. DeepSeek continued relying heavily on third-party websites instead of moving toward the software companies’ own pricing pages.
Turning on its deeper reasoning mode produced a more detailed comparison, but did not eliminate the stale information.
This distinction is useful.
More reasoning can improve how a system analyzes the evidence it has. It cannot guarantee that the evidence itself is current or that retrieval selected the authoritative source.
Reasoning depth and data freshness are separate problems
AI products increasingly offer modes that promise more thinking, deeper research or expanded reasoning.
Those modes can help with complex synthesis, but the Semrush test demonstrates why users should not equate additional reasoning with factual freshness.
A model can reason carefully from an obsolete price and still reach the wrong purchasing conclusion.
For time-sensitive queries, retrieval quality matters at least as much as reasoning quality.
The ideal workflow is not simply “think harder.” It is “find the current primary evidence, establish which facts belong together, then reason over those facts.”
Several other engines passed the details Semrush checked
The review did not find the same pricing problems everywhere.
Perplexity’s important pricing details matched the vendor sites in the examples Paruch verified. Microsoft Copilot’s checked pricing was accurate. The product pricing and feature details checked in Duck.ai were also accurate, and Claude’s checked pricing and product information matched the companies’ official websites.
Those results should not be converted into global accuracy scores.
Passing a handful of checks does not prove that an engine is always correct, just as finding two stale prices does not establish a platform-wide error rate.
The comparison is best read as a set of observed behaviors under one practical research workflow.
Perplexity stood out for source density
Semrush found Perplexity particularly strong when the task required visible supporting sources.
Its accounting-software answer provided enough information and citations that fewer follow-up questions were necessary than with many of the alternatives.
The platform also allows searches to be focused on particular source categories and, on paid tiers, provides access to different AI models and premium research sources.
For the pricing details Paruch checked, the information was accurate.
Source quantity alone does not guarantee correctness, but exposing a broader evidence set can make verification easier because the user has more routes back to the underlying material.
ChatGPT handled conversational refinement well
ChatGPT’s strongest behavior in the test was context retention across a longer research conversation.
Paruch began with a detailed accounting-software request, then asked about individual products, compared leading options and added new requirements.
He rarely needed to repeat the business type, budget or feature constraints.
Semrush also noted that the sources used across the conversation included both comparison pages and software companies’ own websites, making important details easier to inspect.
The trade-off was verbosity. Answers often contained more detail than the reviewer needed, increasing the amount of text required to locate a specific fact.
That is a usability issue rather than an accuracy finding, but it matters in research workflows where users may perform many comparisons in succession.
Copilot was more compelling inside Edge than as a standalone search tool
Microsoft Copilot gave Semrush some of the clearest answers in the comparison, and the pricing details checked in the accounting-software exercise were accurate.
Its distinctive advantage emerged when used inside Microsoft Edge.
Copilot can work with the webpage, PDF or other material already open in the browser and can use related browsing context to answer questions. That changes the workflow from “search the web for me” to “help me understand and compare what I am already researching.”
Semrush found that Copilot used fewer sources than ChatGPT and Perplexity, leaving less supporting material to inspect when validating recommendations.
Again, fewer sources do not automatically mean lower accuracy. They do reduce the visible evidence available to the user.
Duck.ai prioritized privacy and model choice but exposed fewer sources
DuckDuckGo’s Duck.ai returned concise answers and useful follow-up questions, helping the reviewer narrow the accounting-software requirements conversationally.
Its “2nd opinion” function also made it easy to switch models and compare another response.
The pricing and feature details Semrush checked were accurate.
The weakness was source coverage. Paruch found relatively few citations and less variety among the cited websites than in tools such as Perplexity.
Duck.ai’s differentiation is also partly architectural rather than purely search-quality based: it is designed around privacy, removing identifying information before sending prompts to model providers and allowing basic use without creating an account.
Claude produced a smaller shortlist with concentrated sourcing
Claude returned fewer accounting-software recommendations than most of the other general-purpose engines, but Semrush found the shortlist relevant and straightforward.
The product and pricing information checked against official websites was accurate.
The limitation was source diversity.
Although Claude cited numerous pages, many came from the same small set of software marketplaces and comparison sites. Semrush specifically notes repeated use of G2.
This is another reason source count can be misleading. Ten citations from two underlying publishers do not provide the same evidentiary diversity as ten citations distributed across primary documentation, specialist analysis and independent sources.
Consensus is solving a different search problem
Consensus does not fit neatly into a product-shopping accuracy comparison because it is designed for scientific research rather than the open web.
Its database focuses on peer-reviewed literature, and Semrush tested it with a research question about whether influencer marketing increases purchase intention rather than the accounting-software task used for general-purpose engines.
The platform surfaced individual papers and their findings and used its Consensus Meter to show how the research leaned rather than presenting a single answer as if every study agreed.
That specialization highlights an important point in the broader comparison: the “best AI search engine” depends heavily on the evidence universe required by the question.
A tool optimized for peer-reviewed studies should not be judged primarily on restaurant availability or software subscription prices.
Source quality depends on what fact the user needs
The Semrush test reinforces a principle that is easy to lose when AI interfaces compress research into one answer.
There is no universally best source type.
For a software subscription price, the company’s current pricing page is generally authoritative. For whether users find the software difficult to learn, independent reviews can add evidence the vendor cannot provide neutrally. For a scientific claim, peer-reviewed research may be the appropriate standard.
An AI engine therefore needs more than “good sources.” It needs sources appropriate to each claim.
Product-data errors can occur when that matching process fails — for example, when an old third-party pricing table is used instead of a current vendor page.
Follow-up questions are part of the test, not just a convenience
One advantage AI search has over conventional search is conversational refinement.
A user can begin with broad requirements, then add constraints, challenge an answer or ask the engine to compare only two finalists.
Semrush deliberately incorporated that behavior into its testing.
The results show why follow-ups are also a useful reliability probe.
If a user asks “is that the current price?” or “which plan includes this feature?”, a robust search system should ideally re-check authoritative information rather than merely restating its original answer with greater confidence.
DeepSeek’s repeated stale pricing after follow-up questions is notable precisely because the conversation provided opportunities to correct the initial information.
AI recommendations can fail at the point where precision matters most
Broad product discovery tolerates some fuzziness.
If an AI system suggests five credible accounting platforms instead of the perfect four, the user can still continue researching.
Budget constraints and plan features are different.
A business choosing software under $50 per month needs to know whether the required functions are actually included at that price. A recommendation that merges the cheapest plan with premium capabilities can move a product from “ineligible” to “recommended” incorrectly.
As AI assistants become purchasing intermediaries, these small factual relationships can directly affect decisions.
The test is useful precisely because it is not a leaderboard score
Semrush labels the article as a tested and reviewed selection, but it does not present a standardized numerical ranking for accuracy.
That restraint is appropriate.
Nine systems differ in model, retrieval architecture, source access, privacy design, paid features and intended use. Consensus searches scientific literature. Copilot has deep browser integration. Brave Leo emphasizes private browser-based assistance. Perplexity emphasizes research sources. ChatGPT emphasizes conversational research.
A single score would obscure those differences.
The article instead exposes practical failure modes that users can watch for regardless of which engine they prefer.
This is not evidence that one engine is always accurate and another always wrong
The comparison also has obvious limits.
It is based on one reviewer’s hands-on testing across multiple queries, with a shared accounting-software task used where appropriate. It does not disclose a large randomized query corpus or enough repeated trials to calculate statistically meaningful error rates.
AI outputs can also change between runs. Search indexes change, product prices change and providers update models and retrieval systems continuously.
An engine that returned an outdated price during this test may return the correct value tomorrow. An engine that passed every checked detail can still make an error on the next query.
The findings are examples of reliability behavior, not permanent platform grades.
Primary-source verification remains the safest final step
The practical lesson from Semrush’s nine-engine comparison is simple but consequential.
AI search can dramatically accelerate the first stages of research. It can build a shortlist, explain differences, remember constraints and collect sources faster than manually opening dozens of search results.
But the final commercial facts still deserve verification.
Before buying software, check the vendor’s current plan page. Before relying on a scientific conclusion, inspect the underlying research. Before acting on local availability, verify the business’s live information.
The smoother AI search becomes, the easier it is to forget that the generated answer is an interpretation layer sitting between the user and those sources.
Semrush’s test found that layer useful across all nine products in different ways. It also found that some of the most important details could still break inside it.
For product research, that means the best AI search workflow may not be the engine that produces the most confident comparison. It may be the one that makes it easiest to see where each consequential fact came from — and to check it before making the decision.