An Agency Turned “Citability” Into a Standalone Service and KPI

An Agency Turned “Citability” Into a Standalone Service and KPI
Sponsored

A New Zealand digital agency has turned one of AI search’s most slippery concepts into something it can put on a proposal, monitor every month and sell as a recurring service: citability.

What IF Web, a creative studio in Christchurch, says AI referral traffic to its own website increased 300% over a 28-day period compared with the previous 28 days after it combined technical cleanup, prompt-led content and stronger entity consistency across external directories. Domain citations and brand coverage also increased across the AI-search prompts the agency was tracking.

Those results come from a September 3 OtterlyAI case study based on an interview with What IF Web co-founder Isaac Farrow. The traffic figure is eye-catching, but the more interesting part of the case may be the business model built around it. What IF Web explicitly refuses to promise AI traffic or leads. Instead, it sells the ability to become more citable across an agreed set of prompts.

That is a subtle but important shift. Traditional SEO agencies have spent years selling rankings, organic traffic and conversions. What IF Web is treating source selection inside AI answers as a separate measurable outcome—one that can improve before referral traffic becomes large enough to matter commercially.

The agency started with a measurement problem

What IF Web began as a website agency focused on strategy, design, development and ongoing support. As clients started asking about AI search, the team encountered a basic problem: it had no reliable way to answer whether a brand appeared in AI-generated responses.

Manual checks in ChatGPT or Claude could show what happened for one prompt at one moment. Analytics could show some visits arriving from AI products. Neither method provided a stable view of how often a brand was being mentioned or which domains AI engines were citing across a meaningful set of buyer questions.

Farrow told OtterlyAI that AI traffic had effectively been the only KPI available before the agency adopted dedicated monitoring, and that attribution made even that metric difficult to interpret.

The agency’s solution was to move the KPI one step earlier in the user journey.

What IF Web sells citations rather than promising clicks

Farrow describes the offer in unusually narrow terms: the agency does not sell traffic and does not sell leads. It sells citability across the prompt set developed with each client.

Domain citations per prompt are treated as the primary metric, with brand coverage close behind. AI referral traffic sits further down the funnel as a lagging indicator—the number clients naturally care about most, but also the one the agency says it will not guarantee.

This approach recognizes a practical limitation of AI search today. A brand can be visible inside an answer without receiving a click. The user may read the recommendation, remember the company and navigate directly later. The AI product may mention the brand without linking it. A citation can influence discovery even when web analytics cannot connect that influence cleanly to a session.

Citability gives the agency something observable to optimize before those downstream outcomes appear.

The prompt set becomes the equivalent of a ranking universe

A citability KPI only makes sense if everyone agrees on what is being measured. What IF Web therefore builds the tracked prompt set with the client rather than inventing it alone.

The reasoning is sensible: sales teams and marketing managers hear the language real buyers use in conversations. Those questions can be translated into prompts that represent the situations where the company wants to appear.

But this is also one of the model’s biggest limitations. Unlike keyword research, prompt research does not yet have the same mature ecosystem of public volume estimates and query reporting.

Farrow is explicit about that uncertainty, calling prompt research a “very, very educated guess.” The team triangulates Search Console data, client surveys and sales conversations, then treats the resulting prompt set as a working hypothesis rather than a permanent truth.

That caveat is essential. A brand can improve its citation rate dramatically across a badly chosen prompt set and still fail to influence the questions its customers actually ask.

Every engagement starts with a health check

What IF Web has productized the work into an initial review followed by either technical fixes or an ongoing retainer.

The team built a custom Claude skill using OtterlyAI resources, academic research and its own four-pillar assessment framework. The review scores the client’s existing content, tracking data and authority signals before recommendations are made.

Because What IF Web is also a development agency, one-time technical work can be handled directly. The case study mentions robots.txt optimization, sitemap coverage and structural cleanup as examples.

The recurring work—content, authority building, monitoring and further optimization—sits inside the retainer.

Technical SEO was necessary, but usually not the main problem

One of the more useful observations in the case study is that What IF Web does not claim every client has a broken technical foundation.

Most of its clients use Webflow, which already handles several infrastructure basics. Sitemaps are generated automatically, robots.txt can be configured for crawler access, and sites built by the agency generally have the expected technical controls in place.

Farrow says the recurring weakness is content rather than infrastructure.

The agency finds articles that may be genuinely useful but are not written in an answer-led format. Headings appear as statements rather than questions. FAQs are absent. Structured data is missing. Comparison information may be embedded in images or visually constructed with CSS rather than represented as semantic tables.

These observations reflect What IF Web’s working methodology, not official ranking documentation from AI providers. No evidence in the case study establishes that an FAQ or semantic table automatically earns citations. The narrower point is that explicit, machine-readable answers are easier to identify than information hidden inside ambiguous presentation.

Agent analytics changed how the agency thought about crawler access

Farrow also describes an early mistake familiar to the emerging GEO industry. What IF Web added an LLMs.txt file believing it could become an important discovery mechanism for AI systems.

According to the agency’s agent analytics, AI bots were not visiting that file. They were instead reaching familiar entry points such as the sitemap, robots.txt and homepage.

OtterlyAI has made a similar change in its own product philosophy. In a February update to its GEO Audit, the company said it removed its LLMs.txt checker after failing to observe meaningful AI-visibility impact and shifted its attention toward crawler access, static content and structured data.

That does not prove LLMs.txt can never be useful in any future architecture. It does illustrate why AI-search services need measurement rather than relying on whatever tactic happens to be popular in the market.

What IF Web tested the model on itself first

Before selling the service aggressively to clients, the agency used its own website as the longest-running case study.

OtterlyAI reports that AI traffic increased 300% across one 28-day window compared with the previous 28 days. Domain citations and brand coverage also increased across the prompts What IF Web monitored.

The case study identifies three broad levers behind the work: technical cleanup, authority and entity consistency, and content built directly around tracked prompts.

There is no controlled attribution between them. All three were being changed within the same program, so the data cannot tell us whether one lever caused most of the traffic increase or whether the result depended on their combination.

The source also does not provide the underlying session counts behind the 300% increase. A percentage change can look dramatic when the baseline is small, so the figure should be treated as directional evidence from the agency’s own site rather than a benchmark that other companies should expect to reproduce.

Entity consistency became an off-site optimization task

What IF Web’s authority work focused heavily on Clutch because the agency observed that the directory was frequently cited in its category.

The team optimized its profile and made sure company information there matched what appeared on its own website and other directories. It also asked important clients for reviews.

Farrow’s interpretation is notable: he believes profile and entity consistency moved the needle more than simply increasing review volume.

That is an agency observation, not a proven causal rule. But it points to a meaningful difference between conventional on-page optimization and AI recommendation visibility. If an answer engine is trying to determine whether several references describe the same company, conflicting names, service descriptions, locations or category labels can create unnecessary ambiguity.

Keeping factual business information consistent across authoritative external profiles is sensible independently of any AI-ranking theory.

The content strategy starts with prompts rather than a generic blog calendar

The third lever was content written directly against the questions What IF Web was monitoring.

Instead of producing articles because a broad keyword tool suggested an attractive search volume, the agency used its prompt set to identify the decisions and questions where it wanted to become a source.

This is a useful distinction between keyword targeting and answer targeting. A traditional keyword may describe a topic, while a conversational prompt can encode a buyer’s context, constraints and desired comparison.

For example, a prospective client may not simply search for “web design agency.” It may ask an AI assistant which New Zealand agency is suitable for a SaaS marketing website, what tradeoffs exist between Webflow and another stack, or which provider combines branding and development.

Those questions require content that does more than repeat a service keyword. They require explicit answers and evidence.

Location-specific prompts changed the site architecture

What IF Web also used its prompt monitoring to make structural decisions. The agency had moved top-level services away from keyword-led labels such as “Webflow development” and “web design” toward solution-led categories including marketing websites, AI search, branding, web applications and ongoing support.

It preserved keyword-oriented pages as sub-services so existing search visibility would not simply be discarded.

When OtterlyAI data suggested the strongest prompts were location-specific, the agency added targeted pages for web design in New Zealand and Christchurch. Farrow says agent analytics later showed AI crawlers landing on those pages.

That sequence is a good example of AI-search data informing ordinary information architecture rather than creating a separate “AI website.” The same pages can serve human buyers, conventional search and automated retrieval systems when their purpose is clear.

Citability is useful because AI traffic is a lagging and incomplete metric

The strongest argument for citability as a KPI is not that citations are the ultimate business outcome. They are not. A citation does not pay an invoice.

The argument is that referral traffic alone can understate what happens inside AI discovery. Users can see a brand without clicking. AI products can mention a company without producing a trackable referral. Attribution can be lost when users move between devices or return later through branded search or direct navigation.

By monitoring whether the brand and its domain appear across strategically important prompts, an agency can measure an earlier stage in the journey.

That resembles the role rankings have historically played in SEO. Ranking position is not revenue either, but it is an observable intermediate metric between optimization work and organic traffic.

Citability is What IF Web’s attempt to create an analogous metric for generative answers.

It is not yet a standardized KPI

The term sounds precise, but different tools and agencies can define citability differently. What IF Web focuses on domain citations per prompt and brand coverage across a client-specific prompt set. Another provider could use share of voice, citation frequency, answer position or a weighted score across engines.

Prompt selection can also change the result dramatically. So can country, language, model version and the date on which answers are sampled.

That means a citability score should always be accompanied by its measurement frame: which prompts, which platforms, which geography, which period and what counts as a citation.

Without those details, the KPI risks becoming another proprietary marketing score that cannot be reproduced or compared.

OtterlyAI itself warns that AI-search findings are observational

There is another reason to avoid overclaiming from the case study. What IF Web is an agency describing its own commercial program through a vendor whose product it uses.

OtterlyAI’s published research methodology acknowledges a broader limitation of the field: external researchers do not have access to the internal ranking or citation logic of AI platforms. What they can measure is the public output those systems produce.

That makes observed changes valuable, but it does not automatically reveal the mechanism that caused them.

In the What IF Web case, technical changes, directory work, reviews, content and site architecture evolved together. The 300% traffic increase cannot be scientifically assigned to entity consistency, prompt-led blogging or any other single intervention.

The agency is productizing uncertainty rather than pretending it does not exist

One of the more credible parts of What IF Web’s approach is its refusal to guarantee the metric clients most want.

Farrow says offering citability initially felt uncomfortable because there is no guaranteed tangible ROI attached to it. The agency became more confident as it accumulated monitoring data, research and experience on its own site.

That is a healthier framing than promising a predictable number of ChatGPT leads from a set number of GEO articles. AI answer systems change rapidly, prompts are difficult to size and referral attribution remains immature.

A service can still be measurable without pretending its downstream economics are fully solved.

The model could create a new specialization inside agencies

What IF Web currently handles AI search across a small group of generalists, but Farrow expects the work to fragment into specialist roles as the service grows.

He anticipates expertise around authority building, content and client relationships, with possible dedicated ownership for platforms such as YouTube and Reddit.

That prediction makes sense because AI citations can originate from a much wider source ecosystem than a brand’s own website. An effective program may require technical SEO, content strategy, digital PR, directory management, community participation, video and analytics.

In that sense, citability is not merely an on-page optimization score. It can become a way of coordinating several marketing disciplines around the question of whether a brand is represented in the information sources AI systems actually use.

The 300% result is less important than the KPI behind it

What IF Web’s reported traffic increase makes a compelling case-study headline, but it is not the most durable idea in the project. The number covers one 28-day comparison, comes without raw session counts and follows several simultaneous interventions.

The more consequential development is that the agency has created a sellable service around a new intermediate outcome.

Its clients agree on the prompts that matter. The agency measures domain citations and brand coverage across those prompts. Technical work makes content accessible, prompt-led publishing creates candidate answers, and external entity work tries to reinforce the brand across sources that AI engines already cite. Referral traffic is then watched as a downstream result rather than guaranteed as the deliverable.

That framework will need refinement as AI platforms expose better query and referral data. Farrow readily admits that today’s prompt research remains an educated guess.

But the business logic is already clear. SEO turned rankings into a measurable service long before every ranking could be tied perfectly to revenue. What IF Web is attempting the same move for AI search: make “being worth citing” measurable enough to optimize, report and sell.

0%