Search Impact Measurement
Prove what changed, and how much confidence the business should place in it.
A higher AI mention rate is not automatically a business result.
A citation is not revenue.
A branded search increase is not always caused by one campaign.
Hendricks builds an evidence system that connects market exposure, customer behavior, commercial outcomes, and controlled tests without pretending attribution is perfect.
Four Levels of Measurement
Each level answers a different question and carries different weight.
Exposure
What changed in the information environment?
- Search impressions
- Generative-AI visibility where measurable
- Citations
- Cited URLs
- Consideration
- Recommendation
- Rankings
- SERP coverage
- Brand mentions
Behavior
What changed in customer behavior?
- AI-assistant referrals
- Organic visits
- Branded search
- Direct visits
- Returning users
- Decision-content engagement
- Comparison-page use
- Form starts
- Appointment activity
Commercial outcomes
What changed in the business?
- Qualified leads
- Appointments
- Sales-accepted opportunities
- Pipeline
- Win rate
- Closed revenue
- Customer quality
- Partner-sourced revenue
Causal evidence
What evidence suggests the intervention contributed to the change?
- Baseline comparisons
- Staggered rollouts
- Matched demand clusters
- Geographic comparisons
- Segment holdouts
- Landing-page experiments
- Paid-search validation
- Interrupted time-series analysis
Evidence Grades
Every executive conclusion states its evidence grade.
| Grade | Standard |
|---|---|
| A | Controlled experiment combined with first-party CRM or revenue data |
| B | Strong first-party exposure, behavior, and commercial time-series evidence |
| C | Repeated controlled context-panel observations and consistent source patterns |
| D | Directional API, synthetic, or isolated observation |
What Search Impact Measurement produces.
- Measurement-readiness audit
- Event and conversion taxonomy
- Search and AI channel rules
- Search Console and analytics integration
- CRM and pipeline mapping
- BigQuery or equivalent data model
- Branded-demand tracking
- AI-referral analysis
- Impact dashboard
- Experiment plan
- Evidence-graded executive brief
- Impact Ledger
Impact Contract
What gets agreed before any work begins.
At the start of an engagement, Hendricks and the client agree on:
- Primary commercial outcome
- Leading indicators
- Baseline period
- Target customer or segment
- Data sources
- Known limitations
- Planned interventions
- Available controls or comparisons
What Hendricks does not promise.
What we do not promise
- We do not promise that every AI interaction can be traced to an individual buyer.
- We do not classify every direct visit as AI influenced.
- We do not claim causation from a simple before-and-after chart.
We combine direct measurement, leading indicators, customer-source information, commercial data, and controlled tests to create a more defensible body of evidence.
Measurement Questions
Six questions about what impact measurement can and cannot see.
Can GA4 identify all AI traffic?
No. Google Analytics 4 cannot identify all AI traffic, and Hendricks does not report it as though it can. An analytics tool can only classify a visit from what arrives with it, so a referral from an AI assistant is countable in GA4 only when a referrer arrives and is recognized as one.
Three limits stack on top of each other. Assistant referrals are attributed inconsistently across tools and across time. Some visits arrive with no referrer at all and are recorded as direct. And some AI influence never produces a click, because the buyer reads an answer, forms a preference, and arrives later through a branded search or a direct visit that carries no trace of the original exposure.
Hendricks bounds that gap rather than filling it with a guess. Exposure is measured where the answer itself can be observed. Hendricks observes four systems: Google AI Overviews, ChatGPT, Perplexity, and Gemini. AI-assistant referrals are then reported as a floor rather than a total. Branded search, direct visits, decision-content engagement, and self-reported source data are tracked as leading indicators, and every conclusion carries the evidence grade that states how much weight it can hold.
Google AI Mode and Microsoft Copilot are surfaces that exist in the same information environment, and they are named here for that reason alone. Hendricks does not measure, test, monitor, or report on Google AI Mode or Microsoft Copilot. No Hendricks deliverable should be read as covering any of them.
How should self-reported attribution be used?
Self-reported attribution is the answer a buyer gives when asked directly how they found the company, usually in a form field or in the first sales conversation. Hendricks treats it as a leading indicator and a tie-breaker. It is the only signal that can name an influence no analytics tool was able to see, and it is also the weakest signal in the model.
It fails in predictable ways. People misremember. They name the last thing they touched rather than the thing that changed their mind. They pick whichever option sits first in a list. They skip the field entirely. One self-reported answer proves nothing on its own.
Hendricks therefore reads it as a distribution rather than as a record. The question is asked as an open field rather than a fixed pick list, because a pick list teaches the answer. The response is stored on the CRM record rather than on the analytics session, so it survives a buying cycle longer than a session. And a shift in that distribution across a defined period is what gets reported, corroborating exposure and commercial evidence rather than standing in for either. On its own, self-reported attribution is directional evidence and grades accordingly.
How do branded search and direct traffic fit the model?
Branded search and direct traffic sit at the behavior level of the measurement model. Both record that a person already knew the brand before arriving, which is what exposure is supposed to produce. Hendricks reads them as evidence that demand moved, never as proof of which exposure moved it.
The honest counterpoint is that both are easy to over-read. A branded search increase is not always caused by one campaign. Public relations, paid media, an offline conversation, seasonality, a competitor going quiet, and an AI-mediated answer can all lift the same line on the same chart. Hendricks does not classify every direct visit as AI influenced.
Both measures become useful when they are read against three things: a defined baseline period, the exposure record for the same weeks, and a comparison group that did not receive the intervention. The causal evidence level supplies the comparison shapes, including baseline comparisons, staggered rollouts, matched demand clusters, geographic comparisons, and segment holdouts. A branded-search rise that begins when a specific intervention lands, in the segments that intervention targeted and not in the segments it skipped, is a far stronger claim than the same rise reported alone.
What if the sales cycle is long?
A long sales cycle changes what can be measured now. It does not change whether the work can be measured. When closed revenue lands well after the intervention that contributed to it, Hendricks measures the leading indicators in the current period and holds the revenue conclusion until the cycle actually closes.
The four measurement levels mature at different speeds. Exposure moves first and can be observed inside the current period. Behavior follows, in branded search, decision-content engagement, form starts, and appointment activity. Sales-accepted opportunities and pipeline come after that. Closed revenue, win rate, and customer quality arrive last, on the buyer’s schedule rather than the reporting schedule.
The Impact Contract exists to settle this before anyone is disappointed by it. The primary commercial outcome, the leading indicators, the baseline period, and the known limitations are agreed at the start of the engagement, so nobody is asked mid-engagement to accept a proxy that was never part of the plan. Each level is reported with its own evidence grade as that evidence becomes available, and a revenue conclusion is not published on the strength of a pipeline movement.
Can paid search validate demand or messaging?
Yes, within a defined scope. Paid search buys a fast, controlled read on demand and on message that organic and AI-mediated exposure cannot return on the same timeline. Hendricks uses it here as a validation instrument rather than as an acquisition channel, which is why paid-search validation sits in the causal evidence level rather than in the exposure level.
Paid search answers three questions quickly. Does demand exist at the volume the demand model estimated? Does the message earn a click against the alternatives a buyer is shown? Does the landing page convert the traffic that message brings? Each is a controlled test with a known cost and a short read time, and each either supports or contradicts an assumption the wider program depends on.
Paid search does not answer whether an AI assistant will place the brand in a consideration set. Consideration is settled by observation, which is the work of Selection Intelligence, not by media spend. A paid-search validation carries the scope it was run at and no more.
What is the difference between correlation and causation here?
Correlation is two measures moving together. Causation is evidence that one of them produced the other. On a Hendricks engagement the distinction is not academic. It decides which evidence grade a conclusion carries, and therefore what the business is entitled to do with that conclusion.
A before-and-after chart is correlation. A visibility increase and a pipeline increase in the same quarter is correlation. Hendricks does not claim causation from a simple before-and-after chart, because the same movement can be produced by seasonality, a pricing change, a new sales hire, a competitor leaving the market, or a campaign running in another channel at the same time.
Causal evidence requires a comparison that isolates the intervention. The forms Hendricks uses are stated openly: baseline comparisons, staggered rollouts, matched demand clusters, geographic comparisons, segment holdouts, landing-page experiments, paid-search validation, and interrupted time-series analysis. None of them removes doubt. Each of them narrows the set of explanations that survive.
The evidence grades carry the rest. Grade A is the only grade that requires a controlled experiment, and the Hendricks methodology publishes the full scale and what each grade permits. Correlation does not prove causation, and a graded conclusion states plainly which of the two is on offer.