Research and Measurement Standards
Measure the decision, not just the prompt.
Traditional search often centered on the keyword. Early AI-search measurement often centers on the prompt.
The unit of measurement
Hendricks centers on the commercial intent context.
Intent Context
What a unit of measurement is made of.
- Customer Need
- Customer Profile
- Use Case
- Constraints
- Geography
- Decision Stage
- Commercial Value
- Intent Context
An intent context is a realistic customer situation: the need, who has it, their constraints, location, decision stage, and what the decision is worth.
Context Panels
Each panel answers a different question.
- Neutral baselineA defined question without substantial supplied customer context.Question answered: What happens under standardized conditions?
- Customer cohortRepresentative industry, demographic, use-case, budget, geographic, or business constraints.Question answered: Which customer profiles change the outcome?
- Decision journeyMulti-step conversations that become progressively more specific as the customer approaches a decision.Question answered: Does the brand survive as the decision narrows?
- Platform and time panelRepeated observations across relevant search environments and time periods.Question answered: How stable is the observed outcome?
- First-party human researchWith consent, real participants may compare controlled findings with live user experiences.Question answered: Optional
Hendricks observes four systems: Google AI Overviews, ChatGPT, Perplexity, and Gemini.
Outcome Classification
Classify each brand outcome as one or more of:
- Absent
- Referenced
- Cited
- Considered
- Compared
- Recommended
- Preferred
- Inaccurately represented
- Contradicted
- Uncertain
Define classifier rules and human-review thresholds.
Weighting
High-value customer decisions receive more weight than low-value informational questions.
A transparent weighting model can consider:
- Demand
- Commercial intent
- Expected customer value
- Strategic fit
- Eligibility
- Evidence confidence
No weighting model should be presented as universal.
Proof Without False Precision
We separate what is observed, inferred, measured, and proven.
Observed
Responses, citations, sources, rankings, and referrals.
Inferred
The likely relationship between evidence gaps and outcomes.
Measured
Leads, opportunities, pipeline, revenue, and branded demand.
Tested
Baselines, staggered rollouts, matched groups, and holdouts.
A legend of four evidence classes. Observed uses a solid line and a filled dot. Inferred uses a dashed line and a hollow dot. Measured uses a solid rule with a tick. Tested uses a double rule. Every diagram on the page uses these four marks.
Hendricks does not claim access to a model’s hidden reasoning. We study inputs, outputs, sources, interventions, and business outcomes, then state how much confidence the evidence supports.
Evidence Grades
Every conclusion carries the grade of evidence behind it.
| Grade | Evidence |
|---|---|
| A | Controlled experiment combined with first-party CRM or revenue data |
| B | Strong first-party exposure, behavior, and commercial time-series evidence |
| C | Repeated controlled context-panel observations and consistent source patterns |
| D | Directional API, synthetic, or isolated observation |
Metrics
Five measures, each defined before it is reported.
- Observed Consideration Rate
- The commercially weighted percentage of defined test contexts in which the brand is presented as a legitimate option.
- Observed Recommendation Rate
- The commercially weighted percentage of defined test contexts in which the brand is explicitly favored or recommended.
- Selection Stability
- The consistency of consideration or recommendation across reasonable variations in context, wording, platform, location, and time.
- Evidence Coverage
- How much clear, current, and corroborated evidence exists for claims needed to win priority decisions.
- Commercial Selection Gap
- The value-weighted difference between the client’s observed position and the relevant benchmark.
Methodology statement
Hendricks does not claim to reverse-engineer hidden model logic. We observe the information environment, test representative customer contexts, analyze sources and evidence, engineer the conditions a brand controls, and measure what changes.
Absence is not yet a diagnosis. A single answer screen is one observation under one set of conditions.
Reproducibility Requirements
What is stored for each run.
Store for each run where legally and technically permitted:
- Exact question
- Supplied context
- Platform
- Model or search experience
- Date and time
- Location
- Session type
- Response
- Cited sources
- Classifier output
- Confidence
- Human-review status
The published Hendricks self-run, 2026-08-19-110930, measured citation presence only. It is not a full Selection Intelligence baseline. Read the run on the Hendricks Selection Baseline.
Read the Hendricks Selection BaselineLimitations
What this methodology cannot do.
- Personal memory cannot be reproduced universally.
- Model and search behavior changes over time.
- APIs may not reproduce consumer interfaces exactly.
- Not every AI impression is observable.
- Citation does not prove influence.
- Correlation does not prove causation.
- Offline selection may not be attributable.