Skip to main content
Start with a Diagnostic

Research and Measurement Standards

Measure the decision, not just the prompt.

Traditional search often centered on the keyword. Early AI-search measurement often centers on the prompt.

The unit of measurement

Hendricks centers on the commercial intent context.

Intent Context

What a unit of measurement is made of.

  1. Customer Need
  2. Customer Profile
  3. Use Case
  4. Constraints
  5. Geography
  6. Decision Stage
  7. Commercial Value
  8. Intent Context
The intent context formula: seven terms, one unit of measurement.

An intent context is a realistic customer situation: the need, who has it, their constraints, location, decision stage, and what the decision is worth.

Context Panels

Each panel answers a different question.

  1. Neutral baselineA defined question without substantial supplied customer context.Question answered: What happens under standardized conditions?
  2. Customer cohortRepresentative industry, demographic, use-case, budget, geographic, or business constraints.Question answered: Which customer profiles change the outcome?
  3. Decision journeyMulti-step conversations that become progressively more specific as the customer approaches a decision.Question answered: Does the brand survive as the decision narrows?
  4. Platform and time panelRepeated observations across relevant search environments and time periods.Question answered: How stable is the observed outcome?
  5. First-party human researchWith consent, real participants may compare controlled findings with live user experiences.Question answered: Optional

Hendricks observes four systems: Google AI Overviews, ChatGPT, Perplexity, and Gemini.

Outcome Classification

Classify each brand outcome as one or more of:

  • Absent
  • Referenced
  • Cited
  • Considered
  • Compared
  • Recommended
  • Preferred
  • Inaccurately represented
  • Contradicted
  • Uncertain

Define classifier rules and human-review thresholds.

Weighting

High-value customer decisions receive more weight than low-value informational questions.

A transparent weighting model can consider:

  • Demand
  • Commercial intent
  • Expected customer value
  • Strategic fit
  • Eligibility
  • Evidence confidence

No weighting model should be presented as universal.

Proof Without False Precision

We separate what is observed, inferred, measured, and proven.

Observed

Responses, citations, sources, rankings, and referrals.

Inferred

The likely relationship between evidence gaps and outcomes.

Measured

Leads, opportunities, pipeline, revenue, and branded demand.

Tested

Baselines, staggered rollouts, matched groups, and holdouts.

A legend of four evidence classes. Observed uses a solid line and a filled dot. Inferred uses a dashed line and a hollow dot. Measured uses a solid rule with a tick. Tested uses a double rule. Every diagram on the page uses these four marks.

Hendricks does not claim access to a model’s hidden reasoning. We study inputs, outputs, sources, interventions, and business outcomes, then state how much confidence the evidence supports.

Evidence Grades

Every conclusion carries the grade of evidence behind it.

Hendricks evidence grades and the standard each one requires.
GradeEvidence
AControlled experiment combined with first-party CRM or revenue data
BStrong first-party exposure, behavior, and commercial time-series evidence
CRepeated controlled context-panel observations and consistent source patterns
DDirectional API, synthetic, or isolated observation

Metrics

Five measures, each defined before it is reported.

Observed Consideration Rate
The commercially weighted percentage of defined test contexts in which the brand is presented as a legitimate option.
Observed Recommendation Rate
The commercially weighted percentage of defined test contexts in which the brand is explicitly favored or recommended.
Selection Stability
The consistency of consideration or recommendation across reasonable variations in context, wording, platform, location, and time.
Evidence Coverage
How much clear, current, and corroborated evidence exists for claims needed to win priority decisions.
Commercial Selection Gap
The value-weighted difference between the client’s observed position and the relevant benchmark.

Methodology statement

Hendricks does not claim to reverse-engineer hidden model logic. We observe the information environment, test representative customer contexts, analyze sources and evidence, engineer the conditions a brand controls, and measure what changes.

Absence is not yet a diagnosis. A single answer screen is one observation under one set of conditions.

Reproducibility Requirements

What is stored for each run.

Store for each run where legally and technically permitted:

  1. Exact question
  2. Supplied context
  3. Platform
  4. Model or search experience
  5. Date and time
  6. Location
  7. Session type
  8. Response
  9. Cited sources
  10. Classifier output
  11. Confidence
  12. Human-review status

The published Hendricks self-run, 2026-08-19-110930, measured citation presence only. It is not a full Selection Intelligence baseline. Read the run on the Hendricks Selection Baseline.

Read the Hendricks Selection Baseline

Limitations

What this methodology cannot do.

  • Personal memory cannot be reproduced universally.
  • Model and search behavior changes over time.
  • APIs may not reproduce consumer interfaces exactly.
  • Not every AI impression is observable.
  • Citation does not prove influence.
  • Correlation does not prove causation.
  • Offline selection may not be attributable.

Establish a baseline before making claims.