Research Note 001 · Methodology

Designed to separate recall from fit.

The pilot uses fixed buyer situations, repeated runs, separate exposure modes, and a frozen review process so every published number can be traced back to evidence.

Study design

01

Build realistic buyer situations

Twenty-four versioned scenarios cover ecommerce growth, mobile-first products, subscription media, midmarket cross-channel programs, omnichannel retail, and enterprise orchestration. Each specifies company scale, channels, data environment, technical capacity, priority, and constraints.

02

Run two separate modes

Open mode names no vendors and measures unaided recommendation recall. Controlled mode exposes a balanced, scenario-appropriate candidate set using neutral product profiles and measures selection rate when eligible.

03

Repeat across three labs

Claude Sonnet 5, GPT-5.6 Terra, and Gemini 3.1 Pro Preview each evaluated every scenario in both modes three times. A seeded schedule interleaved models, modes, and scenarios; controlled candidate order rotated.

04

Validate and review every response

Responses were checked against a fixed JSON contract and normalized into canonical vendor identities. Every flagged item was adjudicated. Deterministic delimiter repairs were allowed only when the recovered response passed the complete contract.

05

Freeze before reporting

The reviewed dataset was hashed before analysis. Report tables use accepted observations only, retain separate mode denominators, and use controlled opportunity rates instead of incomparable raw win totals.

Review disposition

Source responses
432
Accepted for analysis
417
Excluded
15
Delimiter repairs accepted
56
Vendor identities normalized
5
Unresolved review items
0

Exclusion policy

No silent correction

Fifteen responses broke the predeclared contract: eight exceeded the task-interpretation limit, four exceeded the reason-code limit, two used the wrong response shape, and one omitted a valid confidence value. They remain excluded; content was not truncated or invented to make it pass.

Dataset SHA-256: f160d372918cc5964a3563bdbdb0db172d5e8d1ec07cb1d2eda4cdbde8592f0e

Interpretation rules

Keep modes separate

Open and controlled results answer different questions and are never combined into one rank.

Use the right denominator

Open results use all accepted open responses; controlled rates use only eligible exposures for that platform.

Describe the pilot honestly

Three repetitions support exploratory patterns, not definitive market rankings or narrow significance claims.

Do not infer product quality

The measured outcome is AI recommendation behavior under study conditions, not customer success or commercial performance.