Skip to content

Collecting data · Protocol v2.3, pilot ·

Who AI recommends you for

When someone asks an AI answer engine for software that fits a stated situation, do the brands it names fit, by their own published terms, and does that change when the engine can search?

Results

None yet. The protocol below was written before any data was collected; the figures go here when the analysis is done, whatever they show.

What the study asks

Four AI answer engines are asked, in plain words, for software recommendations in six B2B categories: CRM, email marketing, project management, help desk, website builders and HR software. Each question carries a buying situation. A small startup. A large enterprise. A team of five with a budget under $100 a month, billed monthly. SAML single sign-on. Data hosted in the UK or EU. A named integration.

Every brand an engine names is then checked against its own pricing and documentation pages. Does the brand's published material say it fits the situation it was recommended for? The codes are EXPLICIT (the page meets the exact rule), CONTRARY (the page documents that it does not fit), NOT_FOUND (the pages do not settle it), and a handful of cases in between.

The marketing question underneath it: if what a brand publishes about itself is associated with where an AI engine recommends it, then pricing and documentation pages are doing work in a channel that sits outside the usual attribution.

How it is measured

  • Engines. OpenAI's default tier with native web search, the same model with no tools as a matched memory-only control, Perplexity Sonar, and Gemini Flash with Google grounding. The model strings are frozen before main collection; a served model that differs is excluded and counted.
  • Design. Two prompt templates per category and situation, five runs per engine per wave, all jobs interleaved through time by a recorded seed so an outage truncates every condition evenly.
  • Estimands. For brands that appear in both a fitting and an unsettled context, the difference in how often they are recommended. Reported per engine before any pooled figure, with bootstrap intervals, and labelled by direction and size together: a result can be "positive, negligible".
  • Claims. Observational alignment only. The study does not test an intervention and does not say why an engine recommends what it does.

Where it stands

The pilot on CRM and help desk ran in early October 2026. The codebook is being frozen against the pilot's blinded review pairs before the main waves; the held-out validation set is drawn only after the coding rules stop changing. Pilot data are excluded from the main analysis.

Disclosure

I previously worked at a MiCA-authorised digital-asset firm. No brand in this study's categories is connected to me.