Method
How the lab measures: a frozen panel, four engines through their APIs, seven runs per question and week, every answer stored with its citations, firms matched by a published alias list, frequencies with Wilson intervals, a weekly manual check against the consumer apps. Version 1.0, 2026-09-29.
In one paragraph
We ask the same questions every week, the same way, and count. A panel of questions is written once, hashed and frozen. Each question goes to four engines through their official APIs with web search enabled and an explicit user location, seven times per week. Every answer is stored raw with the URLs it cites. Firms are matched in the answer text against a published list of names and aliases; the list grows from the answers themselves. We publish the share of answers naming each firm with a 95 % interval, the domains cited, the stability of the set of named firms, and the cost of the week. Nothing is edited after the fact: a change in the matching rules re-runs the extraction over the stored answers and carries a new rules version.
Steps
- Panel. Questions in English, one per row, in classes (choosing a provider, a problem, a comparison). Written before any data, frozen by SHA-256 in a journal committed to the open repository; the YAML is revealed on the freeze date. A frozen panel is never edited — a mistake becomes a new class with its own start date.
- Engines. ChatGPT through the OpenAI Responses API with the
web_searchtool anduser_location; Gemini through the Google API with Google Search grounding (no location parameter exists); Perplexity through the Agent API with theweb_searchtool and location; Google AI Mode through DataForSEO with a city code. Models and modes are recorded on every call. - Runs. Seven runs per question, engine and location per ISO week, one per weekday at 06:00 UTC on a small server. A missed day is caught up the next day within the same week and flagged. Runs are never deleted; a failed call is stored with its error and cost and retried the next day.
- Storage. Every call keeps the prompt, engine, model, date, location, the raw JSON answer, the cited URLs in order, the cost and the latency.
- Extraction. Answer text and firm aliases are normalised the same way (case, accents, punctuation, legal forms such as Lda, S.L., Advogados); matching is whole-word inside a sentence; a firm counts once per answer. Persons and products credit their firm. Names and cited domains that match nothing go to a review queue; real firms found there are added to the list with a date, and the week is re-extracted under a new
rules_version. - Metrics. Mention frequency = share of answers naming the firm, per class × engine × location, with the Wilson 95 % interval. Source share = share of citations per domain, typed as firm / directory / forum / official / listicle / other by public rules. Stability = Jaccard overlap of the sets of firms named this week and last week, per engine. Cost per engine. Our own citation count, reported like everyone else’s.
- Manual check. Every week 5–10 key questions are typed by hand into the consumer apps (ChatGPT, Perplexity, Google AI Mode) and compared with the API answers. The difference is itself a reading, not an error to hide.
- Publication. Tables and CSV files per week at stable addresses; readings with a pre-registration (question, metric, threshold, reading date) published before the data; zeros included. Firms are named as the engines name them, without quality judgements. By Machines never appears inside a published table.
Limitations
- API is not the app. Consumer interfaces add memory, personalisation and different retrieval; published comparisons put the overlap between API and app answers at a fraction, not at one. The weekly manual check measures the gap; it does not remove it.
- Answers are not deterministic. The same question yields different lists on different runs; that is why we run seven times and publish intervals instead of a single list.
- Engines differ. Gemini grounding has no location parameter and returns redirect URLs (the domain comes from the source title); Google AI Mode is observed through a third-party SERP provider; Perplexity and OpenAI take a city. Engine columns are therefore not directly comparable with each other, only with themselves over time.
- The firm list is the engines’ list. A firm never named by any engine is absent, not measured as zero, until it is added from a registry control sample.
- Costs cap the panel. The weekly budget is capped; when a week runs over, the panel is cut in a fixed order (documented in the repository) and the cut is logged.
Versions
| Version | Date | Change |
|---|---|---|
| 1.0 | 2026-09-29 | First version: 40-question panel frozen 2026-09-28 (start date moved to 2026-09-29 before any data, prompts unchanged, re-frozen); four engines; seven runs; extraction rules ext-2. |
Reproduce
The panel runner, the extraction and the export are open source (MIT): github.com/OOuph/bymachines-lab. The repository holds the configuration (panels, locations, firm list with aliases, engine prices), the tests and the deployment scripts; the API keys and the raw-answer database are not in it. The CSV files on this site are the unmodified output of lab export.