The judge

Each engine response is read by a language model and classified. This is the step that turns prose into the mention rows the score is computed from.

  • Model: openai/gpt-4o-mini
  • temperature: 0
  • Timeout: 20 seconds
  • Concurrency: 4

What it produces

Per engine × query, a mention row carrying the brand's status in that response — whether it was named, whether it was recommended, its rank, which competitors were named alongside it, and the sentiment.

recommended is a stronger claim than named: it is the subset of mentions where the model actively put the brand forward, and it drives the recommendation rate diagnostic.

Unjudgeable responses are dropped

If the judge cannot classify a response, that observation is discarded, not recorded as invisible. A parsing failure is a failure to measure, and scoring it as absence would manufacture a penalty out of a bug.

The consequence is that dropped responses shrink observations_valid while observations_attempted stays put — so enough of them trip the coverage gate and the score is withheld. The failure surfaces as missing evidence rather than as a bad score.