LLMOcheki launches Search Fire Rate across eight engines, measuring whether AI actually ran a web search
Itera, Inc. (head office: Shinjuku-ku, Tokyo; Representative Director: Takayuki Muto) has begun offering Search Fire Rate in LLMOcheki, its tool for measuring and improving LLMO. The feature measures the proportion of answers for which generative AI ran a web search.
It distinguishes the surfaces answering from learned knowledge from those answering on the basis of search results, which can inform how you allocate investment.
Background: the same “AI answer” comes about in two different ways
A generative AI’s answer is either assembled from learned knowledge alone, or assembled on the basis of a web search run at that moment. The first reflects information up to the point of training; the second reflects what is published now.
That difference determines what works. On a surface where search is run, publishing an article or being listed in a third-party article may be reflected in a relatively short time. On a surface answering from learned knowledge, what you publish may take a long time to be reflected, or may not be reflected at all.
Yet there has been no way to keep track of which is happening for which question. The same budget meets surfaces where it works and surfaces where it does not — and work has been planned without that distinction.
Feature 1: recording and aggregating answer by answer
In the daily brand measurement, whether AI ran a web search while generating an answer is recorded and aggregated answer by answer. These are not estimates or simulations but figures based on the actual responses captured at the time of measurement.
Feature 2: comparison across eight engines
Fire rates can be compared engine by engine across eight AI engines, including ChatGPT, Perplexity, Claude, Gemini, Grok and Copilot. Behaviour differs by engine, so this informs which surface to start with.
Feature 3: working with query fan-out analysis
It can be combined with analysis of the search queries AI actually issued. Where a search was run, you can see how AI broke the original question into sub-questions, and therefore which vocabulary to prepare information around.
Knowing how much search is happening and knowing what is being searched for are connected on the same screen.
On how it is measured
The figures are measured values based on the responses captured in the measurement we run. The same question is run multiple times and the values are processed statistically. We publish an LLMO measurement standard v1.0, and this metric is designed in line with it.
Whether AI runs a search varies with the content of the question, the environment in use, and specification changes on the provider’s side. What the feature shows is a tendency under our measurement conditions; it does not represent the fire rate in every situation of use.
Comment from our Representative Director
“Discussion of AI search work tends to assume every AI behaves the same way. In reality some engines search and then answer, and some answer from what they learned, and the move you should make differs. Unless you can first measure which it is, you keep spending budget on a surface where it does not work.”
LLMOcheki https://llmocheki.com/
Originally published at prtimes.jp