Because AI answers aren't fixed. Ask the same question twice and you can get two different answers — the models are probabilistic. On top of that, your result depends on your location, your account history, which model version you're on, and whether the app is doing a live web search behind the scenes.

Our probes control for that as much as possible: same questions, same setup, repeated runs, tracked over time — so the comparison is apples-to-apples month to month. Your one-off manual test is a single roll of the dice; our report is the average of several, measured the same way each time.

Neither is "wrong." They're answering different questions — yours is "what did I see just now," ours is "what does this engine typically say, and is it trending my way."