We Tested the Leading LLMs on Local Search & Discovery. Here's Where They All Failed

Local Search & Discovery Performance

We tested 4 leading LLMs (Claude, GPT, Gemini, Perplexity) on 345 real-world local search prompts – finding restaurants, checking hours, planning routes, booking tables – each run with and without web search (2,415 evaluations). Every recommended place was verified against Google Search and Maps.

Get the report
Overall scores by provider and configuration

The headline

OpenAI leads (90.7/100), followed by Gemini (86.4), Claude (85.9), and Perplexity (80.4). But no single provider wins everything — rankings shift by task type, and even the best model fails badly 8% of the time.

The scariest finding

Without search, 1 in 5 places Claude recommends doesn't exist, is permanently closed, or is in the wrong location. Even with search, no provider reliably detects closed venues — all 7 configs confidently gave booking guidance to a shuttered Buenos Aires restaurant.

The surprise

Web search helps on factual lookups (+8 points) but hurts on transactional tasks — Claude and Gemini both lose 5+ points on booking prompts when search is enabled. Search returns facts about a place instead of guidance on how to act.

The gap nobody talks about:

Constraint fidelity — whether the recommendations actually match what you asked for — varies 16 points across providers. All models find real places; not all find the right places.

Overall scores, search enabled

Out of 100, every recommended place verified against Google Search and Maps. The full report breaks this down by task category, region and failure type.

ProviderModelScoreCompletenessAccuracyConstraint matchPlausibility
OpenAI gpt-5.2 90.7 100%92%85%91%
Gemini gemini-2.5-flash 86.4 91%92%77%76%
Claude claude-sonnet-4-5 85.9 99%89%76%79%
Perplexity sonar-pro 80.4 97%86%69%71%
Which LLM is best at local search?

With web search enabled, OpenAI (gpt-5.2) scored highest at 90.7/100, ahead of Gemini (86.4), Claude (85.9) and Perplexity (80.4). No provider wins every task type: rankings shift between finding places, checking details, planning routes and contacting a business.

How often do LLMs recommend places that do not exist or have closed?

Even the best configuration recommended a place that was fabricated, permanently closed or in the wrong neighbourhood about 8% of the time. Without web search, roughly 1 in 5 places Claude recommended had such a fatal flaw. No provider reliably detected that a venue had closed.

Does enabling web search make LLMs better at finding places?

It helps discovery (+10 to +21 points) but hurts transactional tasks such as booking a table: Claude and Gemini both lost more than 5 points on those prompts with search on. Search returns facts about a place instead of guidance on how to act.

Get the full report
Get the report
PDF 37-PAGE REPORT

We tested the leading LLMs on local search & discovery. Here's where they all failed.

Q1 26 report is ready to be downloaded!

OPEN PDF

Request a demo