Skip to main content

Benchmark results

SimpleQA accuracy: Ateve 95.19%, Exa 94.38%, Brave 87.31%, Tavily 84.40%

About this benchmark

This benchmark evaluates how accurately models answer factual questions using retrieved web information, based on SimpleQA and FreshQA. SimpleQA contains short, fact-seeking questions, while FreshQA includes questions about changing facts and false premises.

Methodology

  • Dataset: SimpleQA (4,326 questions) and FreshQA (595 questions: 495 TEST + 100 DEV).
  • Answer generation: Answers are grounded in search results retrieved by each provider.
  • Scoring: Accuracy (correct answers / total questions), reported separately for each dataset.
  • Grading: OpenAI’s official SimpleQA grading prompt and FreshQA’s strict grading rules.
  • Retrieval: Up to 10 search results per query. FreshQA results exclude the same five questions with disputed grading for all providers.

Read more

How We Evaluate Search Quality

Read the full methodology, evaluation results, and analysis on the Ateve blog.