Benchmark results
- SimpleQA
- FreshQA
About this benchmark
This benchmark evaluates how accurately models answer factual questions using retrieved web information, based on SimpleQA and FreshQA. SimpleQA contains short, fact-seeking questions, while FreshQA includes questions about changing facts and false premises.Methodology
- Dataset: SimpleQA (4,326 questions) and FreshQA (595 questions: 495 TEST + 100 DEV).
- Answer generation: Answers are grounded in search results retrieved by each provider.
- Scoring: Accuracy (correct answers / total questions), reported separately for each dataset.
- Grading: OpenAI’s official SimpleQA grading prompt and FreshQA’s strict grading rules.
- Retrieval: Up to 10 search results per query. FreshQA results exclude the same five questions with disputed grading for all providers.
Read more
How We Evaluate Search Quality
Read the full methodology, evaluation results, and analysis on the Ateve blog.