OpenAI Web Search 首次进入 Artificial Analysis Search Index,得分 74,位列 26 款产品第 7、搜索提供商第 5,落后 Perplexity、Octen、Parallel 和 Brave。
OpenAI Web Search debuts on the Artificial Analysis Search Index at 74, the 5th best provider behind Perplexity, Octen, Parallel and Brave
@OpenAI Web Search is the first integrated first-party search tool on our board. Instead of an agent calling a Search API in our Stirrup harness, GPT-5.6 Luna (medium reasoning) calls OpenAI's built-in web_search tool, and OpenAI runs the whole search loop inside one Responses API call. All setups use the same underlying model, GPT-5.6 Luna (medium reasoning).
At ~$0.05 per task, OpenAI Web Search costs less than Parallel (advanced, $0.06) and Perplexity (medium, $0.07), but about twice as much as Octen ($0.024), which scores 3 points higher.
Key benchmarking results for OpenAI Web Search (medium search context):
➤ 5th best provider on the Artificial Analysis Search Index: OpenAI Web Search scores 74, a 41-point lift over the same model with no search (33). It places 7th of 26 products, level with You . com (highlights), Nimble (standard) and Exa (auto) at 74, behind the three Perplexity variants (77 to 80), Octen (77), Parallel (advanced) and Brave (LLM context) at 75
➤ OpenAI Web Search performs best on AA-Omniscience, our proprietary factual answer benchmark, achieving 72% accuracy. This ranks 3rd of 26 search provider variants, within 1 point of the leader, Firecrawl (73%). BrowseComp, which needs multi-hop reasoning across several searches, is its weakest benchmark: at 73.5% it ranks 13th of 26, well behind the Perplexity variants and Octen (85% to 87%)
➤ OpenAI Web Search uses fewer tokens for the same tasks: OpenAI bills about 40k input tokens per task, search results included, against 125k for the leanest Search API on the board, Perplexity (low). Its model cost is among the lowest on the board, at about $0.009 per task
➤ Built-in search costs less per task than most Search APIs: At ~$0.05 per task, OpenAI Web Search is cheaper than 17 of the 25 Search API products on the board (median $0.067), including Parallel (advanced, $0.06) and Perplexity (medium, $0.07). Octen costs about half as much ($0.024) and scores 3 points higher
Key details:
➤ Setup: GPT-5.6 Luna (medium reasoning) with OpenAI's web_search tool and search_context_size set to medium, in one Responses API call
➤ Pricing: $10 per 1k web_search calls, plus the search content, which OpenAI bills as model input tokens
➤ Contamination: we block known contamination sources through the tool's domain filter
来源:Artificial Analysis · x.com