Are We There Yet? A Decision Framework for Replacing Term-Based Retrieval with Dense Retrieval Systems
Sebastian Hofstätter Affiliation: TU Wien email: , Nick Craswell Affiliation: Microsoft email: , Bhaskar Mitra Affiliation: Microsoft email: , Hamed Zamani Affiliation: University of Massachusetts Amherst email: and Allan Hanbury Affiliation: TU Wien email:
Abstract
Recently, several dense retrieval (DR) models have demonstrated competitive performance to term-based retrieval that are ubiquitous in search systems. In contrast to term-based matching, DR projects queries and documents into a dense vector space and retrieves results via (approximate) nearest neighbor search. Deploying a new system, such as DR, inevitably involves tradeoffs in aspects of its performance. Established retrieval systems running at scale are usually well understood in terms of effectiveness and costs, such as query latency, indexing throughput, or storage requirements. In this work, we propose a framework with a set of criteria that go beyond simple effectiveness measures to thoroughly compare two retrieval systems with the explicit goal of assessing the readiness of one system to replace the other. This includes careful tradeoff considerations between effectiveness and various cost factors. Furthermore, we describe guardrail criteria, since even a system that is better on average may have systematic failures on a minority of queries. The guardrails check for failures on certain query characteristics and novel failure types that are only possible in dense retrieval sy
原文 arXiv:2206.12993;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2206.12993v1