What are Multi-Vector Embedding Models?
Multi-vector embedding models, also known as late-interaction models, preserve token-level matching information by keeping one vector per token, unlike regular embedding models that compress a whole text into one vector. This approach is useful for applications like semantic search and visual document retrieval. By using the MaxSim operator, these models can score queries against documents more accurately, but at the cost of a bigger index.


The Sentence Transformers library has just gotten a whole lot more interesting with the addition of multi-vector embedding models - also known as late-interaction models. So, how do these models work? Well, they keep one vector per token, rather than squishing a whole text into one vector. This way, they preserve token-level matching information that can easily get lost when you're using regular embedding models. And to score queries against documents, they use the MaxSim operator, which takes the highest similarity between each query token and any document token, and then sums those maxima across the query.
This approach is a total game-changer for applications like semantic search, where you need to match a query against a huge collection of documents. By using multi-vector embedding models, your search results can be way more accurate, since the model can pick up on the nuances of both the query and the documents. Plus, these models can even be used for visual document retrieval - where a text query is matched against page images directly, no OCR needed.
One of the cool things about multi-vector embedding models is that they can be used with the same API as dense, sparse, and reranker models, making it super easy to integrate them into your existing search stack. But, of course, there's a catch - the index can end up being larger, which can impact performance. To get around this, you can use techniques like token pooling and speeding up inference.
All in all, multi-vector embedding models are a powerful tool for applications like semantic search and visual document retrieval. By preserving token-level matching information and using the MaxSim operator, these models can provide way more accurate search results. And while there are some costs associated with using them, the right techniques can definitely help mitigate those costs - making them a fantastic addition to the Sentence Transformers library.
Source: Hugging Face
NO COMMENTS YET
Comments are open. Have a thought or a question? Share it below.