Integrate with Sentence Transformers via MultiVectorEncoder

#1
by tomaarsen HF Staff - opened

Hello @xiaoxiaoshadiao and Team!

Heads up, this PR was AI-generated and human-reviewed.

Pull Request overview

  • Integrate EVIE-8B with Sentence Transformers via MultiVectorEncoder
  • Add a usage example below the existing ColPali instructions

Details

This follows the Preview integration from https://huggingface.co/tencent/EVIE-Preview-4.5B/discussions/1. I've kept ColPali first and grouped its installation and inference steps under a dedicated heading, so each library has its own setup instructions.

The integration uses Transformer -> Dense -> Normalize -> MultiVectorMask, with the trained 4096-to-4096 projection weights and bias extracted from this checkpoint into 1_Dense/model.safetensors. Bidirectional attention is configured through sentence_bert_config.json. A separate retrieval chat template reproduces the ColPali input format while keeping the original chat template unchanged. The processor metadata resolves the standard Qwen3VLProcessor, so no ColPali imports or trust_remote_code are needed.

The existing 16,384-token visual budget is preserved. The README encodes images one at a time to reduce peak memory use and shows how to select the 1,024-token budget used by the evaluation scripts. Unlike EVIE-4.5B, this checkpoint has a fixed 4096-dimensional projection, so no Matryoshka truncation example is included.

This PR is primarily additive. The only changes to existing runtime configuration are processor_class and padding_side. The ColPali path explicitly loads ColQwen3_5Processor and hardcodes left padding in its constructor, so it does not rely on these metadata values. The existing ColPali inference code and model configuration are unchanged, and ColPali behavior should therefore remain the same.

Verified against this repository's ColPali implementation in bf16 with both the pinned Transformers 5.13.1 and Transformers 5.15.0. Within each version, token embeddings are identical and the maximum MaxSim difference is 0 when scoring those embeddings in float32. This avoids differences caused by the libraries' bf16 score accumulation. The example below was verified with the public image URLs and the default model dtype.

To try this integration before the PR is merged, use the PR revision below.

pip install -U "sentence-transformers[image]>=6.0.0"
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("tencent/EVIE-8B", revision="refs/pr/1")

queries = [
    "What is the variable represented on the y-axis of the graph?",
    "Total outlay is maximum in which year?",
]
documents = [
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc1.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc2.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc3.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc4.jpg",
]

query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents, batch_size=1)
print(query_embeddings[0].shape, document_embeddings[0].shape)
# torch.Size([23, 4096]) torch.Size([3161, 4096])

scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[14.3594,  6.1758,  4.5137,  3.2979],
#         [ 1.8079, 10.9180,  1.9668,  1.9541]])

Happy to adjust anything. Please let me know if you have any questions or feedback!

  • Tom Aarsen
tomaarsen changed pull request status to open
Tencent org

Clipboard_Screenshot_1788787515
已经更新

Perfect, thank you! I'll test later to make sure that it works as expected 🤗

tomaarsen changed pull request status to closed

Sign up or log in to comment