Skip to content
All posts

Via6 AI · October 6, 2026 · 5 min read

Best Embedding Model for Multilingual RAG: We Tested 3

We tested EmbeddingGemma 2, EmbeddingGemma 300M and BGE-M3 on 203 English and Arabic search queries. The bigger, newer model did not win.

Key takeaways

  • We tested three open embedding models on 203 real search queries in English and Arabic.
  • Google's small EmbeddingGemma 300M won overall (97% top-1 accuracy), but only after we gave it the task prompts it expects.
  • The newer, larger EmbeddingGemma 2 was slower and worse at English-to-Arabic search.
  • BGE-M3 was the best at Arabic-to-Arabic search.

If your business uses AI to answer questions from your own documents (a support bot, a booking assistant, a voice agent that looks up policies), the quality of that answer starts with one component: the embedding model. It turns your documents and your customers' questions into numbers so the system can find the right passage. Pick a weak one and the AI answers from the wrong page, however smart the chat model on top is.

When Google announced EmbeddingGemma 2, a larger multimodal successor to its small open embedding model, we wanted to know whether it was worth switching. Here is what we measured.

How we tested the embedding models

  • Corpus: 149 published articles, a mix of English and Arabic.
  • Queries: 203 questions, written by an LLM from the articles. They split into English-to-English (95), Arabic-to-Arabic (54) and English-to-Arabic (54, an English question whose answer is in an Arabic article).
  • Score: for each question, did the correct article rank first (R@1) or in the top five (R@5)? We also report MRR, which rewards ranking the right article higher.
  • Hardware: CPU only, input capped at 512 tokens for every model.

Multilingual embedding model results

ModelR@1MRRArabic→Arabic R@1English→Arabic R@1CPU time per doc
EmbeddingGemma 300M (with prompts)0.970.9830.940.930.8 s
BGE-M30.960.9780.960.931.9 s
EmbeddingGemma 2 (with prompts)0.950.9730.930.875.3 s

All three found the right article in the top five for every query (R@5 of 1.00), so the differences are about ranking the right answer first.

Lesson 1: use the prompts, or your test is unfair

Our first run scored EmbeddingGemma 300M at 0.95 R@1 and EmbeddingGemma 2 at 0.92. Both models define a "query" prompt and a "document" prompt, and the first run skipped them. With the prompts applied, the 300M model went to 0.97 and the larger one to 0.95. BGE-M3 needs no prompt, so it was never affected.

If you benchmark an embedding model without reading its model card, you can reach the wrong conclusion. Check whether it expects task prefixes, and use encode_query and encode_document in sentence-transformers rather than a plain encode.

Lesson 2: bigger and newer did not win

EmbeddingGemma 2 has 740M parameters against 300M for its predecessor. On our test it ranked third, took about seven times longer per document on CPU, and was weakest on English questions about Arabic documents (0.87 against 0.93). Its extra size goes into image, video and audio support, which a text-only search system doesn't use.

Lesson 3: Arabic favours BGE-M3

On Arabic-to-Arabic queries BGE-M3 beat both Gemma models (0.96 against 0.94 and 0.93). If most of your customers write in Arabic, that gap may matter more than the overall winner.

What this means for a service business

  1. Start with a small model. A 300M-parameter model matched or beat a model more than twice its size, and it runs on a modest server.
  2. Test on your own questions. Take 100 real customer questions, note which document answers each, and score the models yourself. An hour of this beats any leaderboard.
  3. Check every language you serve. A model that is excellent in English can drop in a cross-language case.
  4. Read the model card. Missing prompts can quietly cost you accuracy.

Limits of this test

  • The queries were written by an LLM from the articles themselves, which can make retrieval easier than real user questions.
  • With 203 queries and every model at or near 1.00 on R@5, a one-point gap in R@1 is roughly two queries. Treat the ranking of the top two as close to a tie.
  • CPU timings are indicative only. They were measured in separate runs and will differ on a GPU.
  • We tested three models, not every option available.

Bottom line

For a bilingual English and Arabic knowledge base, EmbeddingGemma 300M with the correct prompts is a strong, cheap default, and BGE-M3 is the pick if Arabic-only search is your priority. We would not switch to EmbeddingGemma 2 for text retrieval on this evidence.

At Via6 we build voice and chat agents that answer from a business's own documents, so retrieval quality is something we measure rather than assume.

Free AI audit

Want this working in your business?

30 minutes. We map your top 3 automation opportunities and show you the ROI before you commit to anything.

Book Free AI Audit