Skip to content

Add scripts to create bilingual Static Embedding models from Firefox Translations translations models - #1436

Open
tliumozilla wants to merge 6 commits into
mozilla:mainfrom
tliumozilla:embedding_exporter
Open

tliumozilla wants to merge 6 commits into
mozilla:mainfrom
tliumozilla:embedding_exporter

Conversation

@tliumozilla

@tliumozilla tliumozilla commented Jun 17, 2026 •

Copy link
Copy Markdown

Ticket: https://mozilla-hub.atlassian.net/browse/AIAUG-39

How to use

python utils/export_static_embedding.py \ --model-url "https://firefox-settings-attachments.cdn.mozilla.net/main-workspace/translations-models-v2/3a84a714-bfc7-4e52-81a1-b08d5d9421d2.zst" \ --vocab-url "https://firefox-settings-attachments.cdn.mozilla.net/main-workspace/translations-models-v2/671f09a0-12a2-436e-94d8-2575a4594c2a.zst" \ --output-dir static_embedding_export \ --write-model2vec

How to load exported embeddings
pip install model2vec tokenizers

from model2vec import StaticModel
model = StaticModel.from_pretrained("static_embedding_export/model2vec_model")

@tliumozilla
tliumozilla requested a review from a team as a code owner June 17, 2026 17:12
Record fields are not fully documented here. We use field 4 as byte size and assign names
from the name table in order, which is sufficient for exporting Wemb.
"""
with model_bin.open("rb") as f:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This feels very brittle, and I don't see a way to verify that the data is correct?

Is it possible to use the marian sdk directly, either by calling it or extending the code in the translations repo? I extended the code in this test here https://github.com/mozilla/translations/pull/1431/changes#diff-e0f262b6be35a545460aca4b5a826cc2d6644ac1ee6a4e298ca086ea4beb9918 but there is likely a simpler way to do it.

@rolf-moz

Copy link
Copy Markdown

We should have some kind of integration test code that can be run that verifies the static embedding. One way would be looking at cosine distances of common similar tokens like king/queen etc.

@rolf-moz

Copy link
Copy Markdown

We should have some kind of integration test code that can be run that verifies the static embedding. One way would be looking at cosine distances of common similar tokens like king/queen etc.

You could also pull my branch (which we know works ok) and export a few per token embeddings, and compare them with what you have output here.

@evgenyrp

Copy link
Copy Markdown
Collaborator

There's a model registry that exposes paths to released models on GCS: https://storage.googleapis.com/moz-fx-translations-data--303e-prod-translations-data/db/models.json

Then you can just provide a language pair in the arguments and download the models from there.

@evgenyrp

Copy link
Copy Markdown
Collaborator

We discussed that you might be interested in the non-finetuned students or teachers. The paths to these models are stored in the SQLite file that you can download from gs://moz-fx-translations-data--303e-prod-translations-data/db/db.sqlite So, it would be an SQL query to access those. Let me know if you need help with understanding the structure.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants