Add scripts to create bilingual Static Embedding models from Firefox Translations translations models - #1436
tliumozilla wants to merge 6 commits into
Conversation
…Translations translations models ticket: https://mozilla-hub.atlassian.net/browse/AIAUG-39
| Record fields are not fully documented here. We use field 4 as byte size and assign names | ||
| from the name table in order, which is sufficient for exporting Wemb. | ||
| """ | ||
| with model_bin.open("rb") as f: |
There was a problem hiding this comment.
This feels very brittle, and I don't see a way to verify that the data is correct?
Is it possible to use the marian sdk directly, either by calling it or extending the code in the translations repo? I extended the code in this test here https://github.com/mozilla/translations/pull/1431/changes#diff-e0f262b6be35a545460aca4b5a826cc2d6644ac1ee6a4e298ca086ea4beb9918 but there is likely a simpler way to do it.
|
We should have some kind of integration test code that can be run that verifies the static embedding. One way would be looking at cosine distances of common similar tokens like king/queen etc. |
You could also pull my branch (which we know works ok) and export a few per token embeddings, and compare them with what you have output here. |
|
There's a model registry that exposes paths to released models on GCS: https://storage.googleapis.com/moz-fx-translations-data--303e-prod-translations-data/db/models.json Then you can just provide a language pair in the arguments and download the models from there. |
|
We discussed that you might be interested in the non-finetuned students or teachers. The paths to these models are stored in the SQLite file that you can download from gs://moz-fx-translations-data--303e-prod-translations-data/db/db.sqlite So, it would be an SQL query to access those. Let me know if you need help with understanding the structure. |
Ticket: https://mozilla-hub.atlassian.net/browse/AIAUG-39
How to use
python utils/export_static_embedding.py \ --model-url "https://firefox-settings-attachments.cdn.mozilla.net/main-workspace/translations-models-v2/3a84a714-bfc7-4e52-81a1-b08d5d9421d2.zst" \ --vocab-url "https://firefox-settings-attachments.cdn.mozilla.net/main-workspace/translations-models-v2/671f09a0-12a2-436e-94d8-2575a4594c2a.zst" \ --output-dir static_embedding_export \ --write-model2vecHow to load exported embeddings
pip install model2vec tokenizersfrom model2vec import StaticModelmodel = StaticModel.from_pretrained("static_embedding_export/model2vec_model")