Repository navigation
Conversation
A sampled, per-language version of Belebele, Meta FAIR's parallel multiple-choice reading comprehension benchmark. Current AI asked us to host it and run it against their models. - scripts/build_belebele_sample.py pulls facebook/belebele from the Hugging Face datasets-server API. It samples 40 questions once (seeded, 10 per correct letter) and writes the same questions in each language, so scores compare across languages. - One blueprint per language, scored on the answer letter alone, so there are no judge calls. The scoring pattern accepts "B", "(B)", "**B**", "Answer: B" and "B. <option text>". - Each file is marked CC BY-SA 4.0, Belebele's license, not the repo's CC0 default. Don't merge until the license question is settled. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
…em message Apertus on Current AI's router rejects some requests that carry a system message (11 of 40 Catalan questions came back 400), while the same requests with the instruction folded into the user message are accepted. Every model now gets the instruction in the same place. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Blueprint Contribution
Don't merge yet. Belebele is licensed CC BY-SA 4.0, but everything in this repo is dedicated to the public domain (CC0). That needs a decision first (see Notes).
Blueprint Details
blueprints/benchmarks/belebele/belebele-sample-<lang>.yml(17 files)potluck:aisingapore/Qwen-SEA-LION-v4-32B-IT,potluck:aisingapore/Gemma-SEA-LION-v4-27B-IT, and Apertus once hosted), plus a few baseline models for context.models:gets added before merge.What This Blueprint Tests
These are sampled versions of Belebele (Meta FAIR). Each question gives a FLORES-200 passage, a question about it, and four answers. Current AI asked us to host it and run it on their models. The results would feed the per-language model picker in Current AI's product, Alpha.
20260929) with 10 per correct letter, so a model that favours one letter gains nothing. The same 40 questions appear in every language, so scores compare across languages. They are not comparable to full-benchmark Belebele scores, and each file's description says so.$imatcheson the answer letter), with no LLM judge. It acceptsB,B.,(B),**B**,Answer: BandB. <option text>. It rejects other letters and words that merely start with the letter ("Based on…"). A reply that doesn't start with a letter is scored wrong.python3 scripts/build_belebele_sample.py(options--n,--seed,--langs). Needs only PyYAML.Checklist
blueprints/users/<my-github-username>/directory. It's inblueprints/benchmarks/, alongside the other benchmark ports..github/scripts/validate_blueprints.pypasses on all 17 files, and the app'sparseAndNormalizeBlueprintparses them.id. IDs arebb-<hash of source link>-q<n>on purpose: the same ID means the same question in every language.$not_*functions instead ofshould_notblocks where applicable (not needed here)$imatches: all 24 replies were scorable, and each model got 5 of 6 right in each language.Notes
blueprints/users/…paths.🤖 Generated with Claude Code
https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
Generated by Claude Code