Skip to content

Draft: Belebele sample blueprints (17 languages × 40 questions) - #29

Draft
nojibe wants to merge 2 commits into
mainfrom
claude/belebele-sample
Draft

nojibe wants to merge 2 commits into
mainfrom
claude/belebele-sample

Conversation

@nojibe

@nojibe nojibe commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Blueprint Contribution

Don't merge yet. Belebele is licensed CC BY-SA 4.0, but everything in this repo is dedicated to the public domain (CC0). That needs a decision first (see Notes).

Blueprint Details

  • Blueprint ID: one per language, taken from the filename: blueprints/benchmarks/belebele/belebele-sample-<lang>.yml (17 files)
  • Category/Focus: multilingual reading comprehension (multiple choice)
  • Models to test: not set yet. The intended set is the models on Current AI's router (potluck:aisingapore/Qwen-SEA-LION-v4-32B-IT, potluck:aisingapore/Gemma-SEA-LION-v4-27B-IT, and Apertus once hosted), plus a few baseline models for context. models: gets added before merge.

What This Blueprint Tests

These are sampled versions of Belebele (Meta FAIR). Each question gives a FLORES-200 passage, a question about it, and four answers. Current AI asked us to host it and run it on their models. The results would feed the per-language model picker in Current AI's product, Alpha.

  • Sampling: 40 of Belebele's 900 questions, drawn once (seed 20260929) with 10 per correct letter, so a model that favours one letter gains nothing. The same 40 questions appear in every language, so scores compare across languages. They are not comparable to full-benchmark Belebele scores, and each file's description says so.
  • Languages: English, Catalan, Spanish, Basque, Portuguese, French, German, Italian, Indonesian, Malay, Vietnamese, Thai, Tagalog, Swahili, Hindi, Arabic and Chinese (Simplified). Galician was planned, but Belebele doesn't include it.
  • Scoring: deterministic ($imatches on the answer letter), with no LLM judge. It accepts B, B., (B), **B**, Answer: B and B. <option text>. It rejects other letters and words that merely start with the letter ("Based on…"). A reply that doesn't start with a letter is scored wrong.
  • Size: 680 prompts per model across the 17 files.
  • Regenerate: python3 scripts/build_belebele_sample.py (options --n, --seed, --langs). Needs only PyYAML.

Checklist

  • My blueprint is in blueprints/users/<my-github-username>/ directory. It's in blueprints/benchmarks/, alongside the other benchmark ports.
  • Blueprint YAML is valid and follows the blueprint format. .github/scripts/validate_blueprints.py passes on all 17 files, and the app's parseAndNormalizeBlueprint parses them.
  • Each prompt has a meaningful, descriptive id. IDs are bb-<hash of source link>-q<n> on purpose: the same ID means the same question in every language.
  • Blueprint has clear success criteria
  • I've used $not_* functions instead of should_not blocks where applicable (not needed here)
  • I've tested the blueprint locally. I sent 6 Catalan and 6 Thai questions to both SEA-LION models through the app's dispatcher and $imatches: all 24 replies were scorable, and each model got 5 of 6 right in each language.
  • I agree to dedicate my contribution to the public domain under CC0 1.0 Universal. Not possible for this content, see Notes.

Notes

  • License: Belebele is CC BY-SA 4.0 (share-alike), so these files can't be dedicated to the public domain. Each file's header and description state CC BY-SA 4.0 with attribution. Before merging we need to decide whether the repo can hold files with their own license. If so, the README's "all content is CC0" line needs an exception. The WPS v5.1 blueprint merged in feat(blueprints): Add v5.1-evaluating-ai-performance-in-women-peace-security-scenarios.yml #24 has the same issue: its header says CC BY 4.0.
  • Automated PR evaluation: none will run. The PR bot only evaluates blueprints/users/… paths.
  • Router capacity: Current AI's router is a staging one with no fallbacks. We should confirm with them (Sean) before running the full set.

🤖 Generated with Claude Code

https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj


Generated by Claude Code

A sampled, per-language version of Belebele, Meta FAIR's parallel
multiple-choice reading comprehension benchmark. Current AI asked us to
host it and run it against their models.

- scripts/build_belebele_sample.py pulls facebook/belebele from the
  Hugging Face datasets-server API. It samples 40 questions once
  (seeded, 10 per correct letter) and writes the same questions in each
  language, so scores compare across languages.
- One blueprint per language, scored on the answer letter alone, so
  there are no judge calls. The scoring pattern accepts "B", "(B)",
  "**B**", "Answer: B" and "B. <option text>".
- Each file is marked CC BY-SA 4.0, Belebele's license, not the repo's
  CC0 default. Don't merge until the license question is settled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
…em message

Apertus on Current AI's router rejects some requests that carry a
system message (11 of 40 Catalan questions came back 400), while the
same requests with the instruction folded into the user message are
accepted. Every model now gets the instruction in the same place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants