Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/TASKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,7 @@ Languages are **derived in code** — there is no `languages` field to set in th
YAML. A task resolves to a canonical [`lang_Script`](https://en.wikipedia.org/wiki/IETF_language_tag)
code (e.g. `deu_Latn`) from, in order:

1. **`flores200:src-tgt` task names** → the non-English side(s) of the pair.
1. **`flores200:src-tgt` (and `flores200_src-tgt_bpb`) task names** → the non-English side(s) of the pair.
2. **The `{lang}` value** substituted into a `valid_langs` template (preferred
for new multilingual groups — see the template expansion above).
3. **The task's `subset`** (e.g. `de`, `german`, `deu_Latn` all fold to
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Belebele in cloze form: the continuation is the answer text rather than the
# letter "A".."D", so BPB measures comprehension instead of format.
# Per-language tasks live in belebele/.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- belebele_bpb
dataset_path: facebook/belebele
test_split: test
fewshot_split: test
fewshot_config:
sampler: first_n
doc_to_text: "{{flores_passage}}\nQuestion: {{question.strip()}}\nAnswer:"
doc_to_target: "{{['1', '2', '3', '4'].index(correct_answer_num)}}"
doc_to_choice: "{{[mc_answer1, mc_answer2, mc_answer3, mc_answer4]}}"
process_results: !function utils.process_results_belebele
22 changes: 22 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/_bpb_template_yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Shared config for the BPB (bits-per-byte) task variants.
#
# Each task keeps its upstream prompt format and dataset so that BPB is
# measured on exactly the same documents as the accuracy-based task, but
# overrides process_results (see utils.py) so that the gold continuation's
# (loglikelihood, byte_length) pair is emitted for the bits_per_byte
# aggregation. acc/acc_norm come along for free.
tag:
- bpb
output_type: multiple_choice
metric_list:
- metric: bits_per_byte
aggregation: bits_per_byte
higher_is_better: false
- metric: acc
aggregation: mean
higher_is_better: true
- metric: acc_norm
aggregation: mean
higher_is_better: true
metadata:
version: 1.0
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# FLORES-200: BPB of the reference translation given the source sentence,
# using the same "DE: <source> EN:" prompt shape as lighteval's flores200 CF
# formulation. Included by the two direction templates below; per-pair tasks
# live in flores200/.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- flores200_bpb
dataset_path: facebook/flores
test_split: devtest
fewshot_split: dev
fewshot_config:
sampler: first_n
output_type: loglikelihood
# doc_to_target supplies the leading space itself (see utils.py).
target_delimiter: ""
metric_list:
- metric: bits_per_byte
aggregation: bits_per_byte
higher_is_better: false
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: _flores200_bpb_template_yaml
doc_to_target: !function utils.doc_to_target_flores_eng_xx
process_results: !function utils.process_results_flores_eng_xx
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: _flores200_bpb_template_yaml
doc_to_target: !function utils.doc_to_target_flores_xx_eng
process_results: !function utils.process_results_flores_xx_eng
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# Global-MMLU in cloze form, like _mmlu_stem_bpb_template_yaml. Each language
# is a single task over the whole test split (all 57 subjects), so BPB is one
# byte-weighted corpus aggregate per language. Per-language tasks live in
# global_mmlu/.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- global_mmlu_bpb
dataset_path: CohereForAI/Global-MMLU
test_split: test
fewshot_split: dev
fewshot_config:
sampler: first_n
doc_to_text: "Question: {{question.strip()}}\nAnswer:"
doc_to_target: "{{['A', 'B', 'C', 'D'].index(answer.strip())}}"
doc_to_choice: "{{[option_a, option_b, option_c, option_d]}}"
process_results: !function utils.process_results_global_mmlu
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# INCLUDE in cloze form, like _mmlu_stem_bpb_template_yaml. Each language is a
# single task over the whole test split (all domains). Per-language tasks live
# in include/.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- include_bpb
dataset_path: CohereForAI/include-base-44
test_split: test
doc_to_text: "Question: {{question.strip()}}\nAnswer:"
doc_to_target: answer
doc_to_choice: "{{[option_a, option_b, option_c, option_d]}}"
process_results: !function utils.process_results_include
20 changes: 20 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/_mgsm_bpb_template_yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# MGSM: BPB of the final answer number after the upstream lm-eval
# mgsm_direct_<lang> prompt (the test split has no worked solutions).
# Per-language tasks live in mgsm/ and carry the localized prompt.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- mgsm_bpb
dataset_path: jbross-ibm-research/mgsm
training_split: train
test_split: test
output_type: loglikelihood
# doc_to_target supplies the leading space itself (see utils.py).
target_delimiter: ""
doc_to_target: !function utils.doc_to_target_mgsm
process_results: !function utils.process_results_mgsm
metric_list:
- metric: bits_per_byte
aggregation: bits_per_byte
higher_is_better: false
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# MMLU in cloze/continuation form: the continuation is the answer text rather
# than the letter "A".."D", so BPB measures knowledge instead of format.
# Every subject carries the `mmlu_stem_bpb` tag, so a single
# `--tasks mmlu_stem_bpb` invocation runs all 19 STEM subjects in one process
# and writes them to one results JSON.
include: _bpb_template_yaml
tag:
- bpb
- mmlu_stem_bpb
dataset_path: cais/mmlu
test_split: test
fewshot_split: dev
fewshot_config:
sampler: first_n
doc_to_text: "Question: {{question.strip()}}\nAnswer:"
doc_to_target: "{{answer}}"
doc_to_choice: "{{choices}}"
process_results: !function utils.process_results_mmlu
13 changes: 13 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/_xcopa_bpb_template_yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# XCOPA with the upstream lm-eval prompt; the per-language connector comes from
# each task's doc_to_text. Per-language tasks live in xcopa/.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- xcopa_bpb
dataset_path: cambridgeltl/xcopa
validation_split: validation
test_split: test
doc_to_target: label
doc_to_choice: !function utils.doc_to_choice_xcopa
process_results: !function utils.process_results_xcopa
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# XStoryCloze with the upstream lm-eval prompt. Per-language tasks live in
# xstorycloze/.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- xstorycloze_bpb
dataset_path: juletxara/xstory_cloze
training_split: train
validation_split: eval
doc_to_text: "{{[input_sentence_1, input_sentence_2, input_sentence_3, input_sentence_4]|join(' ')}}"
doc_to_target: "{{answer_right_ending-1}}"
doc_to_choice: "{{[sentence_quiz1, sentence_quiz2]}}"
process_results: !function utils.process_results_xstorycloze
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# XWinograd shares the Winogrande schema, so it reuses winogrande_bpb's
# multiple_input functions: BPB is measured on the shared continuation.
# Per-language tasks live in xwinograd/.
include: _bpb_template_yaml
tag:
- bpb
- bpb_multilingual
- xwinograd_bpb
dataset_path: Muennighoff/xwinograd
test_split: test
doc_to_text: !function utils.doc_to_text_winogrande
doc_to_target: !function utils.doc_to_target_winogrande
doc_to_choice: !function utils.doc_to_choice_winogrande
process_results: !function utils.process_results_winogrande
11 changes: 11 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/arc_challenge_bpb.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
include: _bpb_template_yaml
task: arc_challenge_bpb
dataset_path: allenai/ai2_arc
dataset_name: ARC-Challenge
training_split: train
validation_split: validation
test_split: test
doc_to_text: "Question: {{question}}\nAnswer:"
doc_to_target: "{{choices.label.index(answerKey)}}"
doc_to_choice: "{{choices.text}}"
process_results: !function utils.process_results_arc
11 changes: 11 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/arc_easy_bpb.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
include: _bpb_template_yaml
task: arc_easy_bpb
dataset_path: allenai/ai2_arc
dataset_name: ARC-Easy
training_split: train
validation_split: validation
test_split: test
doc_to_text: "Question: {{question}}\nAnswer:"
doc_to_target: "{{choices.label.index(answerKey)}}"
doc_to_choice: "{{choices.text}}"
process_results: !function utils.process_results_arc
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_bul_Cyrl_bpb
dataset_name: bul_Cyrl
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_ces_Latn_bpb
dataset_name: ces_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_dan_Latn_bpb
dataset_name: dan_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_deu_Latn_bpb
dataset_name: deu_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_ell_Grek_bpb
dataset_name: ell_Grek
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_eng_Latn_bpb
dataset_name: eng_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_est_Latn_bpb
dataset_name: est_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_fin_Latn_bpb
dataset_name: fin_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_fra_Latn_bpb
dataset_name: fra_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_hrv_Latn_bpb
dataset_name: hrv_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_hun_Latn_bpb
dataset_name: hun_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_ita_Latn_bpb
dataset_name: ita_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_lit_Latn_bpb
dataset_name: lit_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_lvs_Latn_bpb
dataset_name: lvs_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_mlt_Latn_bpb
dataset_name: mlt_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_nld_Latn_bpb
dataset_name: nld_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_nob_Latn_bpb
dataset_name: nob_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_pol_Latn_bpb
dataset_name: pol_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_por_Latn_bpb
dataset_name: por_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_ron_Latn_bpb
dataset_name: ron_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_slk_Latn_bpb
dataset_name: slk_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_slv_Latn_bpb
dataset_name: slv_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_spa_Latn_bpb
dataset_name: spa_Latn
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
include: ../_belebele_bpb_template_yaml
task: belebele_swe_Latn_bpb
dataset_name: swe_Latn
10 changes: 10 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/boolq_bpb.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
include: _bpb_template_yaml
task: boolq_bpb
dataset_path: aps/super_glue
dataset_name: boolq
training_split: train
validation_split: validation
doc_to_text: "{{passage}}\nQuestion: {{question}}?\nAnswer:"
doc_to_target: label
doc_to_choice: ["no", "yes"]
process_results: !function utils.process_results_boolq
11 changes: 11 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/commonsense_qa_bpb.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
# Cloze form: the continuation is the answer *text*, not the letter "A".."E".
# Scoring BPB on a single letter would measure format-following, not knowledge.
include: _bpb_template_yaml
task: commonsense_qa_bpb
dataset_path: tau/commonsense_qa
training_split: train
validation_split: validation
doc_to_text: "Question: {{ question.strip() }}\nAnswer:"
doc_to_target: "{{choices.label.index(answerKey)}}"
doc_to_choice: "{{choices.text}}"
process_results: !function utils.process_results_commonsense_qa
10 changes: 10 additions & 0 deletions oellm/resources/custom_lm_eval_tasks/bpb/copa_bpb.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
include: _bpb_template_yaml
task: copa_bpb
dataset_path: aps/super_glue
dataset_name: copa
training_split: train
validation_split: validation
doc_to_text: !function utils.doc_to_text_copa
doc_to_target: !function utils.doc_to_target_copa
doc_to_choice: !function utils.doc_to_choice_copa
process_results: !function utils.process_results_copa
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_als_Latn-eng_Latn_bpb
dataset_name: als_Latn-eng_Latn
doc_to_text: 'SQ: {{sentence_als_Latn}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_bos_Latn-eng_Latn_bpb
dataset_name: bos_Latn-eng_Latn
doc_to_text: 'BS: {{sentence_bos_Latn}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_bul_Cyrl-eng_Latn_bpb
dataset_name: bul_Cyrl-eng_Latn
doc_to_text: 'BG: {{sentence_bul_Cyrl}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_cat_Latn-eng_Latn_bpb
dataset_name: cat_Latn-eng_Latn
doc_to_text: 'CA: {{sentence_cat_Latn}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_ces_Latn-eng_Latn_bpb
dataset_name: ces_Latn-eng_Latn
doc_to_text: 'CS: {{sentence_ces_Latn}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_dan_Latn-eng_Latn_bpb
dataset_name: dan_Latn-eng_Latn
doc_to_text: 'DA: {{sentence_dan_Latn}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_deu_Latn-eng_Latn_bpb
dataset_name: deu_Latn-eng_Latn
doc_to_text: 'DE: {{sentence_deu_Latn}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_xx_eng_bpb_template_yaml
task: flores200_ell_Grek-eng_Latn_bpb
dataset_name: ell_Grek-eng_Latn
doc_to_text: 'EL: {{sentence_ell_Grek}} EN:'
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
include: ../_flores200_eng_xx_bpb_template_yaml
task: flores200_eng_Latn-als_Latn_bpb
dataset_name: eng_Latn-als_Latn
doc_to_text: 'EN: {{sentence_eng_Latn}} SQ:'
Loading
Loading