Skip to content

[DRAFT][FOR REVIEW] Merging elliot eval platform with oellm - #107

Closed
islobozhan wants to merge 100 commits into
OpenEuroLLM:mainfrom
elliot-project:main
Closed

islobozhan wants to merge 100 commits into
OpenEuroLLM:mainfrom
elliot-project:main

Conversation

@islobozhan

Copy link
Copy Markdown

As discussed with David, we plan to merge the two repositories at some point. This pull request will be used to compare the two versions and review the changes at a high level, gradually, due to the large number of diff lines.

Foundational refactoring & Image modality
Integration of RegionReasoner: Region-Grounded Multi-Round Visual Reasoning (ICLR 2026) benchmark: https://arxiv.org/pdf/2602.03733. Code is available here: https://github.com/lmsdss/RegionReasoner.
1. Fix HF_HOME crash when env var is unset
2. Validate required cluster variables after loading clusters.yaml
3. Replace hardcoded lmms-eval adapter detection with LMMS_MODEL_ADAPTERS lookup
4. Remove magic-number time budgeting to use TIME_LIMIT from clusters.yaml directly
5. Decompose main.py into constants.py and results.py
1. Add new documentation table grouped by category (Cluster Setup, Environment & Infrastructure, Extending the Platform)
2. Add oellm/contrib/README.md as a registry of community-contributed benchmarks, starting with RegionReasoner
3. Link contrib registry and contributing guide from main README
…arding

- Add --mem=0 to SLURM template for full node memory allocation
- Implement streaming pre-sharding
- Fix CUDA_VISIBLE_DEVICES assignment per shard
- Fix dataset download to fetch all splits (refcocog + refcocoplus)
- Rename region_reasoner → regiondial_bench to match paper terminology
- Allow running splits independently (regiondial-refcocog, regiondial-refcocoplus)
Sync additional task with upstream: 
1. Add missing task (belebele for norwegian)
Merge commit to sync fork with OpenEuroLLM/oellm-cli main branch.

Resolves the "1 commit behind" divergence that persisted after the
previous squash-merge sync.
Expand the HF cache section for Cineca Leonardo with available storage areas, quotas, and purge policies to help users pick the right location.
Sync with upstream - add snellius support
…nch aggregation (#10)

* Fix per-round metric computation in RegionDial-Bench aggregation
* Update tests for RegionDial-Bench
* Fix minor typos
Merge with upstream local runner
Switch CLI to Typer, added runs via config
Add video understanding task group with 5 benchmarks via lmms-eval: 
1. VideoMMMU, 
2. EgoSchema, 
3. VideoMME, 
4. ActivityNet-QA, 
5. LongVideoBench
--batch_size auto was hardcoded for the lm_eval and evalchemy engines. lm_eval's auto search halves on CUDA OOM, so with no GPU nothing bounds it: it probes at batch 64 × prompt length and exhausts RAM. TimeSeriesExam (8 docs) reached 19.7 GB RSS and never produced a score; TabFact behaved the same. Short-prompt tasks are unaffected, which is why this went unnoticed.

The template now reads ${LM_EVAL_BATCH_SIZE:-auto}, resolved by the scheduler to 8 for --local and auto on the cluster, with BATCH_SIZE overriding either — the same convention _resolve_additional_model_args already uses for lighteval, so one variable now covers all three engines. Cluster behaviour is unchanged.
…EU task groups

Merge with upstream Okapi HellaSwag, MMMLU, XCOPA, MultiBLIMP EU task groups
…_norm note

Merge with upstream: robust cluster recognition, lighteval acc_norm note
…lacement (LMMS_MODEL_ARGS), VENV.md install fix
…-all-metrics, bigbench generate_until fix; lmms-eval adapter and device fixes

Merge with upstream 19: per-group metric keys, collect --fetch-all-metrics, bigbench generate_until fix; lmms-eval adapter and device fixes
…in pre-download; keep wsc273 spec for the locked lm-eval 0.4.12
Merge with upstream - drop trust_remote_code on datasets>=4 in pre-download; keep wsc273 spec for the locked lm-eval 0.4.12
…teval ROCm fix, FLORES 4-shot, multilingual-oellm-eu super group
…ix, FLORES 4-shot

Merge with upstream: missing-result failures, LUMI lighteval fix, FLORES 4-shot
…eck, group scores

Merge with upstream : batch lm-eval tasks per call, HF_HOME check, group scores
…ellm-eval push, collect --push)

Push collected results to the dashboard from the login node (oellm-eval push, collect --push)
Merge with upstream
Merge with upstream: SIB-200 scored by acc instead of acc_norm
@islobozhan islobozhan closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant