Repository navigation
[DRAFT][FOR REVIEW] Merging elliot eval platform with oellm - #107
Closed
islobozhan wants to merge 100 commits into
Closed
islobozhan wants to merge 100 commits into
islobozhan wants to merge 100 commits into
Conversation
Foundational refactoring & Image modality
Integration of RegionReasoner: Region-Grounded Multi-Round Visual Reasoning (ICLR 2026) benchmark: https://arxiv.org/pdf/2602.03733. Code is available here: https://github.com/lmsdss/RegionReasoner.
1. Fix HF_HOME crash when env var is unset 2. Validate required cluster variables after loading clusters.yaml 3. Replace hardcoded lmms-eval adapter detection with LMMS_MODEL_ADAPTERS lookup 4. Remove magic-number time budgeting to use TIME_LIMIT from clusters.yaml directly 5. Decompose main.py into constants.py and results.py
1. Add new documentation table grouped by category (Cluster Setup, Environment & Infrastructure, Extending the Platform) 2. Add oellm/contrib/README.md as a registry of community-contributed benchmarks, starting with RegionReasoner 3. Link contrib registry and contributing guide from main README
…arding - Add --mem=0 to SLURM template for full node memory allocation - Implement streaming pre-sharding - Fix CUDA_VISIBLE_DEVICES assignment per shard - Fix dataset download to fetch all splits (refcocog + refcocoplus) - Rename region_reasoner → regiondial_bench to match paper terminology - Allow running splits independently (regiondial-refcocog, regiondial-refcocoplus)
Sync additional task with upstream: 1. Add missing task (belebele for norwegian)
Merge commit to sync fork with OpenEuroLLM/oellm-cli main branch. Resolves the "1 commit behind" divergence that persisted after the previous squash-merge sync.
Expand the HF cache section for Cineca Leonardo with available storage areas, quotas, and purge policies to help users pick the right location.
Sync with upstream - add snellius support
…nch aggregation (#10) * Fix per-round metric computation in RegionDial-Bench aggregation * Update tests for RegionDial-Bench * Fix minor typos
Merge with upstream local runner
Switch CLI to Typer, added runs via config
Add video understanding task group with 5 benchmarks via lmms-eval: 1. VideoMMMU, 2. EgoSchema, 3. VideoMME, 4. ActivityNet-QA, 5. LongVideoBench
--batch_size auto was hardcoded for the lm_eval and evalchemy engines. lm_eval's auto search halves on CUDA OOM, so with no GPU nothing bounds it: it probes at batch 64 × prompt length and exhausts RAM. TimeSeriesExam (8 docs) reached 19.7 GB RSS and never produced a score; TabFact behaved the same. Short-prompt tasks are unaffected, which is why this went unnoticed.
The template now reads ${LM_EVAL_BATCH_SIZE:-auto}, resolved by the scheduler to 8 for --local and auto on the cluster, with BATCH_SIZE overriding either — the same convention _resolve_additional_model_args already uses for lighteval, so one variable now covers all three engines. Cluster behaviour is unchanged.
…EU task groups Merge with upstream Okapi HellaSwag, MMMLU, XCOPA, MultiBLIMP EU task groups
…_norm note Merge with upstream: robust cluster recognition, lighteval acc_norm note
…tric keys, collect --fetch-all-metrics
…lacement (LMMS_MODEL_ARGS), VENV.md install fix
…-all-metrics, bigbench generate_until fix; lmms-eval adapter and device fixes Merge with upstream 19: per-group metric keys, collect --fetch-all-metrics, bigbench generate_until fix; lmms-eval adapter and device fixes
…in pre-download; keep wsc273 spec for the locked lm-eval 0.4.12
Merge with upstream - drop trust_remote_code on datasets>=4 in pre-download; keep wsc273 spec for the locked lm-eval 0.4.12
…teval ROCm fix, FLORES 4-shot, multilingual-oellm-eu super group
…ix, FLORES 4-shot Merge with upstream: missing-result failures, LUMI lighteval fix, FLORES 4-shot
…check, group scores
…eck, group scores Merge with upstream : batch lm-eval tasks per call, HF_HOME check, group scores
…ellm-eval push, collect --push) Push collected results to the dashboard from the login node (oellm-eval push, collect --push)
Merge with upstream
Merge with upstream: SIB-200 scored by acc instead of acc_norm
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
As discussed with David, we plan to merge the two repositories at some point. This pull request will be used to compare the two versions and review the changes at a high level, gradually, due to the large number of diff lines.