Repository navigation
Add img-understanding task group (lmms-eval) - #125
harshraj172 wants to merge 11 commits into
Conversation
Adds two new task groups backed by lmms-eval (a separate harness for vision-language models, since lm-eval-harness/evalchemy don't support image inputs): - vqa: VQAv2, GQA, TextVQA, ScienceQA-img (0-shot) - mmmu: MMMU and MMMU-Pro standard (10-option) (0-shot) template.sbatch gains a new lmms_eval case that resolves the right lmms-eval model class (llava_hf, qwen2_5_vl, qwen2_vl) from the checkpoint's own config.json model_type, since unlike the other suites lmms-eval has no single architecture-agnostic model wrapper. It also sets LMMS_EVAL_DATASETS_CACHE (lmms-eval otherwise silently redirects dataset caching off the shared filesystem on offline compute nodes) and HF_XET_CACHE (the Xet transfer backend ignores HF_HOME and can blow a tight home-directory quota). Validated for parity against public numbers: - llava-hf/llava-1.5-7b-hf on vqa: GQA/TextVQA/ScienceQA-img within a few points of lmms-eval's own reported LLaVA-1.5-7B numbers - Qwen/Qwen2.5-VL-7B-Instruct on mmmu: within range of the model's official card numbers (MMMU 58.6, MMMU-Pro standard 41.0) docs/VENV.md documents venv setup, including the decord workaround needed on linux-aarch64 (no wheel for that platform) and the HF_TOKEN quirk some of the vqa/mmmu dataset loaders need even though the repos are public. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three real bugs surfaced by actually running the new template.sbatch
lmms_eval case end-to-end (not just rendering/syntax-checking it):
1. Model-type detection read config.json directly off disk, which only
works for a local checkpoint directory. --skip_checks (or any model
never locally expanded) leaves model_path as a bare HF Hub id, so
the case fell through to "no known model class" for every job.
Fixed by resolving the architecture via
transformers.AutoConfig.from_pretrained, which works for both local
paths and Hub ids, the same way lm_eval/evalchemy let HF resolve it
internally.
2. llava_hf only auto-assigns device_map from its own per-rank
accelerator index when accelerate launch uses >1 process; with our
normal single-GPU eval jobs (--num-processes 1) it's left empty and
crashes on load ("Device string must not be empty" /
"found ." depending on transformers version). Fixed by explicitly
passing device_map=cuda:0 in model_args when GPUS_PER_NODE == 1.
3. Triton's kernel cache defaults to ~/.triton/cache regardless of
HF_HOME, same category of bug as the HF_XET_CACHE fix already in
this branch. A real MMMU/MMMU-Pro run hit the home-directory quota
9 minutes in (SLURM still reported the job COMPLETED with exit 0,
since the crash happened inside the python process after SLURM's
own bookkeeping). Fixed with a TRITON_CACHE_DIR export next to the
existing HF_XET_CACHE one.
Also fixes oellm.main.collect_results crashing on lmms-eval's textvqa_val
output: that task writes an extra leaderboard-submission JSON (a bare
list) alongside the normal results dict, which broke the unconditional
data.get(...) calls. Now skips any JSON file that isn't a dict.
All four fixes were found and verified via real `oellm-eval schedule`
submissions on JUPITER, not just template rendering: a --limit 4 vqa
live-test now completes end-to-end with sane per-task scores, and the
full mmmu/mmmu-pro run completed and was verified with `oellm-eval
collect` (mmmu_val 51.33%, mmmu_pro_standard 36.59%, vs. Qwen2.5-VL-7B-
Instruct's official card numbers of 58.6%/41.0% -- gaps in the same
range and direction as the vqa-group's llava-hf parity check, likely
from the same root cause: lmms-eval's reimplementation of each
benchmark's prompt/scoring protocol differing from the model authors'
own eval harness).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same six tasks (VQAv2, GQA, TextVQA, ScienceQA-img, MMMU, MMMU-Pro standard), one group instead of two, since they're all 0-shot lmms-eval image-understanding tasks. mmmu_val/mmmu_pro_standard override the group's metric (exact_match) with their own (mmmu_acc). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
…doc gap lmms_eval case now passes trust_remote_code=True in model_args (the model-type probe already used it) and --num_fewshot "$n_shot" (every other suite branch already forwards shot count; this one silently didn't). docs/VENV.md: VQAv2 also needs HF_TOKEN set, not just GQA/ScienceQA-img. An empty HF_TOKEN is indistinguishable from unset to huggingface_hub, so the old doc step silently did nothing on a machine with no cached token; now points at huggingface-cli login instead. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
geoalgo
left a comment
There was a problem hiding this comment.
LGTM, just have one small comment.
| LMMS_MODEL_TYPE=$(python -c " | ||
| import sys | ||
| from transformers import AutoConfig | ||
| try: | ||
| model_type = AutoConfig.from_pretrained(sys.argv[1], trust_remote_code=True).model_type | ||
| except Exception: | ||
| model_type = '' | ||
| print(model_type) | ||
| " "$model_path" 2>/dev/null) |
There was a problem hiding this comment.
lmms-eval doesn't have one generic model class like lm_eval --model hf, each VLM family needs its own specific one, and --model has to match exactly. Since schedule only gets a model path/id, not that class name, we need some way to figure out which one to pass. So, we use AutoConfig to extract the model type
There was a problem hiding this comment.
makes sense can you leave a comment so that it is clear when reading the code?
| ;; | ||
| esac | ||
| if [ -n "$LMMS_EVAL_MODEL" ]; then | ||
| LMMS_MODEL_ARGS="pretrained=$model_path,trust_remote_code=True" |
There was a problem hiding this comment.
Could you please double check that trust_remote_code=True works with the Qwen models? I see it was added in the last commit, and it might be an issue with these models?
qwen2_5_vl.py and qwen2_vl.py both have no trust_remote_code parameter
and assert kwargs == {} in __init__, so passing it there crashes with
AssertionError: Unexpected kwargs: {'trust_remote_code': True}.
llava_hf.py does accept it, so it stays there.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
lm_eval's --output_path treats any path as a directory regardless of
a .json suffix, so a RUN_ID.json path becomes a directory with the
real result nested inside it. rglob("*.json") matched that directory
too, and open() on it crashed with IsADirectoryError.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
mixturevitae2's own checkpoints train on images as SEED-2 discrete tokens spliced into plain text, not through a vision encoder, so none of the existing lmms-eval model classes (built for pixel-input VLMs) can evaluate them. Adds seed2_omni, registered through lmms-eval's own plugin mechanism (LMMS_EVAL_PLUGINS) rather than patching lmms-eval itself: it encodes each image via mixturevitae2's SEED-2 backend, splices the result into the prompt the same way training data does, and runs ordinary text generation. template.sbatch's model-type probe now also checks the tokenizer for the `<seed2_0>` special token, since these checkpoints report model_type: qwen3 in config.json just like any plain text model. Validated end to end on JUPITER: direct lmms_eval eval --model seed2_omni run, then a full oellm-eval schedule submission through the real auto-detection + accelerate launch path, both producing real generations and scores on textvqa_val and all six img-understanding tasks. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
Adds an
img-understandingtask group (VQAv2, GQA, TextVQA, ScienceQA-img, MMMU, MMMU-Pro standard, all 0-shot), running on lmms-eval.Parity results: