Skip to content

Add img-understanding task group (lmms-eval) - #125

Open
harshraj172 wants to merge 11 commits into
mainfrom
harsh/vqa-task-suite
Open

harshraj172 wants to merge 11 commits into
mainfrom
harsh/vqa-task-suite

Conversation

@harshraj172

@harshraj172 harshraj172 commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Adds an img-understanding task group (VQAv2, GQA, TextVQA, ScienceQA-img, MMMU, MMMU-Pro standard, all 0-shot), running on lmms-eval.

Parity results:

Model Task Ours Reference
llava-hf/llava-1.5-7b-hf gqa 60.5 62.0
llava-hf/llava-1.5-7b-hf scienceqa_img 66.1 70.4
llava-hf/llava-1.5-7b-hf textvqa_val 48.1 46.1
Qwen/Qwen2.5-VL-7B-Instruct mmmu_val 51.3 58.6
Qwen/Qwen2.5-VL-7B-Instruct mmmu_pro_standard 36.6 41.0

harshraj172 and others added 3 commits September 30, 2026 16:16
Adds two new task groups backed by lmms-eval (a separate harness for
vision-language models, since lm-eval-harness/evalchemy don't support
image inputs):

- vqa: VQAv2, GQA, TextVQA, ScienceQA-img (0-shot)
- mmmu: MMMU and MMMU-Pro standard (10-option) (0-shot)

template.sbatch gains a new lmms_eval case that resolves the right
lmms-eval model class (llava_hf, qwen2_5_vl, qwen2_vl) from the
checkpoint's own config.json model_type, since unlike the other suites
lmms-eval has no single architecture-agnostic model wrapper. It also
sets LMMS_EVAL_DATASETS_CACHE (lmms-eval otherwise silently redirects
dataset caching off the shared filesystem on offline compute nodes)
and HF_XET_CACHE (the Xet transfer backend ignores HF_HOME and can
blow a tight home-directory quota).

Validated for parity against public numbers:
- llava-hf/llava-1.5-7b-hf on vqa: GQA/TextVQA/ScienceQA-img within a
  few points of lmms-eval's own reported LLaVA-1.5-7B numbers
- Qwen/Qwen2.5-VL-7B-Instruct on mmmu: within range of the model's
  official card numbers (MMMU 58.6, MMMU-Pro standard 41.0)

docs/VENV.md documents venv setup, including the decord workaround
needed on linux-aarch64 (no wheel for that platform) and the HF_TOKEN
quirk some of the vqa/mmmu dataset loaders need even though the repos
are public.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three real bugs surfaced by actually running the new template.sbatch
lmms_eval case end-to-end (not just rendering/syntax-checking it):

1. Model-type detection read config.json directly off disk, which only
   works for a local checkpoint directory. --skip_checks (or any model
   never locally expanded) leaves model_path as a bare HF Hub id, so
   the case fell through to "no known model class" for every job.
   Fixed by resolving the architecture via
   transformers.AutoConfig.from_pretrained, which works for both local
   paths and Hub ids, the same way lm_eval/evalchemy let HF resolve it
   internally.

2. llava_hf only auto-assigns device_map from its own per-rank
   accelerator index when accelerate launch uses >1 process; with our
   normal single-GPU eval jobs (--num-processes 1) it's left empty and
   crashes on load ("Device string must not be empty" /
   "found ." depending on transformers version). Fixed by explicitly
   passing device_map=cuda:0 in model_args when GPUS_PER_NODE == 1.

3. Triton's kernel cache defaults to ~/.triton/cache regardless of
   HF_HOME, same category of bug as the HF_XET_CACHE fix already in
   this branch. A real MMMU/MMMU-Pro run hit the home-directory quota
   9 minutes in (SLURM still reported the job COMPLETED with exit 0,
   since the crash happened inside the python process after SLURM's
   own bookkeeping). Fixed with a TRITON_CACHE_DIR export next to the
   existing HF_XET_CACHE one.

Also fixes oellm.main.collect_results crashing on lmms-eval's textvqa_val
output: that task writes an extra leaderboard-submission JSON (a bare
list) alongside the normal results dict, which broke the unconditional
data.get(...) calls. Now skips any JSON file that isn't a dict.

All four fixes were found and verified via real `oellm-eval schedule`
submissions on JUPITER, not just template rendering: a --limit 4 vqa
live-test now completes end-to-end with sane per-task scores, and the
full mmmu/mmmu-pro run completed and was verified with `oellm-eval
collect` (mmmu_val 51.33%, mmmu_pro_standard 36.59%, vs. Qwen2.5-VL-7B-
Instruct's official card numbers of 58.6%/41.0% -- gaps in the same
range and direction as the vqa-group's llava-hf parity check, likely
from the same root cause: lmms-eval's reimplementation of each
benchmark's prompt/scoring protocol differing from the model authors'
own eval harness).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same six tasks (VQAv2, GQA, TextVQA, ScienceQA-img, MMMU, MMMU-Pro
standard), one group instead of two, since they're all 0-shot lmms-eval
image-understanding tasks. mmmu_val/mmmu_pro_standard override the
group's metric (exact_match) with their own (mmmu_acc).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
@harshraj172 harshraj172 changed the title Add vqa and mmmu task groups (lmms-eval) Add img-understanding task group (lmms-eval) Oct 1, 2026
harshraj172 and others added 4 commits October 1, 2026 15:19
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
…doc gap

lmms_eval case now passes trust_remote_code=True in model_args (the
model-type probe already used it) and --num_fewshot "$n_shot" (every
other suite branch already forwards shot count; this one silently
didn't).

docs/VENV.md: VQAv2 also needs HF_TOKEN set, not just GQA/ScienceQA-img.
An empty HF_TOKEN is indistinguishable from unset to huggingface_hub,
so the old doc step silently did nothing on a machine with no cached
token; now points at huggingface-cli login instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
@harshraj172
harshraj172 requested a review from geoalgo October 1, 2026 14:10

@geoalgo geoalgo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, just have one small comment.

Comment on lines +280 to +288
LMMS_MODEL_TYPE=$(python -c "
import sys
from transformers import AutoConfig
try:
model_type = AutoConfig.from_pretrained(sys.argv[1], trust_remote_code=True).model_type
except Exception:
model_type = ''
print(model_type)
" "$model_path" 2>/dev/null)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need this?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lmms-eval doesn't have one generic model class like lm_eval --model hf, each VLM family needs its own specific one, and --model has to match exactly. Since schedule only gets a model path/id, not that class name, we need some way to figure out which one to pass. So, we use AutoConfig to extract the model type

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

makes sense can you leave a comment so that it is clear when reading the code?

Comment thread oellm/resources/template.sbatch Outdated
;;
esac
if [ -n "$LMMS_EVAL_MODEL" ]; then
LMMS_MODEL_ARGS="pretrained=$model_path,trust_remote_code=True"

@islobozhan islobozhan Oct 1, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please double check that trust_remote_code=True works with the Qwen models? I see it was added in the last commit, and it might be an issue with these models?

qwen2_5_vl.py and qwen2_vl.py both have no trust_remote_code parameter
and assert kwargs == {} in __init__, so passing it there crashes with
AssertionError: Unexpected kwargs: {'trust_remote_code': True}.
llava_hf.py does accept it, so it stays there.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
harshraj172 and others added 3 commits October 2, 2026 04:14
lm_eval's --output_path treats any path as a directory regardless of
a .json suffix, so a RUN_ID.json path becomes a directory with the
real result nested inside it. rglob("*.json") matched that directory
too, and open() on it crashed with IsADirectoryError.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
mixturevitae2's own checkpoints train on images as SEED-2 discrete
tokens spliced into plain text, not through a vision encoder, so none
of the existing lmms-eval model classes (built for pixel-input VLMs)
can evaluate them. Adds seed2_omni, registered through lmms-eval's own
plugin mechanism (LMMS_EVAL_PLUGINS) rather than patching lmms-eval
itself: it encodes each image via mixturevitae2's SEED-2 backend,
splices the result into the prompt the same way training data does,
and runs ordinary text generation.

template.sbatch's model-type probe now also checks the tokenizer for
the `<seed2_0>` special token, since these checkpoints report
model_type: qwen3 in config.json just like any plain text model.

Validated end to end on JUPITER: direct lmms_eval eval --model
seed2_omni run, then a full oellm-eval schedule submission through
the real auto-detection + accelerate launch path, both producing real
generations and scores on textvqa_val and all six img-understanding
tasks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016EhV9asoFcApJnHrxLuRdf
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants