Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
49 changes: 49 additions & 0 deletions docs/VENV.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,3 +119,52 @@ We use [Ali's fork](https://github.com/Ali-Elganzory/evalchemy) which includes a
```

> **Note:** `HF_ALLOW_CODE_EVAL=1` is required because MBPP (run via lm-eval-harness) uses HuggingFace's `code_eval` metric which executes model-generated code. The evalchemy benchmarks (GPQADiamond, MATH500, LiveCodeBench) do not require this variable as they handle code execution safely through internal guards.

## LMMs-Eval (image understanding)

The `img-understanding` task group (VQAv2, GQA, TextVQA, ScienceQA-img, MMMU, MMMU-Pro standard) runs via [lmms-eval](https://github.com/EvolvingLMMs-Lab/lmms-eval), since lm-eval-harness and evalchemy don't handle image inputs. lmms-eval has no single model wrapper the way those suites do; each VLM family needs its own registered model class. The `lmms_eval` case in `template.sbatch` resolves the right one (`llava_hf`, `qwen2_5_vl`, `qwen2_vl`, `seed2_omni`) from the checkpoint's `config.json` or tokenizer. Extend that case if you add a model family it doesn't recognize.

1. Create a venv and install dependencies:
```bash
uv venv --python 3.12 lmms-eval-venv
uv pip install --python lmms-eval-venv/bin/python -r requirements-venv-lmms-eval.txt
```

2. `decord` has no Linux-aarch64 wheel, but `llava_hf` and the Qwen VL model files import it unconditionally, even for image-only tasks. Stub it out; nothing in this task group calls it:
```bash
lmms-eval-venv/bin/python - <<'PY'
import pathlib, sysconfig
site = pathlib.Path(sysconfig.get_paths()["purelib"]) / "decord"
site.mkdir(exist_ok=True)
(site / "__init__.py").write_text('''def cpu(*args, **kwargs):
raise NotImplementedError("decord is not available on this platform")


class VideoReader:
def __init__(self, *args, **kwargs):
raise NotImplementedError("decord is not available on this platform")
''')
print("stubbed decord at:", site)
PY
```

3. VQAv2, GQA, and ScienceQA-img pass `token=True` in their dataset loading code, even though all three repos are public. Without a token this fails before the request ever reaches the network. Any token works, including an expired one, but it must be a real cached token: an empty `HF_TOKEN` is the same as unsetting it, so run `huggingface-cli login` first if you don't already have one, then:
```bash
export HF_TOKEN=$(cat ~/.cache/huggingface/token)
```

`llava_hf` has no real batching support, so the `lmms_eval` case always runs with `--batch_size 1`. Unlike lighteval and evalchemy, this isn't exposed as a CLI option.

4. `seed2_omni` (mixturevitae2's own omni checkpoints, detected via the `<seed2_0>` token in the tokenizer rather than `config.json`, since those checkpoints report `model_type: qwen3` like any plain text model) is registered through lmms-eval's plugin mechanism rather than shipped inside lmms-eval itself. Copy it into the venv:
```bash
cp -r oellm/resources/mv2_lmms_plugin lmms-eval-venv/lib/python3.12/site-packages/
```
It imports `multimodal_processing.model_backends.seed2` from a `mixturevitae2` checkout at runtime, found via `MV2_MULTIMODAL_DIR` (defaults to `/e/project1/jureap59/raj3/mixturevitae2`).

```bash
oellm-eval schedule \
--models llava-hf/llava-1.5-7b-hf \
--task_groups img-understanding \
--venv_path lmms-eval-venv \
--skip_checks true
```
6 changes: 5 additions & 1 deletion oellm/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -718,7 +718,7 @@ def _split_task_and_nshot(name: str) -> tuple[str, int | None]:
# ------------------------------------------------------------------
# 2. Recursively find all JSON result files.
# ------------------------------------------------------------------
json_files = sorted(results_path.rglob("*.json"))
json_files = sorted(p for p in results_path.rglob("*.json") if p.is_file())

if not json_files:
logging.warning(f"No JSON files found under {results_dir}")
Expand All @@ -735,6 +735,10 @@ def _split_task_and_nshot(name: str) -> tuple[str, int | None]:
with open(json_file) as f:
data = json.load(f)

if not isinstance(data, dict):
logging.debug(f"Skipping non-results JSON file: {json_file}")
continue

# Extract model name/path from a few common locations used in different
# versions of the result JSON schema.
model_name = (
Expand Down
Empty file.
1 change: 1 addition & 0 deletions oellm/resources/mv2_lmms_plugin/models/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
AVAILABLE_MODELS = {"seed2_omni": "Seed2Omni"}
96 changes: 96 additions & 0 deletions oellm/resources/mv2_lmms_plugin/models/seed2_omni.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
import os
import sys

import torch
from lmms_eval.api.model import lmms
from lmms_eval.api.registry import register_model
from tqdm import tqdm


@register_model("seed2_omni")
class Seed2Omni(lmms):
def __init__(self, pretrained, device="cuda", batch_size=1, **kwargs):
super().__init__()
assert kwargs == {}, f"Unexpected kwargs: {kwargs}"
assert int(batch_size) == 1, "seed2_omni only supports batch_size=1"

mv2_dir = os.environ.get(
"MV2_MULTIMODAL_DIR", "/e/project1/jureap59/raj3/mixturevitae2"
)
if mv2_dir not in sys.path:
sys.path.insert(0, mv2_dir)
from multimodal_processing.model_backends.seed2 import Backend, seed2_block

self._seed2_block = seed2_block
self._seed2 = Backend()
self._seed2.load()

from transformers import AutoModelForCausalLM, AutoTokenizer

self._device = torch.device(device)
self._tokenizer = AutoTokenizer.from_pretrained(
pretrained, trust_remote_code=True
)
self._model = (
AutoModelForCausalLM.from_pretrained(
pretrained, dtype=torch.bfloat16, trust_remote_code=True
)
.to(self._device)
.eval()
)
im_end_id = self._tokenizer.convert_tokens_to_ids("<|im_end|>")
self._eos_token_ids = [self._tokenizer.eos_token_id, im_end_id]
self.batch_size_per_gpu = 1

@property
def batch_size(self):
return self.batch_size_per_gpu

@property
def device(self):
return self._device

def loglikelihood(self, requests):
raise NotImplementedError("seed2_omni only supports generate_until tasks")

def generate_until_multi_round(self, requests):
raise NotImplementedError("seed2_omni does not support multi-round generation")

def generate_until(self, requests):
res = []
pbar = tqdm(
total=len(requests), disable=(self.rank != 0), desc="seed2_omni responding"
)
for request in requests:
context, gen_kwargs, doc_to_visual, doc_id, task, split = request.args
visuals = doc_to_visual(self.task_dict[task][split][doc_id])
if not isinstance(visuals, list):
visuals = [visuals]

blocks = []
for v in visuals:
ids = self._seed2.encode(v.convert("RGB"))
if ids:
blocks.append(self._seed2_block(ids))
prompt = f"{''.join(blocks)} {context}".strip() if blocks else context
chat = f"<|im_start|>user\n{prompt}<|im_end|>\n<|im_start|>assistant\n<think>\n</think>\n"

inputs = self._tokenizer(chat, return_tensors="pt").to(self._device)
do_sample = gen_kwargs.get("temperature", 0) > 0
with torch.no_grad():
out = self._model.generate(
**inputs,
max_new_tokens=gen_kwargs.get("max_new_tokens", 256),
do_sample=do_sample,
temperature=gen_kwargs.get("temperature") if do_sample else None,
pad_token_id=self._tokenizer.eos_token_id,
eos_token_id=self._eos_token_ids,
)
new_tokens = out[0][inputs["input_ids"].shape[1] :]
res.append(
self._tokenizer.decode(new_tokens, skip_special_tokens=True).strip()
)
self.cache_hook.add_partial("generate_until", (context, gen_kwargs), res[-1])
pbar.update(1)
pbar.close()
return res
Empty file.
26 changes: 26 additions & 0 deletions oellm/resources/task-groups.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -503,6 +503,32 @@ task_groups:
n_shots: [0]
suite: evalchemy

img-understanding:
description: "Image understanding: VQAv2, GQA, TextVQA, ScienceQA, MMMU, MMMU-Pro (standard, 10-option) (all 0-shot)"
suite: lmms-eval
metric: exact_match
tasks:
- task: vqav2_val
n_shots: [0]
dataset: lmms-lab-encoder/VQAv2
- task: gqa
n_shots: [0]
dataset: lmms-lab-encoder/GQA
- task: textvqa_val
n_shots: [0]
dataset: lmms-lab-encoder/textvqa
- task: scienceqa_img
n_shots: [0]
dataset: lmms-lab-encoder/ScienceQA
- task: mmmu_val
n_shots: [0]
dataset: lmms-lab-encoder/MMMU
metric: mmmu_acc
- task: mmmu_pro_standard
n_shots: [0]
dataset: MMMU/MMMU_Pro
metric: mmmu_acc

crows-pairs:
description: "CrowS-Pairs social bias benchmark (English, 0-shot)."
suite: lm-eval-harness
Expand Down
68 changes: 68 additions & 0 deletions oellm/resources/template.sbatch
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,9 @@ TASKS_PER_JOB={tasks_per_job}
export HF_HOME=$HF_HOME
export HF_DATASETS_CACHE="$HF_HOME/datasets"
export HF_HUB_OFFLINE={hf_hub_offline}
export LMMS_EVAL_DATASETS_CACHE="$HF_HOME/datasets"
export HF_XET_CACHE="$HF_HOME/xet_cache"
export TRITON_CACHE_DIR="$HF_HOME/triton_cache"

# Neither harness caches under HF_HOME on its own. lighteval hardcodes
# ~/.cache/huggingface/lighteval and never reads HF_HOME, which fills the user's
Expand Down Expand Up @@ -205,6 +208,7 @@ do
--num_fewshot "$n_shot" \
--output_path "$EVAL_OUT_DIR/$RUN_ID.json" \
--trust_remote_code \
--confirm_run_unsafe_code \
${{LM_EVAL_INCLUDE_PATH:+--include_path $$LM_EVAL_INCLUDE_PATH}} \
--batch_size auto \
${{CACHE_REQUESTS:+--cache_requests $CACHE_REQUESTS}} \
Expand Down Expand Up @@ -286,6 +290,70 @@ do
exit 1
fi
;;
lmms_eval|lmms-eval)
RESULTS_SUBDIR="{evals_dir}/$(openssl rand -hex 5)"
EVAL_OUT_DIR="$RESULTS_SUBDIR"
EVAL_OUT_PAT="*"
if [ -n "$VENV_PATH" ]; then
source "$VENV_PATH/bin/activate"
LMMS_MODEL_TYPE=$(python -c "
import sys
from transformers import AutoConfig, AutoTokenizer
path = sys.argv[1]
result = ''
try:
tok = AutoTokenizer.from_pretrained(path, trust_remote_code=True)
seed2_id = tok.convert_tokens_to_ids('<seed2_0>')
if seed2_id is not None and seed2_id != tok.unk_token_id:
result = 'seed2_omni'
except Exception:
pass
if not result:
try:
result = AutoConfig.from_pretrained(path, trust_remote_code=True).model_type
except Exception:
result = ''
print(result)
" "$model_path" 2>/dev/null)
Comment on lines +299 to +317

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need this?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lmms-eval doesn't have one generic model class like lm_eval --model hf, each VLM family needs its own specific one, and --model has to match exactly. Since schedule only gets a model path/id, not that class name, we need some way to figure out which one to pass. So, we use AutoConfig to extract the model type

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

makes sense can you leave a comment so that it is clear when reading the code?

case "$LMMS_MODEL_TYPE" in
llava|llava_next) LMMS_EVAL_MODEL="llava_hf" ;;
qwen2_5_vl) LMMS_EVAL_MODEL="qwen2_5_vl" ;;
qwen2_vl) LMMS_EVAL_MODEL="qwen2_vl" ;;
seed2_omni) LMMS_EVAL_MODEL="seed2_omni" ;;
*)
echo "[error] lmms-eval suite: no known model class for config.json model_type '$LMMS_MODEL_TYPE' ($model_path). Extend the case in template.sbatch."
EVAL_STATUS=1
LMMS_EVAL_MODEL=""
;;
esac
if [ -n "$LMMS_EVAL_MODEL" ]; then
LMMS_MODEL_ARGS="pretrained=$model_path"
if [ "$LMMS_EVAL_MODEL" = "llava_hf" ]; then
LMMS_MODEL_ARGS="$LMMS_MODEL_ARGS,trust_remote_code=True"
if [ "$GPUS_PER_NODE" -eq 1 ]; then
LMMS_MODEL_ARGS="$LMMS_MODEL_ARGS,device_map=cuda:0"
fi
fi
if [ "$LMMS_EVAL_MODEL" = "seed2_omni" ]; then
export LMMS_EVAL_PLUGINS=mv2_lmms_plugin
fi
accelerate launch --num-processes "$GPUS_PER_NODE" --num-machines 1 \
$MULTI_GPU_FLAG -m lmms_eval eval \
--model "$LMMS_EVAL_MODEL" \
--model_args "$LMMS_MODEL_ARGS" \
--tasks "$task_path" \
--num_fewshot "$n_shot" \
--batch_size 1 \
--output_path "$RESULTS_SUBDIR" \
--log_samples \
${{LIMIT:+--limit $LIMIT}}
EVAL_STATUS=$?
fi
else
echo "[error] lmms-eval suite requires --venv_path (not supported in container mode)."
exit 1
fi
;;
*)
echo "[warning] Unknown evaluation suite '$eval_suite'. Skipping."
EVAL_STATUS=1
Expand Down
15 changes: 15 additions & 0 deletions requirements-venv-lmms-eval.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Dependencies for the lmms-eval suite (image understanding: VQAv2, GQA,
# TextVQA, ScienceQA, MMMU, MMMU-Pro). See docs/VENV.md for setup and the
# platform-specific `decord` workaround this suite needs on linux-aarch64.

lmms-eval==0.7.3
qwen-vl-utils==0.0.14

torch==2.14.0
torchvision==0.29.0
transformers==5.17.0
accelerate==1.15.0
datasets==5.0.1

# seed2_omni's SEED-2 encoder dependency; lmms-eval itself already covers timm/einops/opencv/loguru.
mediapy
Loading