Repository navigation
Add img-understanding task group (lmms-eval) #125
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
harshraj172
wants to merge
11
commits into
main
Choose a base branch
from
harsh/vqa-task-suite
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
11 commits
Select commit
Hold shift + click to select a range
740b46e
Add vqa and mmmu task groups via lmms-eval
harshraj172 7b1a5a4
Fix bugs found by live-testing the lmms-eval branch on real SLURM jobs
harshraj172 45060dc
Merge vqa and mmmu into one img-understanding task group
harshraj172 0da3b34
Trim comments down to one line each
harshraj172 3241d6e
Remove comments added for img-understanding
harshraj172 595f458
Tighten the img-understanding VENV.md section
harshraj172 9c9d0cd
Pass trust_remote_code and n_shot through to lmms_eval, fix HF_TOKEN …
harshraj172 6133c13
Only pass trust_remote_code to llava_hf, not the Qwen VL model classes
harshraj172 c6d7fb7
Skip directories matching *.json in collect_results
harshraj172 ba1aafe
Add seed2_omni model class for mixturevitae2's omni checkpoints
harshraj172 205f40b
Merge branch 'main' into harsh/vqa-task-suite
haideraltahan File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Empty file.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
| AVAILABLE_MODELS = {"seed2_omni": "Seed2Omni"} |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,96 @@ | ||
| import os | ||
| import sys | ||
|
|
||
| import torch | ||
| from lmms_eval.api.model import lmms | ||
| from lmms_eval.api.registry import register_model | ||
| from tqdm import tqdm | ||
|
|
||
|
|
||
| @register_model("seed2_omni") | ||
| class Seed2Omni(lmms): | ||
| def __init__(self, pretrained, device="cuda", batch_size=1, **kwargs): | ||
| super().__init__() | ||
| assert kwargs == {}, f"Unexpected kwargs: {kwargs}" | ||
| assert int(batch_size) == 1, "seed2_omni only supports batch_size=1" | ||
|
|
||
| mv2_dir = os.environ.get( | ||
| "MV2_MULTIMODAL_DIR", "/e/project1/jureap59/raj3/mixturevitae2" | ||
| ) | ||
| if mv2_dir not in sys.path: | ||
| sys.path.insert(0, mv2_dir) | ||
| from multimodal_processing.model_backends.seed2 import Backend, seed2_block | ||
|
|
||
| self._seed2_block = seed2_block | ||
| self._seed2 = Backend() | ||
| self._seed2.load() | ||
|
|
||
| from transformers import AutoModelForCausalLM, AutoTokenizer | ||
|
|
||
| self._device = torch.device(device) | ||
| self._tokenizer = AutoTokenizer.from_pretrained( | ||
| pretrained, trust_remote_code=True | ||
| ) | ||
| self._model = ( | ||
| AutoModelForCausalLM.from_pretrained( | ||
| pretrained, dtype=torch.bfloat16, trust_remote_code=True | ||
| ) | ||
| .to(self._device) | ||
| .eval() | ||
| ) | ||
| im_end_id = self._tokenizer.convert_tokens_to_ids("<|im_end|>") | ||
| self._eos_token_ids = [self._tokenizer.eos_token_id, im_end_id] | ||
| self.batch_size_per_gpu = 1 | ||
|
|
||
| @property | ||
| def batch_size(self): | ||
| return self.batch_size_per_gpu | ||
|
|
||
| @property | ||
| def device(self): | ||
| return self._device | ||
|
|
||
| def loglikelihood(self, requests): | ||
| raise NotImplementedError("seed2_omni only supports generate_until tasks") | ||
|
|
||
| def generate_until_multi_round(self, requests): | ||
| raise NotImplementedError("seed2_omni does not support multi-round generation") | ||
|
|
||
| def generate_until(self, requests): | ||
| res = [] | ||
| pbar = tqdm( | ||
| total=len(requests), disable=(self.rank != 0), desc="seed2_omni responding" | ||
| ) | ||
| for request in requests: | ||
| context, gen_kwargs, doc_to_visual, doc_id, task, split = request.args | ||
| visuals = doc_to_visual(self.task_dict[task][split][doc_id]) | ||
| if not isinstance(visuals, list): | ||
| visuals = [visuals] | ||
|
|
||
| blocks = [] | ||
| for v in visuals: | ||
| ids = self._seed2.encode(v.convert("RGB")) | ||
| if ids: | ||
| blocks.append(self._seed2_block(ids)) | ||
| prompt = f"{''.join(blocks)} {context}".strip() if blocks else context | ||
| chat = f"<|im_start|>user\n{prompt}<|im_end|>\n<|im_start|>assistant\n<think>\n</think>\n" | ||
|
|
||
| inputs = self._tokenizer(chat, return_tensors="pt").to(self._device) | ||
| do_sample = gen_kwargs.get("temperature", 0) > 0 | ||
| with torch.no_grad(): | ||
| out = self._model.generate( | ||
| **inputs, | ||
| max_new_tokens=gen_kwargs.get("max_new_tokens", 256), | ||
| do_sample=do_sample, | ||
| temperature=gen_kwargs.get("temperature") if do_sample else None, | ||
| pad_token_id=self._tokenizer.eos_token_id, | ||
| eos_token_id=self._eos_token_ids, | ||
| ) | ||
| new_tokens = out[0][inputs["input_ids"].shape[1] :] | ||
| res.append( | ||
| self._tokenizer.decode(new_tokens, skip_special_tokens=True).strip() | ||
| ) | ||
| self.cache_hook.add_partial("generate_until", (context, gen_kwargs), res[-1]) | ||
| pbar.update(1) | ||
| pbar.close() | ||
| return res |
Empty file.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,15 @@ | ||
| # Dependencies for the lmms-eval suite (image understanding: VQAv2, GQA, | ||
| # TextVQA, ScienceQA, MMMU, MMMU-Pro). See docs/VENV.md for setup and the | ||
| # platform-specific `decord` workaround this suite needs on linux-aarch64. | ||
|
|
||
| lmms-eval==0.7.3 | ||
| qwen-vl-utils==0.0.14 | ||
|
|
||
| torch==2.14.0 | ||
| torchvision==0.29.0 | ||
| transformers==5.17.0 | ||
| accelerate==1.15.0 | ||
| datasets==5.0.1 | ||
|
|
||
| # seed2_omni's SEED-2 encoder dependency; lmms-eval itself already covers timm/einops/opencv/loguru. | ||
| mediapy |
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
why do we need this?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
lmms-eval doesn't have one generic model class like lm_eval
--model hf, each VLM family needs its own specific one, and--modelhas to match exactly. Sincescheduleonly gets a modelpath/id, not that class name, we need some way to figure out which one to pass. So, we useAutoConfigto extract the model typeThere was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
makes sense can you leave a comment so that it is clear when reading the code?