Repository navigation
Cache both harnesses under HF_HOME - #120
Merged
Merged
Conversation
geoalgo
approved these changes
Sep 22, 2026
lighteval hardcodes ~/.cache/huggingface/lighteval and never reads HF_HOME, so every run fills the user's home directory. On LUMI that ends as "OSError: [Errno 122] Disk quota exceeded" partway through an evaluation, which the job then reports as finished. lm-eval keeps request caches inside its own install directory, which is a read-only filesystem in the container. Export LIGHTEVAL_CACHE_DIR and LM_HARNESS_CACHE_PATH under HF_HOME, pass lighteval its cache_dir as a model argument (the only way it honours one), and forward LM_HARNESS_CACHE_PATH to clusters that run the container with --contain. lighteval caches unconditionally. lm-eval's --cache_requests stays opt-in behind CACHE_REQUESTS, because it makes lm-eval ignore --limit while building contexts.
haideraltahan
force-pushed
the
harness-caches-under-hf-home
branch
from
September 22, 2026 15:21
0c13b47 to
b6e4299
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Neither harness caches under
HF_HOMEon its own.lighteval hardcodes
cache_dir = "~/.cache/huggingface/lighteval"and never readsHF_HOME, so every run writes into the user's home. It caches unconditionally —TransformersModel.__init__constructs aSampleCachewith no flag guarding it — so on LUMI (20 GB home) a run dies during model init:That cost one of four evaluations in a run I would otherwise have recorded as complete.
lm-eval keeps request caches in its own install directory, which is read-only in the container:
The change
lighteval only honours
cache_diras a model argument, so it is passed there in both the venv and container branches.LM_HARNESS_CACHE_PATHis read from the environment, so it is added toSINGULARITY_ENV_ARGSfor the clusters that run with--contain(jupiter, snellius), which forward only the variables named there.Verified on dev-g — lighteval now hits the shared path:
cache_direxists on lighteval'sModelConfigin both 0.13.0 and 0.13.1.dev0.lm-eval's
--cache_requestsstays opt-inIt is wired up behind
CACHE_REQUESTS(trueto enable,refreshto rebuild) rather than on by default, because enabling it makes limited runs several times slower. Fromlm_eval/api/task.py:--limitis ignored while building contexts, so a smoke test builds every document in the task. Same two tasks, same venv, same--limit 4, on dev-g:--cache_requestsThe warm run is no faster than the cold one, and both are 3–4x the uncached baseline. The cache does work — the entry written by run 1 was not rewritten by run 2, which is the hit path — it just does not pay for itself here, and
--limitis what the docs recommend for dev-g.Two things to know if it is ever turned on by default:
requests-{task}-{shots}shot-rank{r}-world_size{w}-tokenizer{name}. It does not cover the task YAML, so editing a task definition silently reuses the old prompts — henceCACHE_REQUESTS=refresh.tokenizer_nameistokenizer.name_or_path, which forrepo,revision=…is just the repo id, so every revision of a checkpoint repo shares one entry. Fine while the tokenizer is identical across revisions.Tests
tests/test_harness_cache_paths.py— both paths exported underHF_HOME,cache_dirpassed to lighteval in both branches,LM_HARNESS_CACHE_PATHforwarded into contained containers, and--cache_requestsabsent unlessCACHE_REQUESTSis set. The first three fail onmain. Suite passes (140),ruff checkandformatclean.Merged
mainto pick up #110; the batching and caching changes touch the same invocation and both survive (test_task_batching.pypasses).