This is a fork of allenai/open-instruct.
This repository is the archived OpenEuroLLM code snapshot for the
TMAX reproduction and transfer study.
It is based on the Open-Instruct tree released in
hamishivi/tmax at 1cce58a2,
which in turn records AllenAI Open-Instruct commit c67b3b50 as its base.
The OpenEuroLLM changes currently published here add the Leonardo Apptainer-based prepared-state backend, related runtime compatibility work, and the exact small Qwen smoke launchers used during bring-up. The Leonardo scripts intentionally retain the paths and scheduler settings from the working environment. They are research snapshots and should be reviewed and adapted before use on another system.
The infrastructure-reproduction stage completed the following:
- implemented a prepared-state Apptainer backend for TMAX environments on Leonardo;
- passed the exact top-200 environment gates;
- passed 270 of 300 tasks in the broader environment diagnostic, with the remaining failures classified;
- completed a real multi-GPU Qwen3.5-2B rollout, one DPPO optimizer update, checkpoint creation, and clean shutdown.
This establishes that the bounded rollout-to-update plumbing executed on the target infrastructure. It does not establish learning or improved task performance.
The study did not complete a nonzero-reward-variance calibration, a meaningful RL pilot, a TMAX performance reproduction, a base-versus-trained capability comparison, or a transfer result. Environment conversion and service portability dominated the work on Leonardo, so the track was stopped. No RL-performance claim is made.
The full final report is recorded in Taskboard issue #351. This repository is retained read-only as a code and infrastructure record.
This public repository contains code only. It does not contain the private experiment-control workspace, job logs, reports, task boards, research notes, prepared caches, datasets, or checkpoints.
It contains fixes on top of upstream for:
- Qwen 3.5 support (fixed hybrid CP-SP training for SFT and RL)
- DPPO Support (new RL loss)
- Terminal agent training (podman-based sandboxes for training)
The training scripts for this fork live under scripts/tmax. Please refer to the README for more details on how to use them. The original scripts target Ai2 infrastructure, while the additional Leonardo scripts exercise the Apptainer prepared-state path. Both sets require adaptation outside their recorded environments. I recommend starting with the smallest relevant smoke script, getting that working, and then scaling up.
For general documentation, usage, and the upstream codebase, refer to the main open-instruct repository. I also recommend checking it for the flags and features.
uvfor dependency management (deps pinned in the repo-rootpyproject.toml/uv.lock).- A Dockerhub login and personal access token (PAT). In particular, you probably need a business account to pull images from Dockerhub at large scale.
This work was supported by the OpenEuroLLM project, co-funded by the Digital Europe Programme under GA no. 101195233.
We acknowledge the EuroHPC Joint Undertaking for awarding this project access to the EuroHPC supercomputer LEONARDO, hosted by CINECA (Italy) and the LEONARDO consortium through a EuroHPC Access call.
The contents presented herein reflect only the authors' view, and the Commission is not responsible for any use that may be made of the information it contains.