Skip to content

[Testing - Not to Land] Add Metal/MLX support with local launcher - #484

Closed
Jack-Khuu wants to merge 723 commits into
gpu-mode:mainfrom
Jack-Khuu:test-mlx
Closed

Jack-Khuu wants to merge 723 commits into
gpu-mode:mainfrom
Jack-Khuu:test-mlx

Conversation

@Jack-Khuu

Copy link
Copy Markdown

Summary

  • Add Metal/MLX support: new MetalGPU enum (M4_Max), LocalLauncher that runs submissions directly on the host machine, and an example MLX vector addition problem
  • Fix macOS compatibility bugs in run_eval.py: handle missing nvidia-smi/rocm-smi//proc/cpuinfo with FileNotFoundError, add MPS/Metal device detection
  • Fix dev leaderboard naming for nested directories (replace / with _ in auto-derived names)

Test plan

  • Start local API server with --api-only and create a dev leaderboard from mlx/example
  • Submit the example MLX kernel for test and benchmark modes via curl
  • Verify run_eval.py system info detection works on macOS without crashing
  • Verify nested directory dev leaderboard names use _ instead of /

ngc92 and others added 30 commits April 27, 2025 23:51
* Remove automatic creation of active-leaderboards channels

* Remove automatic creation of active-leaderboards channels
* Update amd_workflow.yml

* Update amd_workflow.yml

* Update amd_workflow.yml

* Update amd_workflow.yml

* Update amd_workflow.yml

* Update nvidia_workflow.yml
* use workflows timeout

* remove magic number

* lint

* Update nvidia_workflow.yml

* Update nvidia_workflow.yml

* fix verify tests

* make reviewable

* add buffer
* compress github action payload

* allow explicit specification of gh branch

* allow backslashes in code content
* Feat: expect run-id

* Feat: add run_id check in ghrun.trigger

* Remove not needed semaphore

* Update src/discord-cluster-manager/launchers/github.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
…#297)

* [ez] make all connections to the database allow DISABLE_SSL

* lint
* [ez] update references so they run locally

* Apply suggestions from code review

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Mark Saroufim <marksaroufim@meta.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
SinatrasC and others added 21 commits March 13, 2026 06:28
…-mode#459)

* Add HF dataset export logic and admin commands

* Add huggingface-hub and datasets dependencies

* Add tests for HF export and admin API endpoint

* apply review fixes, req hardening and in memory buffer addition

* switch to standalone sql file

* switch to tempfile, hf xet support is not available over inmemorybuffer

---------

Co-authored-by: Sinatras <SinatrasC@users.noreply.github.com>
* Update runs

* Cleanup

* Defaults
Update hf export data timings because of frequent deployments task is missed each day with 24 hours loops
Document the issue where failed ephemeral runners consume maxRunners
slots without being garbage-collected, causing jobs to queue forever.
Includes diagnosis commands and the one-liner fix.
* Add user ban system for blocking submissions

Adds is_banned column to user_info, ban check in prepare_submission
(blocks all entry points: Discord, CLI, Web), admin Discord commands
(/admin ban, /admin unban), and admin API endpoints (POST/DELETE
/admin/ban/{user_id}).

* Fix test mock to set is_user_banned=False for submission tests
Previously minRunners was 0 (scale-to-zero), meaning runners only
appeared on the GitHub runners tab when jobs were queued. Set to 40
so all 5 nodes × 8 GPUs stay online and ready.
* Add KernelGuard integration for submission pre-checks

- Introduced KernelGuard for validating submissions before processing.
- Implemented error handling for rejected submissions in the backend.
- Updated database methods to mark submissions as hacked when flagged.
- Enhanced tests to cover new KernelGuard functionality and error scenarios.
- Added a new kernelguard.py module for managing submission analysis and pre-checks.

* Update Python version requirement and enhance KernelGuard integration

- Updated Python version requirement from 3.10 to 3.11 in pyproject.toml and uv.lock.
- Added `kernelguard` dependency to manage submission pre-checks.
- Enhanced error handling in submission processes to include KernelGuard rejection scenarios.
- Implemented pre-check logic in the submission workflow to prevent blocked submissions from queuing.
- Updated tests to reflect changes in submission handling and pre-check logic.

* ruff fix

---------

Co-authored-by: Sinatras <SinatrasC@users.noreply.github.com>
…pu-mode#464)

Bumps the uv group with 1 update in the / directory: [orjson](https://github.com/ijl/orjson).


Updates `orjson` from 3.11.5 to 3.11.6
- [Release notes](https://github.com/ijl/orjson/releases)
- [Changelog](https://github.com/ijl/orjson/blob/master/CHANGELOG.md)
- [Commits](ijl/orjson@3.11.5...3.11.6)

---
updated-dependencies:
- dependency-name: orjson
  dependency-version: 3.11.6
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
---
updated-dependencies:
- dependency-name: aiohttp
  dependency-version: 3.13.4
  dependency-type: direct:development
  dependency-group: uv
- dependency-name: requests
  dependency-version: 2.33.0
  dependency-type: direct:development
  dependency-group: uv
- dependency-name: cryptography
  dependency-version: 46.0.6
  dependency-type: indirect
  dependency-group: uv
- dependency-name: ujson
  dependency-version: 5.12.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps the uv group with 1 update in the / directory: [cryptography](https://github.com/pyca/cryptography).


Updates `cryptography` from 46.0.6 to 46.0.7
- [Changelog](https://github.com/pyca/cryptography/blob/main/CHANGELOG.rst)
- [Commits](pyca/cryptography@46.0.6...46.0.7)

---
updated-dependencies:
- dependency-name: cryptography
  dependency-version: 46.0.7
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…pu-mode#479)

Bumps the uv group with 1 update in the / directory: [pytest](https://github.com/pytest-dev/pytest).


Updates `pytest` from 8.4.1 to 9.0.3
- [Release notes](https://github.com/pytest-dev/pytest/releases)
- [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst)
- [Commits](pytest-dev/pytest@8.4.1...9.0.3)

---
updated-dependencies:
- dependency-name: pytest
  dependency-version: 9.0.3
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps the uv group with 1 update in the / directory: [python-multipart](https://github.com/Kludex/python-multipart).


Updates `python-multipart` from 0.0.22 to 0.0.26
- [Release notes](https://github.com/Kludex/python-multipart/releases)
- [Changelog](https://github.com/Kludex/python-multipart/blob/master/CHANGELOG.md)
- [Commits](Kludex/python-multipart@0.0.22...0.0.26)

---
updated-dependencies:
- dependency-name: python-multipart
  dependency-version: 0.0.26
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps the uv group with 1 update in the / directory: [python-dotenv](https://github.com/theskumar/python-dotenv).


Updates `python-dotenv` from 1.1.1 to 1.2.2
- [Release notes](https://github.com/theskumar/python-dotenv/releases)
- [Changelog](https://github.com/theskumar/python-dotenv/blob/main/CHANGELOG.md)
- [Commits](theskumar/python-dotenv@v1.1.1...v1.2.2)

---
updated-dependencies:
- dependency-name: python-dotenv
  dependency-version: 1.2.2
  dependency-type: direct:development
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Copilot AI review requested due to automatic review settings May 1, 2026 03:04

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds first-class support for running MLX/Metal-based problems locally (via a new Local launcher and a Metal GPU enum), improves macOS compatibility for system info detection, and fixes dev leaderboard naming for nested problem directories.

Changes:

  • Add MetalGPU (M4_Max) and a LocalLauncher that executes evaluations directly on the host.
  • Make make_system_info() resilient on macOS (MPS detection + tolerate missing nvidia-smi/rocm-smi and /proc/cpuinfo).
  • Sanitize dev leaderboard names derived from nested directories (/ → _) and add an MLX example problem.

Reviewed changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated 6 comments.

Show a summary per file
File Description
src/libkernelbot/run_eval.py Adds MPS detection and macOS-safe fallbacks for missing system tools/files.
src/libkernelbot/launchers/local.py Introduces LocalLauncher to run configs directly on the host machine.
src/libkernelbot/launchers/__init__.py Exports LocalLauncher from the launchers package.
src/libkernelbot/consts.py Adds MetalGPU and registers it in the GPU lookup + SM map.
src/kernelbot/main.py Registers LocalLauncher in the backend launcher set.
src/kernelbot/cogs/admin_cog.py Adds Metal GPUs to Discord admin GPU selection flows.
src/kernelbot/api/main.py Fixes dev leaderboard auto-name derivation for nested directories.
src/envs.txt Adds local environment export helper (currently machine-specific).
instructions.txt Adds manual setup/testing instructions for local MLX runs.
examples/mlx/example/task.yml Defines an MLX vector-add example task runnable on M4_Max.
examples/mlx/example/submission.py Sample MLX submission (baseline vector add).
examples/mlx/example/reference.py Reference implementation + correctness checking for MLX example.
examples/mlx/example/eval.py Evaluation/benchmark harness for the MLX example task.
examples/mlx.yaml Adds an example “problem set” YAML referencing the MLX task.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

from .modal import ModalLauncher

__all__ = [Launcher, GitHubLauncher, ModalLauncher]
__all__ = [Launcher, GitHubLauncher, LocalLauncher, ModalLauncher]
Comment thread src/envs.txt
@@ -0,0 +1,3 @@
export DATABASE_URL="postgresql://$(whoami)@localhost:5432/kernelbot"
export ADMIN_TOKEN="your-admin-token"
export PROBLEM_DEV_DIR="/Users/jackkhuu/Desktop/oss/reference-kernels/problems"
Comment thread src/kernelbot/api/main.py
Comment on lines +647 to 648
leaderboard_name = f"{directory.replace('/', '_')}-dev"
deadline_value = datetime.datetime.now(datetime.timezone.utc) + datetime.timedelta(days=365)
Comment on lines +13 to +30
class LocalLauncher(Launcher):
def __init__(self):
super().__init__("Local", gpus=MetalGPU)

async def run_submission(
self, config: dict, gpu_type: GPU, status: RunProgressReporter
) -> FullResult:
if config["lang"] == "cu":
raise NotImplementedError("CUDA is not supported on Metal GPUs")

logger.info(f"Starting local run for {gpu_type.name}")
await status.push(f"⏳ Running locally on {gpu_type.name}...")

loop = asyncio.get_event_loop()
result = await loop.run_in_executor(None, lambda: run_config(config))

await status.update(f"✅ Local run on {gpu_type.name} complete")
return result
Comment thread src/kernelbot/main.py
Comment on lines 28 to 33
backend.register_launcher(ModalLauncher(consts.MODAL_CUDA_INCLUDE_DIRS))
backend.register_launcher(
GitHubLauncher(env.GITHUB_REPO, env.GITHUB_TOKEN, env.GITHUB_WORKFLOW_BRANCH)
)
backend.register_launcher(LocalLauncher())
return backend
self, config: dict, gpu_type: GPU, status: RunProgressReporter
) -> FullResult:
if config["lang"] == "cu":
raise NotImplementedError("CUDA is not supported on Metal GPUs")
@Jack-Khuu Jack-Khuu changed the title Add Metal/MLX support with local launcher [Testing - Not to Land] Add Metal/MLX support with local launcher May 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.