This repository contains a research compiler for the H16G Apple Neural Engine in the M4. It reads textual MIL, builds a typed graph, lowers supported operations, and writes new HWX objects without calling Apple's compiler.
The project is a canary for the compiler pipeline recovered in Inside the M4 Apple Neural Engine, Part 4b. It shows which parts of that pipeline are understood well enough to reproduce in code and verify on hardware.
- An Apple silicon Mac. Hardware results in this repository were measured on an
M4 (
Mac16,10). - macOS and the Xcode Command Line Tools, which provide
clang++,make, the Foundation SDK, and the IOSurface SDK. - macOS 26.3 build
25D125for the recorded H16G hardware results. Other releases may use different private interfaces or descriptor layouts. - Administrator access for hardware tests. Compilation and software tests do
not require
sudo.
There are no third-party package dependencies.
Install the command-line tools if needed:
xcode-select --installClone the repository and run the software test suite:
git clone https://github.com/maderix/mil-hwx-compiler.git
cd mil-hwx-compiler
make test -j4Build only the compiler:
make build/mil-hwxc -j4Compile the included Conv1x1 and ReLU fixture:
./build/mil-hwxc \
--mil tests/fixtures/conv_relu.mil \
--model-root tests/models/conv_relu \
--target H16G \
--output build/conv-bundleThe output directory contains one or more program-N.hwx files and a
manifest.json file. The manifest records dispatch order, tensor bindings,
physical strides, and shared intermediate surfaces.
To compile and run the Chunked DeltaNet block on an M4:
sudo -v
bash tests/run_chunked_deltanet_hardware.shThe script builds the compiler and runner, emits 58 HWX programs, provisions
them in the macOS aned cache, runs two numerical cases, and prints a warm
latency measurement. The script exits with a failure if compilation,
provisioning, execution, or either output comparison fails.
To compile and run the FP16 attention graph:
sudo -v
bash tests/run_online_reduction_hardware.shThis emits three programs and checks all 16,384 output elements against the CPU reference. To compare it with the output of Apple's compiler on the same MIL, inputs, and synchronous runtime path:
ANE_BENCHMARK_WARMUP=50 \
ANE_BENCHMARK_ITERATIONS=5000 \
ANE_BENCHMARK_BATCHES=5 \
bash tests/run_fa2_ab_hardware.shThe repository provides:
- A compiler path from textual MIL to fresh H16G HWX objects.
- Hardware tests with independent CPU references for every supported family.
- Encoders for the decoded HWX container and Task Descriptor fields.
- A small runtime that loads provisioned objects and submits them through
_ANEClient.
Coverage is limited to the operation and shape rows listed below. The target tables were measured on one M4 and one macOS build. This compiler is not a Core ML or MLX replacement, and the private runtime is unsuitable for App Store software.
MIL source
-> lexer and parser
-> typed SSA graph
-> normalization and decomposition
-> structural fusion
-> H16G legality and numeric-mode selection
-> tiling, liveness, SRAM, and DMA planning
-> H16G task encoding and program composition
-> HWX object and binding manifest
The production path does not load an Apple-compiled HWX file, choose code from
a fixture name, or patch an existing container. HWXObjectWriter creates each
object from compiler data structures. The compiler reparses the completed
object before accepting it.
Multi-operation graphs are partitioned according to target capability tables. Compatible adjacent tasks can share one HWX program. Other tasks remain separate programs connected through manifest-managed IOSurfaces. Unsupported operations, shapes, data types, and axes fail compilation. A transition that cannot be composed stays on the standalone program path.
Composition has two forms. Simple elementwise operations can be folded into a field of an adjacent task. Longer chains use operation-specific SRAM input and output forms so an intermediate remains inside one program. The planner selects these forms from operation, shape, data type, bridge state, and value lifetime. It does not select them from a model or function name.
| Family | Measured H16G coverage |
|---|---|
| Conv1x1 | 16 FP16 channel and spatial geometries, with optional ReLU |
| Regular convolution | C64/C128, S32/S64, K3/K5 |
| Depthwise convolution | C64/C128/C256/C512, S64, K3 |
| Matmul | Square N128/N256/N512 and tiled multiples of 128 through N4096 |
| Binary ALU | Add at N128 through N2048; multiply at N128/N512; max/min at N512 |
| Unary and LUT | ReLU, sigmoid, tanh, GELU, SiLU, exp, log, sqrt, rsqrt, reciprocal |
| Reduction | Sum, mean, and max over measured channel, height, and width axes |
| Layout | S2D and D2S B4/B8, including 64-byte physical row padding |
| Fused layout | S2D, Conv1x1, and D2S at natural C8/C16/C24/C32 |
| W8A8 | Four-layer C64/S64 Conv1x1 chain with packed middle blocks |
| Attention | H4/S64/D64 decomposition; unmasked FP16 S128/D128 forward graph in three programs |
| State updates | Four-step N128 affine scan; FP16 Chunked DeltaNet block at C128/D128 |
The Chunked DeltaNet fixture uses ordinary matmul, add, multiply, and exp operations. The caller supplies normalized Q and K tensors, the transposed K layout, and the fixed triangular matrices. No DeltaNet operation name reaches the planner or H16G encoders.
| Graph | HWX programs | CPU-reference result | Research latency | Apple compiler | ANE power capture |
|---|---|---|---|---|---|
| FP16 attention, S128/D128 | 3 | 16,384/16,384 elements passed; maximum error 0 |
369.458 us |
127.208 us |
61/140 active; 601.9 mW active average; 681 mW peak |
| Matmul, reshape, GELU, N256 | 1 | Maximum error 0.00610352 |
66.94 us |
76.60 us |
8/60 active; 36-134 mW; GPU 0 mW |
| Four-step FP16 affine scan, N128 | 8 | Every stage within 1 FP16 ULP; final output within 3 ULP | Not measured | InvalidMILProgram |
60/60 active; 37-152 mW |
| Chunked DeltaNet block, C128/D128 | 58 | Output relative L2 0.004214; final-state relative L2 0.004443; maximum error 3.69e-05 |
22728.792 us |
InvalidMILProgram |
73/80 active; 19-403 mW; 86.34 mW average; GPU 0 mW |
The attention graph is split at its two external-surface boundaries. Program 0 runs matmul and scale. Program 1 runs reduce-max, subtraction, exponentiation, reduce-sum, reciprocal, and multiply while keeping its intermediate values in SRAM. Program 2 runs the final matmul.
The attention A/B used 50 warmups and five alternating batches of 5,000 evaluations per compiler, for 25,000 measured samples each. Both implementations used the same MIL, inputs, IOSurfaces, QoS, and synchronous completion boundary. The research compiler took 2.986 times the Apple compiler median. A separate research-only profile measured median submission times of 114.458, 110.583, and 113.146 microseconds for the three programs. The median complete chain was 369.417 microseconds without profiling instrumentation.
The matmul-GELU values are medians from two earlier runs. Each run used 20 warmups and five alternating batches of 2,000 evaluations per compiler.
The Chunked DeltaNet result used 10 warmups and five batches of 50 evaluations. Its 58-program schedule is a correctness result and currently carries substantial dispatch and intermediate-surface cost. Apple's compiler rejected the same 14-input MIL program, so an equivalent latency comparison is not available.
Powermetrics sampled the system ANE and GPU rails every 100 ms while the hardware tests ran. The attention power capture ran only research-generated programs. These values confirm ANE activity during the tests. They are system-wide estimates and do not measure per-process energy.
The current attention result is recorded in a benchmark and profile receipt and a powermetrics screenshot. Earlier compiler A/B results are in the original baseline and the first optimization receipt. The current Chunked DeltaNet run is recorded in docs/evidence/chunked-deltanet-m4-2026-09-03.txt.
Each hardware script compiles its MIL fixture, provisions the emitted object, runs it on the M4 ANE, and checks the output. Useful entry points include:
bash tests/run_m4_hardware.sh
bash tests/run_matmul_hardware.sh
bash tests/run_unary_hardware.sh
bash tests/run_reduce_hardware.sh
bash tests/run_layout_hardware.sh
bash tests/run_online_reduction_hardware.sh
bash tests/run_online_reduction_fallback_hardware.sh
bash tests/run_affine_scan_hardware.sh
bash tests/run_matmul_gelu_hardware.sh
bash tests/run_chunked_deltanet_hardware.sh
bash tests/run_compiler_ab_hardware.shStock macOS loads these objects from
/Library/Caches/com.apple.aned/<build>/InMemoryModelCache/<executable>/.
The scripts use sudo -n to create that cache entry and install the generated
HWX file. Run sudo -v first if the current shell has no valid credential
timestamp. The software test suite writes only to build/.
ANEProvisionedRuntime creates IOSurfaces from the binding manifest and calls
the private AppleNeuralEngine.framework runtime. The manifest keeps logical
tensor sizes separate from physical row, plane, batch, and allocation sizes.
This is required for narrow tensors whose rows are padded to 64 bytes.
The compiler does not provide a kernel-driver path or bypass the normal aned
cache requirement. Private interfaces and accepted HWX layouts can change with
macOS releases.
The HWX container, Task Descriptor fields, and compiler stages were recovered by compiling author-written MIL with Apple's compiler, comparing generated objects, tracing compiler execution, and testing edited objects on hardware. No Apple source code was available or used, and no Apple binary code is distributed here.
Some plugins/H16G/Encoding/*EncoderData.inc files contain measured Task
Descriptor words for operation and geometry rows whose field grammar is still
incomplete. They are indexed target measurements rather than copied program
containers. Tests store hashes and decoded field values, not Apple-generated
HWX files.
See DISCLAIMER.md for the full scope and private-API notes.
The original code and documentation in this repository are available under the MIT License.