Diagnose slow PyTorch training with zero-code instrumentation. Catch regressions in CI.
Works automatically with: Hugging Face Trainer · PyTorch Lightning · RF-DETR
Quickstart · What you get · Compare runs · Regression checks · Integrations · Documentation
⭐ If TraceML helps you find a bottleneck, please star the repository.
TraceML shows where each training step goes—input loading, data transfer, forward, backward, and optimizer work—then identifies the bottleneck and saves evidence for local comparison or CI.
If your training framework is already installed:
pip install traceml-ai- python train.py
+ traceml run train.pyFor standard Hugging Face Trainer, PyTorch Lightning, and RF-DETR training,
TraceML instruments the run without changes to the training script. It prints a
diagnosis when training finishes and writes final_summary.json and
final_summary.txt under logs/<run-name>/.
No training script ready? Try the interactive demo or Colab example.
The terminal report identifies the bottleneck, shows the timing and resource evidence behind it, and suggests the next investigation.
+----------------------------------------------------------------------------------------------------------------------------------------------------------+
| TraceML Run Summary |
| bert_finetune · 1 rank · 1 GPU observed · 256 common steps · 52.4s |
+----------------------------------------------------------------------------------------------------------------------------------------------------------+
| |
| Verdict: INPUT-BOUND (CRITICAL) |
| Why: Input Wait took 64% of Step Time. |
| Next: Increase workers, prefetch, or storage throughput. |
| |
| STEP TIMING (Window Average), GPU Clock || STEP MEMORY: BALANCED |
| Step Time 200.4 ms 100% || |
| ├─ Input Wait 128.0 ms 64% ◀ cause || |
| ├─ Compute 68.0 ms 34% || avg per-step peak avg |
| │ ├─ Forward 24.0 ms 12% || Allocated 2.9 GB |
| │ ├─ Backward 38.0 ms 19% || Reserved 3.2 GB |
| │ └─ Optimizer 6.0 ms 3% || |
| ├─ H2D 0.4 ms <1% || |
| └─ Residual 3.6 ms 2% || |
| DataLoader fetch: 120.0 ms (CPU, supplemental) || |
| |
| SYSTEM METRICS: LOW GPU UTIL || PROCESS METRICS: NORMAL |
| Evidence: GPU utilization averaged 24%. || |
| || |
| avg || avg |
| CPU 18% || CPU capacity 14% |
| RAM used 6.2 GB (19%) || RSS used 3.1 GB (10%) |
| GPU util 24% || CUDA allocated 2.9 GB |
| GPU memory/device 3.3 GB (21%) || CUDA reserved 3.2 GB (20%) |
| GPU temperature 42C || |
| GPU power 58W || |
| |
| |
| Full evidence: logs/bert_finetune/final_summary.json (--html-report) |
+----------------------------------------------------------------------------------------------------------------------------------------------------------+
Training with multiple ranks? See a rank-straggler diagnosis
+----------------------------------------------------------------------------------------------------------------------------------------------------------+
| TraceML Run Summary |
| ddp_pretrain · 4/4 ranks · 4 GPUs observed · 2/2 nodes · 250 common steps · 40.1s |
+----------------------------------------------------------------------------------------------------------------------------------------------------------+
| |
| Verdict: INPUT STRAGGLER (CRITICAL) |
| Why: R0/N0 waited 254.5 ms for input; R1/N0 waited 3.8 ms for input. |
| Next: Inspect input wait on the slow rank. |
| Scope: N = node · R = global rank · G = GPU index |
| |
| STEP TIMING (Median R1/N0), GPU Clock || STEP MEMORY: BALANCED · 4/4 ranks |
| Step Time 303.7 ms 100% || |
| ├─ Input Wait 3.8 ms 1% || |
| ├─ Compute 259.5 ms 85% || avg per-step peak median rank avg worst rank avg |
| │ ├─ Forward 80.0 ms 26% || Allocated 8.5 GB 9.4 GB, R2/N1 |
| │ ├─ Backward 169.5 ms 56% || Reserved 8.9 GB 9.8 GB, R2/N1 |
| │ └─ Optimizer 10.0 ms 3% || |
| ├─ H2D 1.1 ms <1% || |
| └─ Residual 39.3 ms 13% || |
| DataLoader fetch: 3.7 ms (CPU, supplemental) || |
| |
| SYSTEM METRICS: LOW GPU UTIL · 2/2 nodes || PROCESS METRICS: NORMAL · 4/4 ranks |
| Evidence: GPU utilization averaged 14%. || |
| || |
| median node avg worst node avg || median rank avg worst rank avg |
| CPU 18% 26%, N1 || CPU capacity 12% 81%, R2/N1 |
| RAM used 16.0 GB (27%) 20.8 GB (35%), N1 || RSS used 3.1 GB (10%) 5.4 GB (17%), R1/N0 |
| GPU util 9% 9%, N1 || CUDA allocated 2.9 GB 4.6 GB, R3/N1 |
| GPU memory/device 5.0 GB (31%) 7.0 GB (44%), N1 || CUDA reserved 3.2 GB (20%) 6.8 GB (43%), R3/N1 |
| GPU temperature 58C 70C, N1 || |
| GPU power 220W 280W, N1 || |
| |
| |
| Full evidence: logs/ddp_pretrain/final_summary.json (--html-report) |
+----------------------------------------------------------------------------------------------------------------------------------------------------------+
TraceML is designed for lightweight, always-on training diagnosis. It shows enough evidence to choose the next investigation; use a kernel profiler when the result points inside GPU compute.
| Diagnosis | Where to investigate |
|---|---|
| Input-bound | DataLoader workers, transforms, tokenization, collation, or storage |
| H2D-bound | Pinned memory, non-blocking copies, batch size, or transfer overlap |
| Compute-bound | Model compute, mixed precision, batch size, or deeper profiling |
| Residual-heavy | Logging, checkpointing, validation, CPU stalls, or unobserved work |
| Rank straggler | Rank-local input, data imbalance, node variance, or networking |
| Memory creep | Retained tensors, logging references, or cached activations |
Read How to Read TraceML Output for definitions, diagnosis rules, and evidence limits.
After changing the DataLoader, batch size, model, or infrastructure, compare two completed runs:
traceml compare before/final_summary.json after/final_summary.json+--------------------------------------------------------------------------------------+
| TraceML Compare |
+--------------------------------------------------------------------------------------+
| A: before_dataloader_fix |
| B: after_dataloader_fix |
| Primary diagnosis: INPUT-BOUND -> COMPUTE-BOUND (changed) |
| |
| Verdict: IMPROVEMENT |
| Why: GPU Step Time decreased by 59.9%. |
+--------------------------------------------------------------------------------------+
TraceML keeps the comparison evidence in JSON and text so the result is easy to review, archive, or use in automation. See Compare Runs.
Use the same comparison as a local or CI gate for compatible runs:
traceml compare \
logs/reference/final_summary.json \
logs/candidate/final_summary.json \
--max-step-time-regression-pct 5 \
--output compare/reference-vs-candidateTraceML writes the evidence before returning the CI exit code. The regression guard is an experimental pilot; read the Regression Guard for comparability requirements and supported environments.
Launch one TraceML process per training process:
- torchrun --nproc-per-node=4 train.py
+ traceml run train.py --nproc-per-node=4The final summary aligns common steps across ranks and can identify the worker most likely to be holding up the run. See Distributed Training, DDP rank stragglers, and Slurm.
TraceML is currently focused on single-device training and DDP on one machine. You can launch multi-node runs, but that path remains experimental. Distributed GPU comparisons assume homogeneous hardware across ranks. See the integration support matrix for framework-specific details, including FSDP.
| Investigation | Result |
|---|---|
| ResNet-18 input pipeline | See an input-bound run become compute-bound after changing only DataLoader settings. |
| RF-DETR Nano training | Examine single-GPU phase timing and four-GPU DDP scaling on real COCO batches. |
| RF-DETR release regression | Attribute a release-to-release slowdown to input waiting and verify the fixed release. |
The zero-code command works with these standard trainer APIs:
| Framework | Training API |
|---|---|
| Hugging Face Trainer | trainer.train() |
| PyTorch Lightning | trainer.fit(...) |
| RF-DETR | model.train(...) |
Custom loops and other training paths use their existing TraceML integration:
| Training path | Setup |
|---|---|
| Plain PyTorch or a custom loop | traceml.init() and trace_step(...) |
| Hugging Face Accelerate | Explicit step instrumentation |
| Ray Train and Ray Data | TraceML Trainer/config wrappers |
| DeepSpeed | Explicit step instrumentation |
| MONAI | TraceML handler setup |
Plain PyTorch example
import traceml_ai as traceml
traceml.init(mode="auto")
for batch in dataloader:
with traceml.trace_step(model):
optimizer.zero_grad(set_to_none=True)
outputs = model(batch["x"])
loss = criterion(outputs, batch["y"])
loss.backward()
optimizer.step()Run it with traceml run train.py.
Manual APIs remain available for custom loops and advanced timing boundaries. See the Public API and integration support matrix.
Summary mode is the default. For live diagnostics in the terminal, use:
traceml run train.py --mode=cliFor the browser dashboard, install the optional dashboard dependencies:
pip install "traceml-ai[dashboard]"
traceml run train.py --mode=dashboardTraceML can also export a self-contained HTML report and send its compact summary to an existing W&B or MLflow run. See W&B and MLflow and the complete quickstart.
Questions, contributions, and real-world slowdown reports are welcome:
Apache 2.0. See LICENSE.