Continuous Tracing with GPU Monitoring | Datadog

Continuous Tracing with GPU Monitoring

About This Program

Troubleshooting large, distributed training jobs running on Ray and Anyscale can be time consuming and cumbersome. So we want to present our latest addition to the GPU Monitoring product that help you debug training job issues quickly. We’ve now added continuous tracing for these training jobs to help you deep dive into bottlenecks where we can directly tie CUDA or NCCL operations back to your actual model and PyTorch operations.

Related Resources

Interested in more of our latest features?

Help make the next releases of Datadog products our best yet.