AI Watch

Fault tolerant distributed training on Amazon EKS using NVRx

AWS Machine Learning Aravind Neelakantan 15 min de lecture Anglais

Résumé de la publication. Les conditions de AWS Machine Learning ne permettent pas de reproduire l'article en intégralité : retrouvez le texte complet sur le site d'origine.

Résumé de la source (Anglais)

Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%+ training efficiency and second-scale recovery.

Lire l'article complet

aws.amazon.com

Voir sur AWS Machine Learning (nouvel onglet)