Distributed DL Training on Spot Instances with SageMaker

0

Hi,

The documentation on SageMaker suggests that one can do distributed deep learning training (multi-node) [1]. It is also possible to use Spot instances with Sagemaker [2]. Is it possible to combine these features and do multi-node distributed training on Spot instances? If yes, what is the failure semantics whenever peers drop out before submitting their gradients? I could not find any documentation on that matter.

[1] https://docs.aws.amazon.com/sagemaker/latest/dg/distributed-training.html [2] https://docs.aws.amazon.com/sagemaker/latest/dg/model-managed-spot-training.html

Keine Antworten

Du bist nicht angemeldet. Anmelden um eine Antwort zu veröffentlichen.

Eine gute Antwort beantwortet die Frage klar, gibt konstruktives Feedback und fördert die berufliche Weiterentwicklung des Fragenstellers.

Richtlinien für die Beantwortung von Fragen