Distributed DL Training on Spot Instances with SageMaker

0

Hi,

The documentation on SageMaker suggests that one can do distributed deep learning training (multi-node) [1]. It is also possible to use Spot instances with Sagemaker [2]. Is it possible to combine these features and do multi-node distributed training on Spot instances? If yes, what is the failure semantics whenever peers drop out before submitting their gradients? I could not find any documentation on that matter.

[1] https://docs.aws.amazon.com/sagemaker/latest/dg/distributed-training.html [2] https://docs.aws.amazon.com/sagemaker/latest/dg/model-managed-spot-training.html

Nessuna risposta

Accesso non effettuato. Accedi per postare una risposta.

Una buona risposta soddisfa chiaramente la domanda, fornisce un feedback costruttivo e incoraggia la crescita professionale del richiedente.

Linee guida per rispondere alle domande