Checkpointing an instance.

0

Hi, I am trying to run a scientific calculation job on an instance. It is a large calculation which should take a lot of time to compute. Due to the dependency among the calculations, it cannot be parallelized.My concern is that if the instance fails, I will lose all the progress. What are the best practices where if the instance fails, I can resume the computation without having to restart from the beginning.I will prefer to checkpoint periodically and launch another instance from that checkpoint. Does AWS have any built in mechanisms that I can use to checkpoint? thanks

AG
질문됨 7달 전209회 조회
1개 답변
0

You can snapshot the EBS volumes and then restore that to a new volume and resume. It will take some development on your part, and you'll need to ensure you supply the correct permissions to your instance role to take snapshots and restore them. See https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/EBSSnapshots.html and https://docs.aws.amazon.com/AWSEC2/latest/WindowsGuide/ebs-creating-snapshot.html and https://docs.aws.amazon.com/prescriptive-guidance/latest/backup-recovery/restore.html

profile pictureAWS
답변함 7달 전
profile picture
전문가
검토됨 2달 전

로그인하지 않았습니다. 로그인해야 답변을 게시할 수 있습니다.

좋은 답변은 질문에 명확하게 답하고 건설적인 피드백을 제공하며 질문자의 전문적인 성장을 장려합니다.

질문 답변하기에 대한 가이드라인

관련 콘텐츠