跳至內容

EMR HDFS data restore

0

Hello Experts,

Technically speaking, EBS volumes assigned to the EMR core nodes are persistent storage and I have specifically created them to not delete on cluster termination. Then, I have attached the same volume to new EMR cluster and mounted them back.

After restarting the data node service, I got an exception stating "clusterID" version is not matched. To resolve this, I matched the VERSION file same as master node. However I unable to see the data that stored in the hdfs volume. I m not hadoop expert though, but I think there is some gap here.

Question is why I can't reuse the ebs volume in emr cluster

Thanks in advance.

已提問 2 年前檢視次數 751 次

1 個回答
4
已接受的答案

Hi,

Basically, the data is stored as HDFS blocks on these disks and only the NameNode is aware of these blocks (stores these blocks details as metadata)

Hypothetically, even if you were able to re-attach the EBS volumes to another cluster, the newer or other cluster Namenode is unaware of these HDFS blocks. Please note each data block(HDFS) is tied to a hadoop cluster and has its own unique block id and block location on disks differs. (HDFS blocks != OS blocks)

Because of the above framework limitation, it is not possible to switch EBS volumes since EBS is not a persistent store in EMR, meaning it is deleted once the EMR cluster is terminated. Even if we can attach additional EBS volumes to a node in EMR cluster, it is not possible to reuse them with another cluster. Please refer this document for details. I hope this answers your question.

AWS
支援工程師

已回答 2 年前

專家

已審閱 2 年前

專家

已審閱 2 年前

您尚未登入。 登入 去張貼答案。

一個好的回答可以清楚地回答問題並提供建設性的意見回饋,同時有助於提問者的專業成長。