How to replicate Glue bookmark in custom spark aws glue script, where i am not able to use dynamic frame to read the data but need to process only new files in the source S3 folder

0

How to replicate Glue bookmark in custom spark aws glue script, where i am not able to use dynamic frame to read the data but need to process only new files in the source S3 folder

KG
已提問 6 個月前檢視次數 186 次
1 個回答
0

You can't use bookmarks without DynamicFrame, why are you enable to use it (you can convert to DataFrame as soon as you read)?
Otherwise you have to do your own listing of files and passing that truncated list to Spark

profile pictureAWS
專家
已回答 6 個月前
  • the files i am trying to read have a gzip compression and also base64 encoded then its a json file, i am unable to read the file in a dynamic frame , and i want to read only new files which come in the source folder

  • I think the issue there is the base64, that's not standard and you would need to read as text, then decode then parse as json. You can do that with plain Spark but then you would lose the bookmarks. What you could do is read with DynamicFrame as csv (even though is not csv), all the encoding will be in one column and if you don't have memory issues you could decode and parse.

您尚未登入。 登入 去張貼答案。

一個好的回答可以清楚地回答問題並提供建設性的意見回饋,同時有助於提問者的專業成長。

回答問題指南