Skip to content

Articles tagged with Amazon EMR

Amazon EMR is a cloud big data platform for running large-scale distributed data processing jobs, interactive SQL queries, and machine learning (ML) applications using open-source analytics frameworks such as Apache Spark, Apache Hive, and Presto.

Content language: English

Filter articles
Select tags to filter
Sort by
Sort by most recent

Browse through articles or filter your results using the tools displayed.

22 results
In this guide, I will provide the EMR cluster configuration and Spark config required to launch it in FTA mode. EMR Spark and Glue Spark are available in FTA and FGAC modes, to work with AWS Lake For...
This article aims to provide a practical guide for cross account setup
Amazon EMR and AWS Glue Spark provides two modes for Apache Spark workloads that access tables in the AWS Glue Data Catalog that are governed by AWS Lake Formation. Learn when to use what configuratio...
When you enable Lake Formation runtime roles on an EMR 7.x cluster, all Spark steps may silently time out after 300 seconds. This happens because S3 locations registered with the default Service-Linke...
Running Spark on EMR with KMS-encrypted S3 data? Every object read triggers a kms:Decrypt API call — and at scale, those costs add up fast. If your compliance requirements prevent switching to S3 Buck...
A field guide for PyTorch, TensorFlow, Spark, and Kubernetes workloads reading training data from Amazon S3 Express One Zone directory buckets.
I want to view the Spark UI for my AWS Glue job runs, but I cannot use Docker on my local machine. I need an alternative way to run the Apache Spark History Server natively on macOS to read Spark even...
I'm trying to install Python modules in my AWS Glue Python Shell job using wheel files stored in Amazon Simple Storage Service (Amazon S3). My job runs in a private Virtual Private Cloud (VPC) with Am...
This framework provides a structured approach for migrating analytics workloads from EMR on EC2 to EMR Serverless in enterprise environments. It guides organizations through the complete migration lif...
Enterprises struggle with EMR version upgrades, facing challenges like production downtime, performance degradation, and compliance risks. Without a structured approach, organizations often experience...
Performance testing for big data analytics tools and engines at petabyte scale is an increasingly challenging avenue. Using traditional sample test datasets may not reflect the actual production-grade...
AWS

Yokesh NKSUPPORT ENGINEER

published 2 years ago3 votes1.9K views

This article offers instructions on how to set up and access Delta tables from SQL Explorer in EMR JupyterHub. SQL Explorer utilizes the Presto engine configured within the EMR cluster to process data...
Amazon EMR
  • 1
  • 2
  • Page size
    12 / page