Skip to content

All Content tagged with Amazon EMR

Amazon EMR is a cloud big data platform for running large-scale distributed data processing jobs, interactive SQL queries, and machine learning (ML) applications using open-source analytics frameworks such as Apache Spark, Apache Hive, and Presto.

Content language: English

Filter content
Select tags to filter
Sort by
Sort by most recent
465 results
When you enable Lake Formation runtime roles on an EMR 7.x cluster, all Spark steps may silently time out after 300 seconds. This happens because S3 locations registered with the default Service-Linke...
Running Spark on EMR with KMS-encrypted S3 data? Every object read triggers a kms:Decrypt API call — and at scale, those costs add up fast. If your compliance requirements prevent switching to S3 Buck...
A field guide for PyTorch, TensorFlow, Spark, and Kubernetes workloads reading training data from Amazon S3 Express One Zone directory buckets.
**Service and environment** * Product: Amazon EMR * Release label: emr-7.13.0 * Applications: Spark, Livy (and others as applicable) * Region: [e.g. us-east-1] * Workload: PySpark / Livy submitting Py...
2
answers
2
votes
268
views

asked 3 months ago

I want to view the Spark UI for my AWS Glue job runs, but I cannot use Docker on my local machine. I need an alternative way to run the Apache Spark History Server natively on macOS to read Spark even...
I'm trying to install Python modules in my AWS Glue Python Shell job using wheel files stored in Amazon Simple Storage Service (Amazon S3). My job runs in a private Virtual Private Cloud (VPC) with Am...
This framework provides a structured approach for migrating analytics workloads from EMR on EC2 to EMR Serverless in enterprise environments. It guides organizations through the complete migration lif...
Enterprises struggle with EMR version upgrades, facing challenges like production downtime, performance degradation, and compliance risks. Without a structured approach, organizations often experience...
Hello, I am running an Apache Spark job on Amazon EMR that needs to connect to an Amazon MSK cluster configured with IAM authentication. The EMR cluster has an IAM role with full MSK permissions, and...
1
answers
0
votes
355
views

asked 10 months ago

Hi Team, I'm trying to set up the Amazon Q on the EMR studio notebook workspace and followed this guide: https://docs.aws.amazon.com/amazonq/latest/qdeveloper-ug/emr-setup.html?trk=769a1a2b-8c19-4976...
1
answers
0
votes
158
views

asked 10 months ago

Getting this issue in Amazon EMR during a pyspark job execution. ``` df = spark.read.parquet("s3a://test/raw-billing-cor-data/cur2/123456789/cid-cur2/data/BILLING_PERIOD=2025-08/") py4j.protocol.Py4...
1
answers
0
votes
241
views

asked 10 months ago

  • 1
  • 2
  • 3
  • 4
  • 5
  • •••
  • 39
  • Page size
    12 / page