Skip to content

Instantly share code, notes, and snippets.

View Anshuman-02905's full-sized avatar

Anshuman Mandal Anshuman-02905

View GitHub Profile
Glue Feature How EMR Uses It
πŸ—‚οΈ Glue Data Catalog Acts as EMR's Hive Metastore - so Spark SQL and Hive queries can discover schema & tables stored in S3.
🧩 Glue Crawlers Automatically infer schema from S3 data β†’ EMR Spark can read these tables instantly.
πŸͺ„ Glue ETL Jobs You can mix Glue ETL jobs (serverless) with EMR PySpark workflows - e.g., Glue for light jobs, EMR for heavy transformations.
πŸ”„ Glue Job Bookmarks Track incremental load progress (used in both Glue and EMR workflows).
πŸ” **Glue + L
Service Tagline Core Role
KDS 🧩 "Ingest Everything" Real-time event ingestion & buffering
KDA βš™οΈ "Process It in Motion" Real-time analytics & enrichment
KDF 🚚 "Deliver Everything" Reliable managed delivery to storage
KVS πŸŽ₯ "See Everything" Real-time video/audio ingestion & analytics
Feature Kinesis Data Streams (KDS) Kinesis Data Firehose (KDF)
Purpose Custom, scalable real-time stream processing Managed delivery of streaming data
Control Developer-managed (shards, consumers, scaling) Fully managed (auto-scaling, no shards)
Replay Support Yes (1–7 days retention) No (fire-and-forget delivery)
Processing Consumer applications process in near real-time Lambda transformations or format conversions
Destinations Any consumer (custom apps, Firehose, Analytics) S3, Redshift, OpenSearch, HTTP
Ideal For Real-time computation, analytics, custom logic Near-real-time ingestion, ETL, data lake loading
Feature Amazon EMR Azure Databricks
Engine Spark, Hive, Hadoop, Presto Spark
Cluster Management Fully managed, autoscaling Managed via workspace
Cost Optimization Spot + transient clusters Job clusters, autoscaling
Integration S3, Glue Catalog, Redshift ADLS, Synapse
Customization Root access, bootstrap scripts Limited (managed runtime)

AWS Glue vs Azure Equivalent

| Feature | AWS Glue | Azure Equivalent |

| ----------------- | ------------------- | --------------------------------------- |

| ETL Engine | Spark (Serverless) | Data Factory Data Flows / Synapse Spark |

| Metadata | Glue Catalog | Azure Purview |

Feature Kinesis Data Streams (AWS) Azure Event Hubs
Purpose Collect and store real-time event data for custom processing Same β€” ingest event, log, or telemetry data for downstream processing
Data Retention 1–7 days (default 24 hrs, extendable to 7) β€” optionally up to 365 days with extended retention 1–7 days (standard), up to 90 days with premium tiers
Scalability Unit Shards (~1 MB/s in, 2 MB/s out each) Throughput Units (1 MB/s in, 2 MB/s out)
Consumption Model Multiple consumers can read the same stream independently (enhanced fan-out) Multiple consumer groups can read the same event hub independently
Integration Feeds into Firehose, Lambda, KDA (Flink), or custom Spark apps Feeds into Stream Analytics, Functions, Databricks, or custom apps
Use Case Real-time ingestion for IoT, logs, clickstreams, metrics, CDC pipelines Same β€” telemetry, log ing
Feature AppFlow Azure Data Factory
Focus SaaS integrations Enterprise-scale ETL/ELT orchestration
Control Minimal setup, low code Full control, dynamic pipelines
Connectors Deep SaaS partnerships Broad hybrid connector library
Transformations Basic (mapping, filters) Complex data flows (Mapping Data Flows)
Orchestration ❌ Not supported βœ… Triggers, branching, retries
Best For Plug-and-play SaaS β†’ S3/Redshift Metadata-driven ETL pipelines
Feature AWS DMS Azure DMS
Migration Mode Full + Continuous (CDC) Full + Continuous (CDC)
Engine Support Heterogeneous (Oracle β†’ PostgreSQL) Heterogeneous (Oracle β†’ Azure SQL)
Monitoring CloudWatch Azure Portal
Security IAM Roles, VPC ExpressRoute, VPN
Integration S3, Redshift, Kinesis ADLS, Synapse
Layer Purpose Example AWS Services Azure Equivalent
Ingestion Bring data from SaaS, databases, or streams DMS, AppFlow, Kinesis ADF, Azure DMS, Event Hubs
Transformation Clean, enrich, and process raw data Glue, EMR Data Factory (Mapping Data Flows), Databricks
Storage / Modeling Persist structured and semi-structured data S3, Redshift, Lake Formation ADLS Gen2, Synapse, Purview
Consumption Analytics, BI, ML Athena, QuickSight, SageMaker Synapse, Power BI, Azure ML
Service Purpose Characteristics
Kinesis Data Streams (KDS) Raw event ingestion Flexible schema, developer-managed
Kinesis Data Firehose Reliable delivery to storage Schema-bound, delivery-optimized
Kinesis Data Analytics (KDA) Stream processing SQL or Apache Flink runtime
Kinesis Video Streams (KVS) Unstructured media streaming Handles live video/audio data