| Glue Feature | How EMR Uses It |
|---|---|
| ποΈ Glue Data Catalog | Acts as EMR's Hive Metastore - so Spark SQL and Hive queries can discover schema & tables stored in S3. |
| π§© Glue Crawlers | Automatically infer schema from S3 data β EMR Spark can read these tables instantly. |
| πͺ Glue ETL Jobs | You can mix Glue ETL jobs (serverless) with EMR PySpark workflows - e.g., Glue for light jobs, EMR for heavy transformations. |
| π Glue Job Bookmarks | Track incremental load progress (used in both Glue and EMR workflows). |
| π **Glue + L |
| Service | Tagline | Core Role |
|---|---|---|
| KDS | π§© "Ingest Everything" | Real-time event ingestion & buffering |
| KDA | βοΈ "Process It in Motion" | Real-time analytics & enrichment |
| KDF | π "Deliver Everything" | Reliable managed delivery to storage |
| KVS | π₯ "See Everything" | Real-time video/audio ingestion & analytics |
| Feature | Kinesis Data Streams (KDS) | Kinesis Data Firehose (KDF) |
|---|---|---|
| Purpose | Custom, scalable real-time stream processing | Managed delivery of streaming data |
| Control | Developer-managed (shards, consumers, scaling) | Fully managed (auto-scaling, no shards) |
| Replay Support | Yes (1β7 days retention) | No (fire-and-forget delivery) |
| Processing | Consumer applications process in near real-time | Lambda transformations or format conversions |
| Destinations | Any consumer (custom apps, Firehose, Analytics) | S3, Redshift, OpenSearch, HTTP |
| Ideal For | Real-time computation, analytics, custom logic | Near-real-time ingestion, ETL, data lake loading |
| Feature | Amazon EMR | Azure Databricks |
|---|---|---|
| Engine | Spark, Hive, Hadoop, Presto | Spark |
| Cluster Management | Fully managed, autoscaling | Managed via workspace |
| Cost Optimization | Spot + transient clusters | Job clusters, autoscaling |
| Integration | S3, Glue Catalog, Redshift | ADLS, Synapse |
| Customization | Root access, bootstrap scripts | Limited (managed runtime) |
| Feature | Kinesis Data Streams (AWS) | Azure Event Hubs |
|---|---|---|
| Purpose | Collect and store real-time event data for custom processing | Same β ingest event, log, or telemetry data for downstream processing |
| Data Retention | 1β7 days (default 24 hrs, extendable to 7) β optionally up to 365 days with extended retention | 1β7 days (standard), up to 90 days with premium tiers |
| Scalability Unit | Shards (~1 MB/s in, 2 MB/s out each) | Throughput Units (1 MB/s in, 2 MB/s out) |
| Consumption Model | Multiple consumers can read the same stream independently (enhanced fan-out) | Multiple consumer groups can read the same event hub independently |
| Integration | Feeds into Firehose, Lambda, KDA (Flink), or custom Spark apps | Feeds into Stream Analytics, Functions, Databricks, or custom apps |
| Use Case | Real-time ingestion for IoT, logs, clickstreams, metrics, CDC pipelines | Same β telemetry, log ing |
| Feature | AppFlow | Azure Data Factory |
|---|---|---|
| Focus | SaaS integrations | Enterprise-scale ETL/ELT orchestration |
| Control | Minimal setup, low code | Full control, dynamic pipelines |
| Connectors | Deep SaaS partnerships | Broad hybrid connector library |
| Transformations | Basic (mapping, filters) | Complex data flows (Mapping Data Flows) |
| Orchestration | β Not supported | β Triggers, branching, retries |
| Best For | Plug-and-play SaaS β S3/Redshift | Metadata-driven ETL pipelines |
| Feature | AWS DMS | Azure DMS |
|---|---|---|
| Migration Mode | Full + Continuous (CDC) | Full + Continuous (CDC) |
| Engine Support | Heterogeneous (Oracle β PostgreSQL) | Heterogeneous (Oracle β Azure SQL) |
| Monitoring | CloudWatch | Azure Portal |
| Security | IAM Roles, VPC | ExpressRoute, VPN |
| Integration | S3, Redshift, Kinesis | ADLS, Synapse |
| Layer | Purpose | Example AWS Services | Azure Equivalent |
|---|---|---|---|
| Ingestion | Bring data from SaaS, databases, or streams | DMS, AppFlow, Kinesis | ADF, Azure DMS, Event Hubs |
| Transformation | Clean, enrich, and process raw data | Glue, EMR | Data Factory (Mapping Data Flows), Databricks |
| Storage / Modeling | Persist structured and semi-structured data | S3, Redshift, Lake Formation | ADLS Gen2, Synapse, Purview |
| Consumption | Analytics, BI, ML | Athena, QuickSight, SageMaker | Synapse, Power BI, Azure ML |
| Service | Purpose | Characteristics |
|---|---|---|
| Kinesis Data Streams (KDS) | Raw event ingestion | Flexible schema, developer-managed |
| Kinesis Data Firehose | Reliable delivery to storage | Schema-bound, delivery-optimized |
| Kinesis Data Analytics (KDA) | Stream processing | SQL or Apache Flink runtime |
| Kinesis Video Streams (KVS) | Unstructured media streaming | Handles live video/audio data |