Skip to content

Instantly share code, notes, and snippets.

View Anshuman-02905's full-sized avatar

Anshuman Mandal Anshuman-02905

View GitHub Profile
Feature AppFlow Azure Data Factory
Focus SaaS integrations Enterprise-scale ETL/ELT orchestration
Control Minimal setup, low code Full control, dynamic pipelines
Connectors Deep SaaS partnerships Broad hybrid connector library
Transformations Basic (mapping, filters) Complex data flows (Mapping Data Flows)
Orchestration ❌ Not supported ✅ Triggers, branching, retries
Best For Plug-and-play SaaS → S3/Redshift Metadata-driven ETL pipelines
Feature Kinesis Data Streams (AWS) Azure Event Hubs
Purpose Collect and store real-time event data for custom processing Same — ingest event, log, or telemetry data for downstream processing
Data Retention 1–7 days (default 24 hrs, extendable to 7) — optionally up to 365 days with extended retention 1–7 days (standard), up to 90 days with premium tiers
Scalability Unit Shards (~1 MB/s in, 2 MB/s out each) Throughput Units (1 MB/s in, 2 MB/s out)
Consumption Model Multiple consumers can read the same stream independently (enhanced fan-out) Multiple consumer groups can read the same event hub independently
Integration Feeds into Firehose, Lambda, KDA (Flink), or custom Spark apps Feeds into Stream Analytics, Functions, Databricks, or custom apps
Use Case Real-time ingestion for IoT, logs, clickstreams, metrics, CDC pipelines Same — telemetry, log ing

AWS Glue vs Azure Equivalent

| Feature | AWS Glue | Azure Equivalent |

| ----------------- | ------------------- | --------------------------------------- |

| ETL Engine | Spark (Serverless) | Data Factory Data Flows / Synapse Spark |

| Metadata | Glue Catalog | Azure Purview |

Feature Amazon EMR Azure Databricks
Engine Spark, Hive, Hadoop, Presto Spark
Cluster Management Fully managed, autoscaling Managed via workspace
Cost Optimization Spot + transient clusters Job clusters, autoscaling
Integration S3, Glue Catalog, Redshift ADLS, Synapse
Customization Root access, bootstrap scripts Limited (managed runtime)
Feature Kinesis Data Streams (KDS) Kinesis Data Firehose (KDF)
Purpose Custom, scalable real-time stream processing Managed delivery of streaming data
Control Developer-managed (shards, consumers, scaling) Fully managed (auto-scaling, no shards)
Replay Support Yes (1–7 days retention) No (fire-and-forget delivery)
Processing Consumer applications process in near real-time Lambda transformations or format conversions
Destinations Any consumer (custom apps, Firehose, Analytics) S3, Redshift, OpenSearch, HTTP
Ideal For Real-time computation, analytics, custom logic Near-real-time ingestion, ETL, data lake loading
Service Tagline Core Role
KDS 🧩 "Ingest Everything" Real-time event ingestion & buffering
KDA ⚙️ "Process It in Motion" Real-time analytics & enrichment
KDF 🚚 "Deliver Everything" Reliable managed delivery to storage
KVS 🎥 "See Everything" Real-time video/audio ingestion & analytics
Glue Feature How EMR Uses It
🗂️ Glue Data Catalog Acts as EMR's Hive Metastore - so Spark SQL and Hive queries can discover schema & tables stored in S3.
🧩 Glue Crawlers Automatically infer schema from S3 data → EMR Spark can read these tables instantly.
🪄 Glue ETL Jobs You can mix Glue ETL jobs (serverless) with EMR PySpark workflows - e.g., Glue for light jobs, EMR for heavy transformations.
🔄 Glue Job Bookmarks Track incremental load progress (used in both Glue and EMR workflows).
🔐 **Glue + L
Step Service Purpose Azure Counterpart
1️⃣ AWS Glue Serverless ETL with PySpark ADF Mapping Data Flows
2️⃣ Amazon EMR Scalable, managed Spark/Hive clusters Azure Databricks
3️⃣ AWS Lambda Event-driven or micro-transformations Azure Functions
4️⃣ Step Functions (optional) Pipeline orchestration & retries ADF Pipelines / Logic Apps
Scenario Best AWS Service Azure Equivalent
Metadata-driven, serverless ETL Glue ADF Data Flows
Heavy Spark-based transformation EMR Databricks
Small real-time transformations Lambda Functions
Multi-step workflow orchestration Step Functions ADF Pipelines
Capability AWS Service Azure Equivalent Notes
Data Lake Storage S3 ADLS Gen2 Core storage for raw → curated data
Data Warehouse Redshift Synapse (Dedicated) MPP engine for structured analytics
Serverless Querying Athena Synapse Serverless Query directly on data lake
Catalog & Governance Lake Formation + Glue Catalog Purview / Unity Catalog Fine-grained security, metadata lineage