Skip to content

Instantly share code, notes, and snippets.

@steniowagner
Created August 7, 2026 14:06
Show Gist options
  • Select an option

  • Save steniowagner/7f8831c77964fea622fbd82067426213 to your computer and use it in GitHub Desktop.

Select an option

Save steniowagner/7f8831c77964fea622fbd82067426213 to your computer and use it in GitHub Desktop.

João Victor de Souza Duarte

Senior Data Engineer

📍 Rio de Janeiro, Brazil
📧 jduarte.dev@gmail.com
📱 +55 21 96810-2733
🔗 LinkedIn: https://www.linkedin.com/in/jvsduarte
💻 GitHub: https://github.com/joaovictordesouzaduarte


Professional Summary

Senior Data Engineer specializing in building scalable cloud-native data platforms on AWS and Databricks. Experienced in architecting modern Lakehouse solutions, optimizing distributed Spark workloads, and delivering production-grade ETL/ELT pipelines that power business-critical analytics.

Strong expertise in Data Platform Architecture, Change Data Capture (CDC), Infrastructure as Code, Data Governance, and distributed computing, with a focus on building reliable, high-performance data platforms that enable data-driven decision making.


Selected Achievements

  • Architected an AWS Lakehouse platform enabling near real-time Change Data Capture (CDC) while improving analytical query performance by 30%.
  • Reduced distributed Spark processing time by 75% through execution plan analysis, partition redesign, cluster tuning, and storage optimization.
  • Delivered end-to-end data products that reduced manual operational effort by 80%, enabling self-service analytics across business teams.
  • Designed and implemented production-grade cloud data platforms leveraging AWS, Databricks, Apache Spark, and Infrastructure as Code.

Technical Expertise

Cloud Platforms

  • AWS (S3, Glue, Athena, ECS, Lambda, EMR, IAM, Redshift)

Data Platforms

  • Databricks
  • Apache Spark
  • PySpark
  • Spark SQL
  • Delta Lake
  • Apache Iceberg
  • Unity Catalog

Data Engineering

  • Apache Airflow
  • dbt
  • ETL / ELT
  • Data Lakehouse
  • Medallion Architecture
  • Change Data Capture (CDC)
  • Structured Streaming
  • Incremental Processing
  • Data Modeling
  • Data Quality
  • Performance Optimization

Programming

  • Python
  • SQL

Infrastructure & DevOps

  • Terraform
  • Docker
  • GitHub Actions
  • CI/CD

Visualization

  • Amazon QuickSight
  • Plotly
  • D3.js

Professional Experience

Data Engineer

JGP Asset Management
October 2022 – Present

Designed, built, and evolved modern cloud-native data platforms supporting financial analytics, business intelligence, and operational reporting.

Data Platform Architecture

  • Architected and owned the implementation of a cloud-native AWS Lakehouse platform using AWS DMS, Amazon S3, AWS Glue, Amazon Athena, and Apache Iceberg, establishing a scalable analytical foundation with near real-time Change Data Capture (CDC) ingestion while improving analytical query performance by 30%.

  • Designed and standardized enterprise-grade Databricks pipelines using PySpark, Spark SQL, Delta Lake, Unity Catalog, and Medallion Architecture, improving data quality, governance, and maintainability while enabling scalable downstream analytics.

Distributed Data Processing

  • Led Spark performance optimization initiatives by analyzing execution plans, eliminating data skew, redesigning partition strategies, compacting small files, and tuning compute clusters, reducing distributed processing time by 75% while improving platform efficiency and operational reliability.

Data Integration & Engineering

  • Built resilient ingestion frameworks orchestrated with Apache Airflow, integrating structured and unstructured data from relational databases, APIs, spreadsheets, XML files, email services, flat files, and external web platforms through automated, fault-tolerant pipelines.

  • Developed reusable SQL models, database objects, and analytical datasets supporting business-critical reporting and operational analytics while improving query performance and simplifying data consumption across the organization.

Data Products & Business Impact

  • Delivered end-to-end data products by combining scalable ETL pipelines with interactive analytical dashboards, reducing manual operational effort by 80% while enabling self-service analytics and accelerating business decision-making.

Projects

AWS Lakehouse Platform

Production-grade cloud-native data platform demonstrating modern Data Engineering best practices from ingestion through business intelligence.

Highlights

  • Architected a scalable AWS Lakehouse implementing Change Data Capture (CDC) from PostgreSQL into Amazon S3 using AWS DMS.
  • Designed a Medallion Architecture powered by dbt and Apache Iceberg, transforming raw operational data into analytics-ready datasets.
  • Implemented Infrastructure as Code using CloudFormation, provisioning AWS resources including S3, RDS, DMS, Glue, Athena, IAM, ECS, and ECR.
  • Built a fully automated CI/CD pipeline using GitHub Actions.
  • Orchestrated scheduled transformations on Amazon ECS Fargate through EventBridge Scheduler.
  • Delivered business-ready analytics through Amazon Athena and QuickSight dashboards.

Technologies

AWS • Apache Iceberg • dbt • ECS Fargate • Athena • Glue • CloudFormation • GitHub Actions • PostgreSQL • DMS • QuickSight

GitHub

https://github.com/joaovictordesouzaduarte/aws_datalakehouse_platform


MercadoLibre Serverless Data Pipeline

Production-inspired serverless data ingestion platform built on AWS for automated web data extraction and cloud-native processing.

Highlights

  • Built a containerized Selenium-based scraping solution running on AWS Lambda with headless Chrome.
  • Automated product extraction, CSV generation, and Amazon S3 ingestion.
  • Provisioned the entire cloud infrastructure using Terraform.
  • Implemented container image publishing through Amazon ECR, enabling reproducible deployments.
  • Designed a fully serverless architecture requiring minimal operational overhead.

Technologies

AWS Lambda • Selenium • Docker • Terraform • Amazon S3 • Amazon ECR

GitHub

https://github.com/joaovictordesouzaduarte/mercadolibre_scrapper


Education

MBA in Data Engineering (Big Data)

INFNET
2025 – Present


Bachelor of Science in Economics

Federal University of Rio de Janeiro (UFRJ)
2022

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment