Senior Data Engineer
📍 Rio de Janeiro, Brazil
📧 jduarte.dev@gmail.com
📱 +55 21 96810-2733
🔗 LinkedIn: https://www.linkedin.com/in/jvsduarte
💻 GitHub: https://github.com/joaovictordesouzaduarte
Senior Data Engineer specializing in building scalable cloud-native data platforms on AWS and Databricks. Experienced in architecting modern Lakehouse solutions, optimizing distributed Spark workloads, and delivering production-grade ETL/ELT pipelines that power business-critical analytics.
Strong expertise in Data Platform Architecture, Change Data Capture (CDC), Infrastructure as Code, Data Governance, and distributed computing, with a focus on building reliable, high-performance data platforms that enable data-driven decision making.
- Architected an AWS Lakehouse platform enabling near real-time Change Data Capture (CDC) while improving analytical query performance by 30%.
- Reduced distributed Spark processing time by 75% through execution plan analysis, partition redesign, cluster tuning, and storage optimization.
- Delivered end-to-end data products that reduced manual operational effort by 80%, enabling self-service analytics across business teams.
- Designed and implemented production-grade cloud data platforms leveraging AWS, Databricks, Apache Spark, and Infrastructure as Code.
- AWS (S3, Glue, Athena, ECS, Lambda, EMR, IAM, Redshift)
- Databricks
- Apache Spark
- PySpark
- Spark SQL
- Delta Lake
- Apache Iceberg
- Unity Catalog
- Apache Airflow
- dbt
- ETL / ELT
- Data Lakehouse
- Medallion Architecture
- Change Data Capture (CDC)
- Structured Streaming
- Incremental Processing
- Data Modeling
- Data Quality
- Performance Optimization
- Python
- SQL
- Terraform
- Docker
- GitHub Actions
- CI/CD
- Amazon QuickSight
- Plotly
- D3.js
JGP Asset Management
October 2022 – Present
Designed, built, and evolved modern cloud-native data platforms supporting financial analytics, business intelligence, and operational reporting.
-
Architected and owned the implementation of a cloud-native AWS Lakehouse platform using AWS DMS, Amazon S3, AWS Glue, Amazon Athena, and Apache Iceberg, establishing a scalable analytical foundation with near real-time Change Data Capture (CDC) ingestion while improving analytical query performance by 30%.
-
Designed and standardized enterprise-grade Databricks pipelines using PySpark, Spark SQL, Delta Lake, Unity Catalog, and Medallion Architecture, improving data quality, governance, and maintainability while enabling scalable downstream analytics.
- Led Spark performance optimization initiatives by analyzing execution plans, eliminating data skew, redesigning partition strategies, compacting small files, and tuning compute clusters, reducing distributed processing time by 75% while improving platform efficiency and operational reliability.
-
Built resilient ingestion frameworks orchestrated with Apache Airflow, integrating structured and unstructured data from relational databases, APIs, spreadsheets, XML files, email services, flat files, and external web platforms through automated, fault-tolerant pipelines.
-
Developed reusable SQL models, database objects, and analytical datasets supporting business-critical reporting and operational analytics while improving query performance and simplifying data consumption across the organization.
- Delivered end-to-end data products by combining scalable ETL pipelines with interactive analytical dashboards, reducing manual operational effort by 80% while enabling self-service analytics and accelerating business decision-making.
Production-grade cloud-native data platform demonstrating modern Data Engineering best practices from ingestion through business intelligence.
- Architected a scalable AWS Lakehouse implementing Change Data Capture (CDC) from PostgreSQL into Amazon S3 using AWS DMS.
- Designed a Medallion Architecture powered by dbt and Apache Iceberg, transforming raw operational data into analytics-ready datasets.
- Implemented Infrastructure as Code using CloudFormation, provisioning AWS resources including S3, RDS, DMS, Glue, Athena, IAM, ECS, and ECR.
- Built a fully automated CI/CD pipeline using GitHub Actions.
- Orchestrated scheduled transformations on Amazon ECS Fargate through EventBridge Scheduler.
- Delivered business-ready analytics through Amazon Athena and QuickSight dashboards.
Technologies
AWS • Apache Iceberg • dbt • ECS Fargate • Athena • Glue • CloudFormation • GitHub Actions • PostgreSQL • DMS • QuickSight
GitHub
https://github.com/joaovictordesouzaduarte/aws_datalakehouse_platform
Production-inspired serverless data ingestion platform built on AWS for automated web data extraction and cloud-native processing.
- Built a containerized Selenium-based scraping solution running on AWS Lambda with headless Chrome.
- Automated product extraction, CSV generation, and Amazon S3 ingestion.
- Provisioned the entire cloud infrastructure using Terraform.
- Implemented container image publishing through Amazon ECR, enabling reproducible deployments.
- Designed a fully serverless architecture requiring minimal operational overhead.
Technologies
AWS Lambda • Selenium • Docker • Terraform • Amazon S3 • Amazon ECR
GitHub
https://github.com/joaovictordesouzaduarte/mercadolibre_scrapper
INFNET
2025 – Present
Federal University of Rio de Janeiro (UFRJ)
2022