Data Engineer with extensive expertise in building end-to-end batch data pipelines and real-time streaming solutions. Specialized in cloud platforms (GCP, AWS, Azure) and modern data technologies. Passionate about building scalable, reliable, and cost-efficient data solutions that drive business value.
- Google Cloud Platform (GCP) - Databricks, BigQuery, Dataproc, Pub/Sub, Cloud Run, Cloud Composer, and GCS
- Amazon Web Services (AWS) - S3, DynamoDB, Redshift, Kinesis, Lambda, Airflow, EC2, and Glue
- Microsoft Azure - ADLS, Synapse, ADF, CosmosDB, MS-Fabric, Event Hub, Stream Analytics
- Orchestration: Apache Airflow, Cloud Composer, Azure Data Factory, and SSIS
- Streaming: Apache Kafka, Confluent Cloud, Azure Event Hub, Apache Flink, Redpanda
- Data Warehousing: Snowflake, BigQuery, Redshift, Azure Synapse, Databricks
- Data Processing: PySpark, Apache Spark, dbt, Microsoft Fabric, Delta Lake
- Transformation: dbt, Apache Beam, Dataflow, Azure Data Factory
- Storage: GCS, S3, ADLS, Apache Iceberg, Delta Lake
- Languages: Python, SQL (T-SQL), and Pyspark
- Tools: dbt, SSIS, SSRS, PowerBI, ThoughtSpot, Looker, Streamlit, n8n
- Infrastructure: Docker, GitHub Actions, Azure DevOps CI/CD, ARM Templates
- Version Control: Git, GitHub, TFS
- Kafka Stock Market Pipeline - Real-time stock market data processing
- Azure Event Hub BookMyShow Pipeline - Real-time event booking and payment streams with Stream Analytics
- GCP Uber Car Idle Alerts - Pub/Sub to BigQuery real-time alerts using Dataflow
- Databricks UPI Transactions CDC - Real-time transaction change data capture with PySpark Streaming
- Confluent Kafka MongoDB Streaming - Kafka to MongoDB real-time pipelines
- AWS Kinesis Spark Streaming - Real-time food delivery pipeline with Spark Streaming
- Flink Real-Time Processing - Redpanda, Flink, and Postgres integration for real-time analytics
- Flight Booking Pipeline - Airflow β GCS β BigQuery β Looker
- Credit Card Processing - End-to-end credit card data pipeline with Looker dashboards
- YouTube Trends Pipeline - GCS Iceberg tables to BigQuery analytics
- Weather Map API Integration - API data extraction and PySpark transformation
- Bigtable CRUD Operations - NoSQL database operations with Python
- News API GCS to Snowflake - Multi-warehouse pipeline with Airflow
- Airflow Dataproc PySpark - Ephemeral Dataproc cluster orchestration
- Pyspark Backfill Pipeline - Backfill DAGs with PySpark jobs
- Real-time Gaming Leaderboard - Low-latency analytics with Dataflow
- Snowflake Car Rental - GCP Airflow to Snowflake integration
- SCD Type 2 CI/CD - GitHub Actions automation for data processing
- Airline Data Ingestion - S3 to Redshift pipeline with data quality checks
- Movie Quality Pipeline - Data quality framework and monitoring
- S3 to Redshift Pipeline - Airflow orchestration with Redshift
- DBT Snowflake Preset - dbt transformations with Preset dashboarding
- Airbnb CosmosDB Pipeline - Near real-time pipeline with CosmosDB and n8n workflows
- Airport ADF with CI/CD - Data Factory with DevOps automation and ARM templates
- Fintech SQL Pipeline - Multi-warehouse architecture with Synapse
- Analysis Services Model - OLAP cube development from SQL Server
- Olympic Data Engineering - End-to-end analytics project
- Snowflake Movies Streamlit App - Dynamic tables with interactive Streamlit dashboards
- Databricks Travel Booking SCD2 - Slowly Changing Dimensions Type 2 implementation
- Databricks dbt Project - dbt frameworks on Databricks
- dbt Snowflake - dbt with Snowflake integration
- Microsoft Fabric End-to-End - Bronze-Silver-Gold architecture on Fabric
- Microsoft Fabric Uber Analytics - Comprehensive analytics platform
- Microsoft Fabric API PowerBI - API ingestion with PowerBI reporting
- GCP PySpark Streaming Iceberg - Open table format implementation
- Databricks Ecommerce Event-Driven - Event-driven architecture on Databricks
- Databricks Healthcare DLT - Delta Live Tables for healthcare data
- Airflow GCP Complete Stack - Complete GCP visualization stack
- Yahoo Finance API - Financial data extraction with Airflow
- Data Engineer Handbook - Comprehensive data engineering resources and best practices
- Awesome Data Engineering - Curated list of tools, frameworks, and resources
- System Design Academy - System design principles for scalability
- Data Structures & Algorithms (Python) - DSA templates, solutions, and practical projects
- Python Web Scraping Projects - Web data extraction techniques and examples
- PySpark Project Files - Spark learning resources and implementations
- Linux Project Files - Linux system administration knowledge base
- βοΈ Azure Cloud Professional Certifications
- βοΈ GCP Data Engineering Certifications
- βοΈ AWS Solutions Architect Certifications
- π Data Engineering Specialized Certifications
- π Technology Stack Specific Training
- π Advanced Tools & Frameworks Certifications
- π Multiple Udemy courses on data platforms
- π Official cloud provider training programs
- ποΈ Awards and Recognitions in IT career
- ποΈ Extensive hands-on project experience
- ποΈ Community contributions and knowledge sharing
| Category | Count | Focus |
|---|---|---|
| GCP Projects | 15+ | Cloud data pipelines, BigQuery, Dataproc, Airflow |
| AWS Projects | 10+ | Redshift, S3, Kinesis, Lambda pipelines |
| Azure Projects | 12+ | Synapse, Data Factory, ADLS, CosmosDB |
| Databricks Projects | 5+ | dbt, Delta, Streaming, MLOps |
| Streaming Projects | 8+ | Kafka, Kinesis, Event Hub, Flink, Pub/Sub |
| Microsoft Fabric | 3+ | End-to-end lakehouse solutions |
| Learning Resources | 10+ | DSA, Python, Linux, System Design |
| Total Repositories | 75+ | Production-grade implementations |
β
Multi-Cloud Architecture - Design and implement solutions across GCP, AWS, Azure
β
Real-Time Streaming - Kafka, Kinesis, Event Hub, Flink, Pub/Sub expertise
β
Data Warehousing - Snowflake, BigQuery, Redshift, Synapse optimization
β
Data Transformation - dbt, PySpark, Apache Spark, SQL optimization
β
Orchestration - Apache Airflow, Cloud Composer, Data Factory automation
β
CI/CD & DevOps - GitHub Actions, Azure DevOps, Infrastructure as Code
β
Data Modeling - SCD, Dimensional modeling, Star Schema, Lakehouse architecture
β
Analytics & BI - PowerBI, Looker, Streamlit dashboard development
β
Full-Stack Pipelines - End-to-end implementations from ingestion to visualization
β
Documentation & Knowledge Sharing - Best practices, tutorials, mentoring
I believe in building production-grade data solutions that embody:
- Scalability - Handle data at any volume with optimal performance
- Reliability - Robust error handling, monitoring, and alerting
- Maintainability - Clean, well-documented, and testable code
- Cost-Efficiency - Optimized resource utilization and cloud spending
- Future-Proof - Aligned with industry best practices and emerging technologies
- GitHub Profile: github.com/ViinayKumaarMamidi
- All Repositories: 75+ Active Projects
I'm passionate about:
- π Designing scalable data architectures
- π Building reliable data pipelines both batch and real time
- βοΈ Leveraging cloud technologies effectively
- π Transforming data into actionable insights
- π€ Sharing knowledge and mentoring others
- βοΈ Exploring AI tools and planning on leveraging AI to build data pipelines in cloud environments
Interested in collaborating? Please feel free to explore my repositories, fork projects, or reach out with ideas or any suggestions!

