Scala Developer
October 7, 2026
Experience: 5–7 Years
Budget: 1L-1.4L
JD:
Job Overview
We are looking for an experienced Data Engineer with 5–7 years of hands-on experience in building scalable and reliable data pipelines using Python, SQL, PySpark, and AWS data services.
The ideal candidate should have strong experience in designing raw, standardized, and curated data pipelines, implementing data quality and reconciliation frameworks, and preparing data for analytics, dashboards, feature engineering, and machine learning consumption.
Experience working with retail, commerce, media, or CPG datasets is preferred.
Key Responsibilities
- Design, develop, and maintain scalable raw, standardized, and curated data pipelines across growing data volumes.
- Build reusable and well-documented data products for ASIN, SKU, product, brand, category, inventory, search, media, and vendor performance.
- Develop large-scale data processing solutions using Python, SQL, and PySpark.
- Work extensively with AWS data services including Amazon S3, AWS Glue, Athena, Lake Formation, Step Functions, and Lambda.
- Implement API-based and batch data ingestion patterns from multiple source systems.
- Develop robust data quality, reconciliation, completeness, and freshness checks to identify data issues before downstream consumption.
- Support ASIN/SKU/product identity mapping and source crosswalks across disparate systems.
- Apply appropriate data modelling, partitioning, metadata management, and data lineage strategies.
- Prepare and structure data for analytics, dashboards, feature engineering, and machine learning/model consumption.
- Monitor pipeline health, investigate failures, troubleshoot data issues, and improve pipeline reliability and performance.
- Collaborate with Data Analysts, Data Scientists, Business Stakeholders, and other Engineering teams to understand data requirements and translate them into scalable technical solutions.
- Continuously optimize data pipelines for performance, scalability, reliability, and maintainability.
- Create technical documentation for data pipelines, data products, source mappings, and data quality processes.
Mandatory Skills
- 5–7 years of Data Engineering experience
- Strong hands-on experience with SQL, Python, and PySpark
- Strong experience with AWS data services
- Amazon S3
- AWS Glue
- Amazon Athena
- AWS Lake Formation
- AWS Step Functions
- AWS Lambda
- Experience building ETL/ELT and scalable data pipelines
- Experience with raw, standardized, and curated data layers
- Strong understanding of data modelling and partitioning strategies
- Experience with API-based and batch ingestion
- Experience implementing data quality, reconciliation, and freshness checks
- Understanding of metadata management and data lineage
- Experience with pipeline monitoring, troubleshooting, and performance optimization
Good to Have
- Experience with Snowflake or Databricks
- Experience with retail, e-commerce, commerce, media, or CPG datasets
- Experience with ASIN/SKU/product identity mapping
- Knowledge of analytics and BI data preparation
- Experience supporting feature engineering and ML data pipelines
- AWS certification such as:
- AWS Certified Data Engineer
- AWS Certified Solutions Architect
Preferred Candidate Profile
The ideal candidate is someone who can independently own data engineering workflows, work with large and complex datasets, troubleshoot production pipeline issues, and collaborate effectively with technical and business stakeholders.
Strong problem-solving, analytical, communication, documentation, and ownership skills are essential.