Padmi

Data Engineer - Python/SQL

Delhi NCRPosted 3 months ago
Software engineeringSeniorFull Time; Regular
Apply at Cantik Technologies

Opens the source posting on shine.com

Source description

About the role

View original

As a Data Engineer, you will be responsible for designing, building, and operating secure, reliable, and cost-efficient data pipelines on AWS. Your role will cover batch, streaming, and CDC ingestion through to curated datasets ready for analytics. In addition, you will embed testing, governance, observability, and CI/CD practices across the data lifecycle. Key Responsibilities: - Implement pipelines using AWS Lambda, EMR for batch/streaming; manage schema evolution and backfills. - Build workflows with Step Functions and/or MWAA (Airflow) including retries, alerts, and SLAs. - Model datasets on S3 (Parquet/Iceberg), Glue Data Catalog, Athena; handle partitioning, compaction, and performance tuning. - Add automated tests (Great Expectations/Deequ, PySpark unit tests), data contracts, and CI gates for data quality and testing. - Implement IAM, KMS, VPC endpoints, Lake Formation policies, PII handling, audit trails (CloudTrail), and RBAC/RLS for security and governance. - Set up CloudWatch metrics/alerts, lineage, usage dashboards; optimize costs through S3 lifecycle, compression, and job sizing for observability and FinOps. - Provision resources with Terraform/CloudFormation/CDK, build pipelines with GitHub Actions, and manage environment promotion and rollbacks for CI/CD and IaC. - Maintain pipeline diagrams, SLAs/SLOs, and incident playbooks for documentation and runbooks. - Be able to adopt capabilities from other teams instead of building duplicate capabilities and contribute back to the community. Outcomes (first 60-90 days): - Productionize one end-to-end pipeline with tests, monitoring, and alerting. - Establish governance baseline (Lake Formation, tagging/classification, encryption) and data contracts for two key sources. - Implement CI/CD and IaC for data services; reduce at least one cost driver via storage/compute optimization. Skills and Experience Required: - 8+ years of experience building data platforms on AWS; proficiency in Python (including PySpark) and SQL. - Hands-on experience with DMS, Kinesis/MSK, Glue/EMR, Lambda, Step Functions/MWAA, S3/Parquet/Iceberg, Glue Catalog, Athena, Redshift/Serverless. - Expertise in IaC (Terraform) and CI/CD (GitHub Actions). - Familiarity with data quality frameworks (Great Expectations/Deequ), schema evolution, and backfills. - Knowledge of security and compliance practices such as IAM/KMS, private networking, Lake Formation, audit/lineage. - Experience in performance and cost tuning across storage and compute. Nice to Have: - Familiarity with Snowflake, Apache Flink, Redshift Spectrum, OpenLineage, OPA/policy-as-code. As a Data Engineer, you will be responsible for designing, building, and operating secure, reliable, and cost-efficient data pipelines on AWS. Your role will cover batch, streaming, and CDC ingestion through to curated datasets ready for analytics. In addition, you will embed testing, governance, observability, and CI/CD practices across the data lifecycle. Key Responsibilities: - Implement pipelines using AWS Lambda, EMR for batch/streaming; manage schema evolution and backfills. - Build workflows with Step Functions and/or MWAA (Airflow) including retries, alerts, and SLAs. - Model datasets on S3 (Parquet/Iceberg), Glue Data Catalog, Athena; handle partitioning, compaction, and performance tuning. - Add automated tests (Great Expectations/Deequ, PySpark unit tests), data contracts, and CI gates for data quality and testing. - Implement IAM, KMS, VPC endpoints, Lake Formation policies, PII handling, audit trails (CloudTrail), and RBAC/RLS for security and governance. - Set up CloudWatch metrics/alerts, lineage, usage dashboards; optimize costs through S3 lifecycle, compression, and job sizing for observability and FinOps. - Provision resources with Terraform/CloudFormation/CDK, build pipelines with GitHub Actions, and manage environment promotion and rollbacks for CI/CD and IaC. - Maintain pipeline diagrams, SLAs/SLOs, and incident playbooks for documentation and runbooks. - Be able to adopt capabilities from other teams instead of building duplicate capabilities and contribute back to the community. Outcomes (first 60-90 days): - Productionize one end-to-end pipeline with tests, monitoring, and alerting. - Establish governance baseline (Lake Formation, tagging/classification, encryption) and data contracts for two key sources. - Implement CI/CD and IaC for data services; reduce at least one cost driver via storage/compute optimization. Skills and Experience Required: - 8+ years of experience building data platforms on AWS; proficiency in Python (including PySpark) and SQL. - Hands-on experience with DMS, Kinesis/MSK, Glue/EMR, Lambda, Step Functions/MWAA, S3/Parquet/Iceberg, Glue Catalog, Athena, Redshift/Serverless. - Expertise in IaC (Terraform) and CI/CD (GitHub Actions). - Familiarity with data quality frameworks (Great Expectations/Deequ), schema evolution, and backfills. - Knowledge of security and compliance practices such

One address, no account. We’ll tell you when matching roles go live.

More at Cantik Technologies

Related open roles

View all roles