Source description
About the role
As a Data Engineer, you will be responsible for owning the architecture and maintenance of data pipelines. This will involve orchestrating complex workflows using Apache Airflow to ensure seamless data flow from various on-site sources into the central data warehouse. Your role will be crucial in transforming raw construction data into actionable insights that drive project timelines and efficiency. Key Responsibilities: - Design, build, and maintain robust data pipelines using Apache Airflow. - Write and optimize Directed Acyclic Graphs (DAGs) to automate ETL processes. - Develop and manage end-to-end ETL (Extract, Transform, Load) lifecycles. - Ensure accurate extraction of data from various APIs and databases, followed by transformation based on business logic, and loading into the data warehouse. - Maintain and enhance data architecture to support high-volume data ingestion and processing. - Implement monitoring and alerting for data pipelines to ensure high availability and accuracy. - Troubleshoot and resolve data issues promptly. - Work with AWS services (S3, Redshift, Glue, Lambda) to build scalable cloud-based data solutions if applicable. - Collaborate closely with construction tech team, data analysts, and software engineers to understand data requirements for Digital Twin and operational dashboards. Qualifications Required: - Minimum of 3 years of professional experience in Data Engineering. - Strong hands-on experience with Apache Airflow, including writing custom operators, managing DAG dependencies, and troubleshooting scheduling issues. - Proven track record of building and optimizing ETL/ELT pipelines. - Strong proficiency in Python for scripting and Airflow, as well as SQL for database querying and transformation. - Experience with relational databases such as PostgreSQL, MySQL, and Data Warehousing concepts. Additional Details: Experience with AWS ecosystem, specifically S3, Redshift, Glue, or EMR, and familiarity with containerization tools like Docker or Kubernetes will be a plus. Familiarity with Real Estate/Construction domain data or IoT data streams is also advantageous. As a Data Engineer, you will be responsible for owning the architecture and maintenance of data pipelines. This will involve orchestrating complex workflows using Apache Airflow to ensure seamless data flow from various on-site sources into the central data warehouse. Your role will be crucial in transforming raw construction data into actionable insights that drive project timelines and efficiency. Key Responsibilities: - Design, build, and maintain robust data pipelines using Apache Airflow. - Write and optimize Directed Acyclic Graphs (DAGs) to automate ETL processes. - Develop and manage end-to-end ETL (Extract, Transform, Load) lifecycles. - Ensure accurate extraction of data from various APIs and databases, followed by transformation based on business logic, and loading into the data warehouse. - Maintain and enhance data architecture to support high-volume data ingestion and processing. - Implement monitoring and alerting for data pipelines to ensure high availability and accuracy. - Troubleshoot and resolve data issues promptly. - Work with AWS services (S3, Redshift, Glue, Lambda) to build scalable cloud-based data solutions if applicable. - Collaborate closely with construction tech team, data analysts, and software engineers to understand data requirements for Digital Twin and operational dashboards. Qualifications Required: - Minimum of 3 years of professional experience in Data Engineering. - Strong hands-on experience with Apache Airflow, including writing custom operators, managing DAG dependencies, and troubleshooting scheduling issues. - Proven track record of building and optimizing ETL/ELT pipelines. - Strong proficiency in Python for scripting and Airflow, as well as SQL for database querying and transformation. - Experience with relational databases such as PostgreSQL, MySQL, and Data Warehousing concepts. Additional Details: Experience with AWS ecosystem, specifically S3, Redshift, Glue, or EMR, and familiarity with containerization tools like Docker or Kubernetes will be a plus. Familiarity with Real Estate/Construction domain data or IoT data streams is also advantageous.
More at Neemtree