Source description
About the role
Job description Key Responsibilities Design, develop, and maintain large-scale data processing systems using Spark, Hive, and SQL. Lead end-to-end data pipeline development (ingestion, transformation, validation, and delivery). Collaborate with business stakeholders, analysts, and downstream teams to understand data requirements. Ensure performance optimization of queries and ETL jobs. Review and enforce coding standards, design patterns, and best practices within the team. Troubleshoot and resolve complex production issues in data workflows. Automate workflows using Shell scripting and scheduling tools. Drive data quality initiatives including validation, monitoring, and alerting frameworks. Guide and mentor junior engineers perform code reviews and technical evaluations. Participate in architecture discussions and contribute to technical/PI roadmap planning Education Minimum bachelors degree in computer science is required, masters degree preferred Experience (In Years) Advanced knowledge of the principles of computer application including software engineering, system analysis and design, design and digital systems, computer system architecture, database management system, and problem solving and computer programming, among others. MIn 13 + years of software design and development experience Core hands on development skills in Bigdata ecosystem i.e. Spark Scala, SOLR, Pig, Hive, Kafka and HBase. Minimum 4+ years in a Tech Lead or senior technical role Proven experience handling large-scale, distributed data systems Optimize and tune queries, job flows on Hadoop environments to meet performance requirements. Good understanding of release management, CI/CD Work with Bigdata developers designing scalable supportable Application development. Programming using Shell script/Python to create automation scripts. Operate in various development environments (Agile, Waterfall, etc.) while collaborating with key stakeholders Responsible for developing and maintaining the operate runbooks. Preferred: 4+ years of experience in deploying Hadoop components Hive, Spark, SQL,shellscript Strong communication skills to lead interactions with business and technology leaders across the global lines of businesses. 3 to 5 years of functional experience in Insurance industry Cloud experience is plus with Azure, Experience working in cross-functional, multi-location teams. Excellent analytical and problem-solving skills Technical Skills Mandatory: Strong hands-on experience with: Apache Spark (Core & SQL) Hive / SQL Scala Shell Scripting Deep understanding of: Data warehousing concepts ETL/ELT pipeline design Big data ecosystems (Hadoop, HDFS) Experience in performance tuning of Spark jobs and Hive queries Familiarity with workflow schedulers (IBM Maestro, Airflow, Oozie, etc.) Preferred: Exposure to cloud platforms (Azure / AWS / GCP) & AI Tools. Knowledge of data lake architecture and modern data platforms Experience in version control (Git) and CI/CD pipelines Understanding of data governance and security practices Other Critical Requirements Like Voice/ Non-Voice for Insurance Ops Non-Voice (primarily backend/data engineering role) Strong analytical and problem-solving skills Excellent collaboration and stakeholder communication skills Ability to work in a fast-paced, agile environment Ownership mindset with focus on delivery and quality
More at Virtuoso Staffing Solutions