Padmi

Data Engineer

MumbaiPosted 2 months ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at ETCIO - Cloud Data Center

Opens the source posting on shine.com

Source description

About the role

View original

As an insight-driven Data Engineer at Proximity Works, you will play a crucial role in building and scaling data pipelines and core analytical datasets that power analytics, AI model evaluation, safety systems, and business decision-making across Bharat AI's agentic AI platform. Your work will directly influence how AI systems learn, adapt, and improve, as well as how leaders make product and business decisions. You will collaborate closely with various teams such as Product, Data Science, Infrastructure, Marketing, Finance, and AI/Research to ensure data reliability, interpretability, and outcome orientation as the platform scales rapidly. Responsibilities: - Design, build, and manage scalable data pipelines to ensure high-fidelity user event and system data ingestion - Develop and maintain analytics-ready datasets for tracking key product and business metrics - Structure data to enable insights, experimentation, and feedback loops for improving models and product decisions - Collaborate with cross-functional teams to translate business questions into trustworthy datasets - Implement robust systems for batch and streaming data processing - Actively participate in data architecture decisions and ensure data security, integrity, and compliance - Monitor and troubleshoot pipeline and streaming job health to improve reliability, performance, and data quality What Matters (Non-Negotiable): - Thinks in architecture, feedback loops, and outcomes - Ability to structure data models that improve AI model evaluation, decision-making, and learning velocity - Communicates in terms of impact and insights Requirements: - 3-5 years of professional experience as a Data Engineer or in a similar role - Strong proficiency in Python for data processing and orchestration - Hands-on experience with Apache Spark and distributed data processing systems - Understanding of data pipeline design for analytics, reporting, and ML workflows - Strong problem-solving skills and the ability to work across teams with varied data requirements Desired Skills: - Experience working with Databricks in production environments - Familiarity with the GCP data stack and data quality frameworks - Exposure to data validation, schema management tools, and analytics use cases - Strong Plus: GCP Dataflow, BigQuery, Google Analytics, experience designing datasets for experimentation or ML feedback loops Benefits: - Best in class compensation - Proximity Talks for learning from experienced professionals - High-impact work directly influencing AI models and business decisions - Continuous learning in a collaborative team environment At Proximity Works, we are a global team of coders, designers, product managers, and experts dedicated to solving complex problems and building cutting-edge technology at scale. Your impact on our success will be significant, and you will have the opportunity to work with experienced leaders in the tech, data, and product fields. About Us: - Proximity Works is an AI engineering company headquartered in San Francisco - We build production AI systems for the world's largest sports, media, and entertainment platforms - Six years in production, reaching over 200 countries In conclusion, as a Data Engineer at Proximity Works, you will have the opportunity to work on high-impact projects, collaborate with diverse teams, and continuously grow your data engineering skills in a supportive environment. As an insight-driven Data Engineer at Proximity Works, you will play a crucial role in building and scaling data pipelines and core analytical datasets that power analytics, AI model evaluation, safety systems, and business decision-making across Bharat AI's agentic AI platform. Your work will directly influence how AI systems learn, adapt, and improve, as well as how leaders make product and business decisions. You will collaborate closely with various teams such as Product, Data Science, Infrastructure, Marketing, Finance, and AI/Research to ensure data reliability, interpretability, and outcome orientation as the platform scales rapidly. Responsibilities: - Design, build, and manage scalable data pipelines to ensure high-fidelity user event and system data ingestion - Develop and maintain analytics-ready datasets for tracking key product and business metrics - Structure data to enable insights, experimentation, and feedback loops for improving models and product decisions - Collaborate with cross-functional teams to translate business questions into trustworthy datasets - Implement robust systems for batch and streaming data processing - Actively participate in data architecture decisions and ensure data security, integrity, and compliance - Monitor and troubleshoot pipeline and streaming job health to improve reliability, performance, and data quality What Matters (Non-Negotiable): - Thinks in architecture, feedback loops, and outcomes - Ability to structure data models that improve AI model evaluati

One address, no account. We’ll tell you when matching roles go live.

More at ETCIO - Cloud Data Center

Related open roles

View all roles