Source description
About the role
\nAnalyze, manipulate, and process large sets of structured and semi-structured data including healthcare datasets (claims, patient, and provider data) using Python (PySpark), SQL, and Spark SQL within the Azure data platform.\nApply data mining, data modeling, and statistical analysis techniques to extract and analyze information from large datasets stored in Azure Data Lake Storage (ADLS) and Delta Lake across Bronze, Silver, and Gold layers.\nIdentify business problems and management objectives that can be addressed through data analysis; translate business requirements into analytical workflows and data solutions using Azure Data Factory and Azure Databricks.\nDevelop and maintain data models and analytical datasets to support reporting, business intelligence, and downstream data consumption; identify relationships, trends, and factors that could affect the results of analysis.\nApply feature selection and data profiling methods to identify patterns, anomalies, and data quality issues that may impact analytical outcomes; recommend data-driven improvements.\nClean, manipulate, and prepare raw data for analysis using ETL/ELT processes; implement data quality frameworks including validation rules, cleansing, and reconciliation techniques to ensure accuracy and consistency.\nWrite new functions and applications in Python (PySpark) and SQL to conduct data transformation, validation, and analysis of large-scale datasets; optimize code for performance and efficiency.\nTest, validate, and reformulate data pipelines and datasets to ensure accurate and reliable data delivery; monitor, troubleshoot, and resolve data and performance issues using Azure monitoring tools.\nAnalyze data patterns and recommend data-driven solutions to key stakeholders; deliver oral and written presentations of data analysis findings to management and end users to support informed decision-making.\nSupport data governance, security, and compliance practices aligned with HIPAA requirements; prepare documentation for data processes and analytical workflows to support governance and audits.\nDevelop and manage workflow orchestration, scheduling, and automation for consistent and timely data delivery using Azure-based tools.\nImplement version control and deployment processes using Azure DevOps and Git to support reproducible and reliable data workflows.\nIntegrate data from multiple sources including APIs, databases, and flat files into a unified, analytics-ready data platform.\nRead technical articles, research publications, and conference papers to identify emerging analytic trends and technologies; continuously evaluate new tools and methodologies to improve data processing, analysis, and platform efficiency.\nCollaborate with business stakeholders, analysts, and data scientists to gather requirements, advise on appropriate analytical techniques, and recommend data-driven solutions aligned with organizational objectives.\n
Labor Condition Application (ETA Form 9035 & 9035E) filed with the United States Department of Labor, Office of Foreign Labor Certification for a H-1B non-immigrant worker with the following details:\n1. Number of H1B non-immigrant workers included in LCA: \n1(One)\n2. Job Position: \nAzure Data Engineer\n3. Wages Offered: \n$83221.00 - $83222.00 per year\n4. Period of Employment: \n04/06/2026 to 04/05/2029\n5. Locations where H1B non-immigrant worker will work:\n1) Louisville, Kentucky - 40202\nComplaints alleging misrepresentation of material facts in the Labor Condition Application and/ or failure to comply with the terms of the Labor Condition Application may be filed with any office of the Wage & Hour Division of the United States Department of Labor.\nThe verification of Labor Condition Application (ETA Form 9035 & 9035E) will be available for all to review. \nhttps://tinyurl.com/yrvtrhct\n
More at Ventois