Source description
About the role
Senior Data Analyst (Machine Learning) role:
Interview Format: In-person interview only (no video interviews). Candidates should be prepared for live SQL and PySpark problem-solving, as many candidates struggle with hands-on coding in person.- PLEASE MAKE SURE YOUR CANDIDATES ARE LOCAL AND CAN COME ONSITE.
Required Certifications: Databricks Data Engineer certification is MOST IMPORTANT but can be completed post-onboarding. Google Cloud certification is also acceptable post-hire.
SQL: THIS IS HER TOP TOP MUST HAVE!! Must be able to write and execute SQL queries independently. Strong hands-on ability is required, including data extraction and basic to intermediate transformations.
PySpark (Highest Priority): Core requirement of the role. Candidates must have strong hands-on experience building data pipelines and performing large-scale data processing.
Python / Pandas: Required for data cleaning, transformation, and analysis.
Machine Learning (20% of role): Candidates must already have hands-on ML experience (classification, regression, clustering, model evaluation). No training will be provided on ML concepts.
Data Analysis Focus: Majority of the role involves SQL-based data extraction and PySpark/Pandas-driven analysis and insights
ML + Analytics Integration: Candidates should understand end-to-end workflows combining SQL, PySpark, and ML.
Team Structure: Approximately 10 Data Engineers and 6 Data Scientists within the team.
Please ensure candidates are screened heavily for hands-on PySpark, PANDA, SQL coding ability, and real ML implementation experience before submission.
Must Have
Applied Machine Learning
Azure Databricks
Big Data Analytics
Databricks Certified Data Engineer Associate
Data Structures
google cloud certified machine learning engineer
Machine Learning Operations
Pandas Python Library
PySpark
LOCALS ONLY PLEASE
Job Title - Data Analyst Senior-MUST HAVE STRONG MACHINE LEARNING
Job Location - 4 Irving Place, New York, New York, United States of America, 10003
Work Model: Hybrid (Onsite presence required)- MUST ALSO DO IN PERSON INTERVIEW
Experience Level: 7+ years
Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, Statistics, Mathematics, or related field.
Strong experience in machine learning algorithms, predictive modeling, and data mining.
Proficiency in Pyspark, Python pandas (required) for data science workloads.
Strong SQL (required) knowledge and experience with relational databases.
Minimum 3 years of experience with data visualization tools such as Power BI, Dax Queries, and best practices.
Experience with Azure Databricks, Google Cloud, and modern data science libraries (e.g., scikit-learn, pandas, NumPy).
Experience with GenAI and large language models.
Ability to interpret complex datasets and produce actionable insights.
Must know how to analyze the root cause of dashboard errors.
Have experience in ML Ops and have strong coding background.
Have experience with Natural Language Processing (NLP).
Knowledge or experience with A/B Testing.
Working knowledge of designing, training, and implementing machine learning models.
Familiarity with cloud-based infrastructure
Excellent communication and problem-solving skills.
7 or more years of experience in data science and machine learning engineering.
Additional Skills (Skills that are a plus, but not required)
Knowledge of statistical methods and experimental design.
Responsibilities
-
Key Responsibilities
-
Advanced Analytics & Machine Learning
-
Design, develop, and optimize machine learning models (forecasting, classification, clustering).
-
Apply data mining techniques to uncover patterns and insights in large datasets.
-
Perform feature engineering, model validation, and performance tuning.
-
Explore and deploy modern AI and ML approaches to enhance automation and analytics.
-
Data Preparation & Quality
-
Prepare structured and unstructured data for modeling and advanced analysis.
-
Develop scripts and tools for data cleansing, validation, and enrichment.
-
Collaborate with Data Engineering to maintain efficient data pipelines.
-
Identify data quality issues and propose remediation.
-
Analytics, Insights & Reporting
-
Conduct deep-dive analyses to identify trends and improvement opportunities.
-
Communicate complex findings in clear, concise ways to technical and non-technical stakeholders.
-
Support the development of dashboards, metrics, and analytical solutions.
-
Cross-Team Collaboration
-
Work with architects, engineers, and analysts to define analytical requirements.
-
Contribute to conceptual data model design and workflow optimization.
-
Promote best practices in machine learning, analytics, and data governance.
More at 3B Staffing