Source description
About the role
Title: Data Scientist Duration: 6 month contract to hire Location: Remote Role of the team: • Are there specific projects the team is focused on? o Machine Learning with an angle towards specific use cases – Predicting claims data, can they foresee an outcome and have to do something manual about it later? o Collect data, feature engineering, context of data being used, experimental/research heavy o Setting up system to use large language models to parse a digest of slack messages for the company – Ask that has come out of a few areas, still experimental • Is the team more focused on Building models or interpreting data? o Interpreting data. Goal is to build and make use of models. • How much of what the Data Science team does is self-directed and how much is direction is coming from somewhere else in the company? o Elsewhere in the company Project Information / Candidate Responsibilities: • Is there a specific skill level do you need, are they leading a team of Junior resources? (Junior / Mid / Senior / Lead) o Need to be able to mentor a junior Data Scientist – Be able to be independent (mid to senior Data Science Environment: • Will this role focus more on Data Analysis or Data Analytics? (Analysis focuses on what happened. Analytics focuses on why it happened and what will happen next) o Data Analytics • Can you tell me about the types of Data Science Models that need to be built? o What types of data will be used in the models? ? Structured data ? Would you like the candidate to have experience with NLP Natural Language Processing? (Python, R, etc.) ? Not a strict requirement, but Slack project could make use of it o Will this person need specific tool experience for Analysis and Modeling? (Python, R, PySpark. SAS, R-Studio, Excel, Pandas, GeoPandas, scikit-learn, NumPy, Tableau) ? Python, SQL, PySpark, Databricks • Will this person need to build any data pipelines from relational or NoSQL big data? o Yes, will need to build data pipelines from relational and NoSQL Big Data o Will any Big Data technologies be useful in this role? (Hadoop, HDFS, Hive, Sqoop, Spark, SparkSQL, SparkML, MapReduce, Python, Pyspark, Databricks, or Kafka Streaming) ? PySpark ? Ascend would be a nice to have • Which Cloud environment(s) are you currently using or planning to use? Azure (Microsoft), AWS (Amazon), GCP (Google Cloud Provider)) o AWS – Integration with Azure on the United side – AWS is good, but both is better Data Science Methods and Techniques: • Will the candidate be building any Predictive Modeling to predict user behavior? (Neural Networks, Decision Tree, Logistic Regression, Linear to Multiple Regression Analysis, Segmentation, Clustering, Geospatial Mapping, or Classification) o Claims project will be • What types of Machine Learning are you planning to implement (Supervised, Unsupervised, Semi-Supervised, and/or Reinforcement)? o Supervised • Can you tell me which ML techniques and tools will be most important in this role? o Packages to build, train, and deploy ML models? (TensorFlow, Keras, PyTorch, Scikit-learn or similar packages) Are you open to any tool? o Machine Learning Libraries? (Azure ML, AWS ML, SparkML lib) o ML notebooks? (Jupyter Notebooks, Azure Machine Learning Notebook, R Notebook, or other tools) Likes seeing Github links in profiles • Doesn’t like ‘add files’ – Likes to see actual coding instead of dumping code into projects
More at Thoughtwave Software and Solutions