Source description
About the role
JOB DESCRIPTION:
Job Title: Spark Developer/ Data Consultant
Job Location: London, UK
Job Type: Permanent
Design and Dockerized Ingestion of streaming data from kafka ,spark, Cassandra framework.
ngesting stale data to Cassandra applying framework.
Design and Develop Client Model using SparkML, K-Means, random forest algorithm.
SparkML pipelines workflow String indexer, Hot encoder, Vector.
Designing and writing scala programs and automate the application using apache airflow.
Big Data Migration
Developed Transforming data via Filtering, Cleaning ingested HDFS data using Spark batch jobs.
It includes complete ETL handling.
Extractions →Developed a generic Spark framework for handling various types of input formats like Avro, Bytes, Delta, Delimited.
Transformation →Developed generic spark framework to handle different types of transformation phases including JSON transform, Metadata Filter Transform(SQL based transformation/filtering/, Typing transformation (to centrally convert data-types into expected formats), Formatter transformation (to sanitize data to have similar datatype for e.g. similar date formats, similar JSON structures etc..)
Validation/Loading →Developed framework which performs target specific validation that the dataset has been written correctly and loading into target destination (of various types).
Responsible for delivery end to end lifecycle of Big Data Platform Using open source databases technologies including Apache No SQL Cassandra, Hive, Spark, Jenkins.
Big Data Cloud Integration
Enabling Big Data on AWS IAAS and AWS SAAS hybrid cloud.
Developed/deployed and automated triggers of lambda functions using pyspark scripting to ingest the data into S3.
Migrated on Premises data and queries to AWS using EMR: Used Tech stack of (Spark, Hadoop, YARN, HBase), Athena (S3 Queries)
Job type: Permanent Experience level: 5 Years Reference: 20-00288
More at eTeam