Source description
About the role
About the Team
Citi is looking for a Lead AI/ML Data Scientist to join the Olympus Data Reconciliation and Engineering team, where you will shape the next generation of AI and machine learning capabilities powering enterprise-scale reconciliation across global processing hubs.
In this role, you will drive the full lifecycle of ML model development — from ideation and architecture through to deployment and adoption — delivering measurable impact across Capital Markets operations, risk, and finance. Your work will sit at the intersection of advanced data science and real-world financial systems, influencing outcomes at a global scale.
Responsibilities
-
Design, build, and deploy AI and machine learning models — including Agentic AI and Generative AI solutions — to solve complex reconciliation and data engineering challenges at enterprise scale.
-
Lead the end-to-end ML model development lifecycle, from requirements gathering and data preprocessing through to ensemble modeling, validation, and production integration.
-
Analyze large volumes of structured and unstructured financial data to uncover trends, patterns, and opportunities for optimization across banking platforms.
-
Define and deliver ML model roadmaps in collaboration with technical and business teams, ensuring alignment with project timelines, budgets, and Citi's architecture standards.
-
Translate complex data findings into clear visualizations and strategic recommendations that inform decisions made by senior business and technology leaders.
-
Partner with engineering, operations, and cross-functional teams to ensure seamless model integration, long-term scalability, and reliable performance in production environments.
-
Identify and communicate technology risks and their business implications, developing mitigation strategies and maintaining transparency with stakeholders at all levels.
-
Maintain comprehensive model documentation and support knowledge transfer to ensure continuity and adoption across teams.
-
Required Qualifications & Skills:
-
Technical Expertise: 10+ years hands-on experience in AI/ML development and big data engineering within Financial Services, Insurance, or Telecom environments
-
Expert-level proficiency in Python (scikit-learn, TensorFlow, PyTorch, Pandas, NumPy), R (caret, tidyverse, mlr3), and SQL (PostgreSQL, Oracle, MySQL)
-
Deep technical knowledge implementing supervised and unsupervised ML algorithms: linear/logistic regression, neural networks (CNN, RNN, LSTM, Transformers), k-means clustering, DBSCAN, decision trees (CART, C4.5), and ensemble methods (Random Forest, XGBoost, LightGBM, CatBoost)
-
Proven experience building and deploying Agentic AI and LLM-based solutions using: LangGraph for complex agent orchestration and state management
-
LangChain for chain-of-thought reasoning and retrieval-augmented generation (RAG)
-
Agent Development Kit (ADK) for enterprise-grade autonomous agent development
-
Production-level experience with MLOps frameworks and infrastructure : Apache Airflow for ML pipeline orchestration and workflow automation
-
Kubernetes for containerized model deployment and scaling
-
Docker for reproducible ML environments
-
Advanced proficiency with distributed computing technologies : Apache Spark (PySpark, Spark MLlib) for large-scale data processing
-
Hadoop ecosystem (HDFS, MapReduce, YARN)
-
Apache Hive for data warehousing and SQL-on-Hadoop
-
Expertise with cloud-native data platforms : AWS S3 for scalable data lake storage
-
Amazon Redshift for enterprise data warehousing
-
AWS SageMaker , Azure ML , or Google Vertex AI (beneficial)
-
Strong background in data reconciliation frameworks , data quality validation , and ETL/ELT pipelines for financial data processing at enterprise scale
-
Beneficial Skills & Qualifications: Hands-on experience with advanced statistical modeling: Generalized Linear Models (GLM) , Random Forest , Gradient Boosting (AdaBoost, XGBoost), and Natural Language Processing (NLP) techniques including text mining, topic modeling (LDA), and sentiment analysis
-
Experience with model versioning and experiment tracking tools (Mlflow, Weights & Biases, DVC)
-
Proficiency with Git/GitHub/Bitbucket for version control and collaborative development
-
Knowledge of CI/CD pipelines for ML model deployment (Jenkins, GitLab CI, GitHub Actions)
-
Familiarity with data visualization libraries (Matplotlib, Seaborn, Plotly) and BI tools (Tableau, Power BI)
-
Experience with real-time streaming data frameworks (Kafka, Kinesis)
-
Passion for staying current with emerging AI/ML frameworks, research papers, and open-source contributions
-
Education: Bachelor’s or Master’s degree in Computer Science , Data Science , Software Engineering , Information Systems , Mathematics , Statistics or related fields of study.
Job Family Group: Technology ------------------------------------------------------ Job Family: Data Science ------------------------------------------------------ Time Type: Full time ------------------------------------------------------ Most Relevant Skills Please see the requirements listed above. ------------------------------------------------------ Other Relevant Skills For complementary skills, please see above and/or contact the recruiter. ------------------------------------------------------ Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi .
View Citi’s EEO Policy Statement and the Know Your Rights poster.
More at Citi