Source description
About the role
Responsibilities: Create and extend software frameworks and tools to run, monitor, and report on automated tests. Have experience building and maintaining Automation frameworks and CI/CD pipeline.Review testing strategies, approaches, and test plans and provide insight to development teams.Architect, design and build software to automate a variety of testing methods (e. g. API validation, performance/load testing, fuzzing, etc. )Develop and maintain test frameworks for GenAI-based and AI/ML applications, including model validation, LLM prompt-response consistency, hallucination detection, and regression testing for AI inference APIs.Implement evaluation pipelines to test LLM accuracy, latency, bias, toxicity, and reliability.Work as part of a team of highly skilled software professionals.Continuously learn and expand your technical horizonsOur Technology Stack: Selenium, Playwright, Docker, Kubernetes, AWS (Bedrock, Lambda, SQS, S3 etc. ), LangChain, LLM evaluation frameworks (e. g., Ragas), OpenAPI, Postgres, Linux. Requirements: Are excited to work at the intersection of AIML, Big Data and the Cybersecurity problem space.Have 5+ years of experience shipping quality software.Love driving high software quality.Have experience deploying services on cloud computing platforms (e. g, AWS, Azure, GCP).Are an expert using common tools for executing functional, load and fuzz testing (e. g. Locust, CATS).Have worked on distributed systems and microservices architecture (preferred). Responsibilities: Create and extend software frameworks and tools to run, monitor, and report on automated tests. Have experience building and maintaining Automation frameworks and CI/CD pipeline.Review testing strategies, approaches, and test plans and provide insight to development teams.Architect, design and build software to automate a variety of testing methods (e. g. API validation, performance/load testing, fuzzing, etc. )Develop and maintain test frameworks for GenAI-based and AI/ML applications, including model validation, LLM prompt-response consistency, hallucination detection, and regression testing for AI inference APIs.Implement evaluation pipelines to test LLM accuracy, latency, bias, toxicity, and reliability.Work as part of a team of highly skilled software professionals.Continuously learn and expand your technical horizonsOur Technology Stack: Selenium, Playwright, Docker, Kubernetes, AWS (Bedrock, Lambda, SQS, S3 etc. ), LangChain, LLM evaluation frameworks (e. g., Ragas), OpenAPI, Postgres, Linux. Requirements: Are excited to work at the intersection of AIML, Big Data and the Cybersecurity problem space.Have 5+ years of experience shipping quality software.Love driving high software quality.Have experience deploying services on cloud computing platforms (e. g, AWS, Azure, GCP).Are an expert using common tools for executing functional, load and fuzz testing (e. g. Locust, CATS).Have worked on distributed systems and microservices architecture (preferred).
More at Arctic Wolf