Responsibilities:
* Develop and deploy end-to-end microservices-based solutions for batch and real-time algorithms, including monitoring, logging, automated testing, and performance testing.
* Design, implement, and optimize MLOps pipelines using tools such as Kubeflow, Seldon, MLFlow, Docker, and Kubernetes.
* Collaborate with Data Scientists to enhance the ML model development process and ensure performance improvements.
* Ensure scalability, maintainability, and robustness of deployed machine learning models.
* Monitor and troubleshoot ML model performance and infrastructure issues in production (experience with Prometheus and Grafana is valuable).
* Support and enhance ML software infrastructure, including CI/CD, data storage, cloud services, security, and system monitoring.
* Work with cloud platforms, particularly GCP and Azure, to optimize resource allocation and costs.
* Stay up to date with the latest trends and best practices in MLOps.
Qualifications:
* Bachelor's or Master’s degree in Computer Science, Engineering, or a related field.
* 5+ years of experience as a Machine Learning Engineer or in a similar role.
* Proficiency in Python and experience with ML frameworks like TensorFlow, PyTorch, and scikit-learn.
* Strong understanding of MLOps best practices and tools, including Kubeflow, Seldon, MLFlow, Docker, and Kubernetes.
* Experience working with cloud platforms, especially GCP.
* Knowledge of data processing, ETL, and feature engineering techniques.
* Strong problem-solving skills and ability to work in a fast-paced, collaborative environment.
* Excellent communication and interpersonal skills.