A2 Global Consulting
Senior Machine Learning Operations Engineer
Chandigarh · remote
via TheirStack
First seen Sep 28 · seen live today · via TheirStack
Skills mentioned
pythonflaskfastapirestawsdockerkubernetesterraformtensorflowpytorchmachine learningdata scienceci/cd
The posting, as published
**Position Overview**
We are looking for a skilled Senior MLOps Engineer to join our growing efforts in machine learning and AI engineering. In this role, youll collaborate with Data Scientists, Machine Learning Engineers, Legal Knowledge Experts, Developers, and other MLOps Engineers to design, build, and maintain robust production machine learning infrastructure and tooling. You will play a key role in enabling scalable, efficient, and reliable delivery of customer facing machine learning solutions.
**Job Responsibilities**
- Collaborate with cross-functional teams to build and maintain end-to-end machine learning pipelines, from data ingestion to model deployment and monitoring
- Design, implement, and optimize infrastructure for rapid prototyping, continuous integration, deployment, and model evaluation.
- Monitor and maintain production machine learning systems to ensure reliability, scalability, and performance.
- Build and support production-grade AI engineering workflows, including LLM applications, retrieval-augmented generation, agentic systems, prompt orchestration, and evaluation frameworks.
- Design and maintain agentic systems integrations using emerging interoperability standards and patterns such as Model Context Protocol (MCP) and Agent-to-Agent (A2A) communication.
- Develop and maintain infrastructure as code to provision, configure, and manage cloud-based machine learning and AI infrastructure reliably and repeatably
- Provide technical guidance and mentorship to junior team members and foster knowledge sharing within the MLOps and Data Science teams
- Stay updated with the latest advancements in MLOps, cloud technologies, and AI engineering to identify and implement best practices
- Set and enforce standards for code quality and best practices across the data science and engineering organizations to ensure maintainability, scalability, and robustness of systems.
- Other duties as assigned
**A little bit about you**
- Proven experience deploying and maintaining machine learning models in production environments
- Demonstrated ability to gather requirements, design systems, and scope and plan projects effectively, with a focus on the entire machine learning lifecycle
- Experience building or supporting production generative AI or LLM-based applications, including workflows such as RAG, prompt management, evaluation, or agent orchestration
- Familiarity with MCP, A2A, or related patterns for connecting AI systems, tools, services, and agents in secure and maintainable ways
- Experience with LLM observability concepts and tools, including tracing, evaluation, monitoring, cost tracking, and quality measurement
- Experience with REST API design and implementation, preferably using frameworks such as Flask or FastAPI
- Proficiency in Python, including both general-purpose programming and machine learning frameworks such as scikit-learn, TensorFlow, PyTorch, or similar
- Experience in building and scaling machine learning pipelines and infrastructure, including data gathering, feature engineering, model training, and deployment workflows
- Proficiency with cloud platforms, preferably AWS, and tools like SageMaker, Lambda, or similar
- Experience with IaC tools such as Terraform, AWS CDK, CloudFormation or similar
- Strong communication skills to effectively collaborate with diverse teams, including product managers, engineers, and data scientists
- Familiarity with CI/CD tools and practices as they apply to machine learning workflows
- Experience with modern containerization tools such as Docker and Kubernetes