VuNet Systems
Senior Machine Learning Engineer
Bengaluru, Karnataka
via TheirStack
First seen Oct 8 · seen live today · via TheirStack
Skills mentioned
gokafkamachine learning
The posting, as published
# **Join Our Journey at VuNet**
**VuNet** is a pioneer in *Business Journey Observability*, leveraging Big Data and Machine Learning to transform digital experiences across the financial services. Our deep-tech platform provides end-to-end visibility into customer journeys — empowering proactive issue resolution, operational resilience, and superior user satisfaction.
If you’ve ever used instant payment systems like UPI, chances are you’ve already experienced the power of our platform — we monitor over *28 billion digital transactions monthly*(that’s equal to watching 3 years of tik-tok videos)**,** *touching 400 million users* with leading banks and financial institutions*.*
VuNet is Series B funded, part of NASSCOM DeepTech Club, awarded NASSCOM’s AI Gamechanger, recognized in Forbes DGEMS 200 and by several global analysts including Gartner, Omdia.
We’re building a new category of observability purpose-built for complex digital journeys — across payments, lending, core banking and more — already powering some of the largest banks in India and MEA.
# **Your Role:** **Senior Machine Learning Engineer**
We are looking for a Senior Machine Learning Engineer who combines strong mathematical and statistical foundations with solid software engineering.
This is not primarily a GenAI, prompt-engineering or AI-API integration role.
You will build the underlying intelligence used by our observability platform: statistical models, machine-learning algorithms and analytical techniques that operate continuously on large volumes of time-series and operational data.
*We are looking for someone who can reason from first principles, formulate a problem mathematically, evaluate alternative approaches, understand their assumptions and failure modes, implement the solution and take it all the way into reliable production.*
You will work closely with platform engineering, SRE, product and domain teams, but will be expected to independently identify opportunities where better algorithms can materially improve the product.
# **Roles & Responsibilities**
ML Platform & Production Engineering
- Enhance and productionize VuNet's ML/MLOps platform for building, orchestrating, deploying, validating and monitoring ML workloads.
- Work with Temporal-based workflows and VuNet's ML frameworks for model execution, scheduling and lifecycle management.
- Build mechanisms for model versioning, experimentation, benchmarking, validation, safe rollout and performance monitoring.
- Engineer ML capabilities for continuous production workloads across thousands to tens of thousands of time series and monitored entities.
- Design for high-throughput streaming data, distributed computation, low-latency inference, horizontal scalability and resource efficiency.
Anomaly Detection & Behavioural Intelligence
- Develop adaptive anomaly-detection algorithms across infrastructure, application, transaction and business time-series data.
- Build techniques that automatically account for seasonality, trends, changing baselines, noise and workload patterns.
- Explore change-point detection and distribution-shift techniques to identify meaningful behavioural changes.
- Explore multivariate approaches where anomalies emerge from relationships among signals rather than from individual metrics.
Forecasting & Predictive Operations
- Develop forecasting algorithms for capacity planning, workload forecasting, resource exhaustion and proactive operations.
- Build models capable of handling multiple seasonalities, incomplete/noisy telemetry and changing operational behaviour.
- Quantify prediction uncertainty and identify leading indicators that provide early warning of degradation or saturation.
Root Cause, Dependency and Incident Intelligence
- Develop analytical and ML techniques that distinguish probable root causes from downstream symptoms.
- Use service topology, dependencies, temporal relationships, behavioural correlation and statistical evidence for RCA.
- Correlate alerts, anomalies and operational events into meaningful incidents using time, topology, entity relationships and behavioural similarity.
- Develop techniques for event clustering, incident evolution, impact analysis, blast-radius identification and probable-cause ranking.
- Explore causal inference and dependency-aware approaches where they provide measurable value.
Algorithm Validation, Simulation & Quality
- Build systematic benchmarking and validation frameworks for anomaly detection, forecasting, correlation and RCA algorithms.
- Define evaluation measures including precision, recall, false-positive rate, detection delay, stability, forecasting error and operational usefulness.
- Evaluate algorithms under seasonality, concept drift, noisy or missing telemetry, cold-start conditions and changing workloads.
- Develop simulation and synthetic-data frameworks for realistic metrics, anomalies, incidents, failures and dependency scenarios.
- Validate algorithms against labelled datasets, historical production data and controlled or simulated scenarios to prevent regressions.
# **What You Bring**
### **Mandatory Skills**
- Strong software engineering skills, particularly in Python.
- Strong grounding in probability, statistics, statistical inference and machine learning.
- Good understanding of time-series analysis including trends, seasonality, decomposition, forecasting and change detection.
- Ability to reason about algorithms beyond library APIs: assumptions, trade-offs, computational complexity, failure modes and appropriate evaluation methods.
- Experience building or materially adapting ML/statistical algorithms rather than only integrating pre-built AI services.
- Experience taking algorithms from experimentation into reliable production systems.
- Strong debugging and analytical problem-solving skills.
- Understanding of distributed systems and high-volume data processing.
- Ability to work with imperfect, noisy and evolving real-world datasets.
- *High agency: identifies meaningful problems, forms hypotheses, prototypes solutions, validates them against data and drives successful approaches into production without waiting for detailed task definitions.*
**Good to Have Skills**
- Go experience for production services or high-performance components.
- Experience with Apache Flink or similar streaming/data-processing frameworks.
- Experience working with observability, monitoring, SRE, telemetry or AIOps systems.
- Familiarity with metrics, logs, traces, service topology and operational event data.
- Experience with Kafka or other streaming platforms.
- Exposure to techniques such as Bayesian methods, clustering, dimensionality reduction, probabilistic models, causal inference or optimization.
- Experience building large-scale simulation, benchmarking or synthetic-data frameworks.
- Experience working on systems where false positives, model drift, latency and explainability have direct operational consequences.
# **What We Offer**
### **Life at VuNet: Building the Future Together**
At VuNet, we’re building a world-class observability platform, proudly **Made in India —** and we're just getting started.
We’re a team of passionate problem-solvers who love tackling complex challenges. We learn fast, adapt quickly, and stay curious — especially when it comes to exploring and staying ahead of the curve with emerging technologies likeGen AI**.**
More than just a tech company, VuNet is a place where collaboration, learning, and innovation are part of everyday life. We believe in working together, taking ownership, and growing as a team.
If you’re looking to work on cutting-edge technology, make a real impact, and grow with a supportive team — you’ll feel right at home at **VuNet**.
# **Benefits For You**
- Health insurance coverage for you, your parents, and dependents.
- Mental wellness and 1:1 counselling support.
- A learning culture that promotes growth, innovation, and ownership.
- Transparent, Inclusive