
Lead Software Engineer, Cloud Site Reliability (SRE)
Jobgether • India
No Relocation
Posted: September 24, 2026
Additional Content
Job Description
- This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Software Engineer, Cloud Site Reliability (SRE) based in India. This role leads cloud reliability and 24x7 site reliability operations for critical technology environments, with a strong focus on Azure infrastructure and cloud-native platforms. You will take ownership of major incidents, drive operational excellence, and ensure high availability and SLA adherence across production systems. The position combines hands-on cloud engineering with observability, automation, incident management, and reliability improvements. You will work extensively with Azure, AKS, Kubernetes, Docker, Datadog, and infrastructure-as-code technologies to build resilient and scalable environments. The role also provides an opportunity to advance proactive monitoring, anomaly detection, AIOps, and self-healing capabilities. As a technical leader, you will mentor engineers, collaborate with development teams, and communicate operational performance to stakeholders and leadership.
- Accountabilities: Lead 24x7 NOC and site reliability operations through rotational shifts, ensuring system availability, operational stability, and adherence to service-level agreements. Act as Major Incident Manager for P1 and P2 incidents, coordinating triage activities, war rooms, technical teams, and stakeholder communications through resolution. Manage and troubleshoot Azure infrastructure, including virtual machines, networking, storage, and related cloud services. Administer Azure Kubernetes Service (AKS), Kubernetes, and Docker environments, including scaling, troubleshooting, performance optimization, and reliability improvements. Establish and enhance observability practices across logs, metrics, and traces using platforms such as Datadog and Azure Monitor. Drive proactive monitoring, alert optimization, anomaly detection, and AIOps initiatives to identify and address potential reliability issues before they impact users. Build automation and self-healing workflows using Terraform, ARM templates, Helm, Power Automate, PowerShell, Python, Bash, and other appropriate technologies. Collaborate with engineering teams to strengthen deployment pipelines, improve system reliability, and advance cloud-native architecture and operational practices. Develop operational dashboards and reports using Power BI and ServiceNow to provide visibility into incidents, reliability, and service performance. Lead monthly business reviews and provide clear operational reporting and insights to leadership and stakeholders. Mentor team members, promote knowledge sharing, standardize operational processes, and drive continuous improvements across the reliability function. Support broader initiatives involving multi-cloud environments, predictive monitoring, and self-healing systems where appropriate. Requirements 7–12 years of professional experience in CloudOps, Site Reliability Engineering, NOC, or comparable 24x7 operations environments. Strong hands-on expertise with Azure infrastructure, particularly virtual machines, networking, storage, and related IaaS services. Extensive experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker, including troubleshooting, scaling, and performance tuning. Strong experience with monitoring and observability platforms such as Datadog and Azure Monitor, including integrations, alerting, dashboards, and operational analysis. Proven experience in incident management and major incident handling, including P1/P2 coordination, stakeholder communication, root-cause analysis, and operational reporting. Experience with Infrastructure as Code technologies such as Terraform, ARM templates, and Helm. Strong scripting capabilities using PowerShell, Python, Bash, or similar automation technologies. Experience working with ServiceNow, particularly Incident, Problem, and Change Management modules and associated dashboards. Good understanding of distributed systems, cloud-native architectures, and reliability engineering principles. Excellent communication, leadership, coordination, and problem-solving skills, with the ability to work effectively across technical and business teams. Experience working in multi-cloud environments, particularly Azure and AWS, is a plus. Exposure to AIOps, predictive monitoring, anomaly detection, or self-healing systems is desirable. Relevant certifications in Azure, Datadog, Kubernetes, or related cloud and reliability technologies are advantageous. Bachelor’s degree or equivalent technical education and professional experience. Benefits Fully remote opportunity based in India, with a rotational shift structure supporting 24x7 operations. Leadership responsibility across cloud reliability, infrastructure operations, incident management, and operational excellence. Hands-on exposure to Azure IaaS, AKS, Kubernetes, Docker, Datadog, Azure Monitor, and Infrastructure as Code. Opportunity to develop advanced observability, AIOps, predictive monitoring, automation, and self-healing capabilities. Collaboration with engineering teams on cloud-native architecture, deployment pipelines, scalability, and reliability initiatives. Opportunities to mentor team members and influence operational standards and engineering practices. Exposure to multi-cloud technologies and large-scale distributed systems. An inclusive work environment focused on teamwork, openness, respect, fairness, and continuous improvement. Support for professional development through exposure to modern cloud and reliability technologies and relevant certification paths. Opportunities to participate in leadership reporting, business reviews, and cross-functional initiatives with broad organizational visibility.
- How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
- We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
- apply for this job