Site Reliability Engineer 3 @Granicus India
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago

[Hiring] Site Reliability Engineer 3 @Granicus India

2mths ago - Granicus India is hiring a remote Site Reliability Engineer 3. πŸ’Έ Salary: unspecified πŸ“Location: India

Role Description

Granicus is seeking a Site Reliability Engineer 3 (SRE) with strong AIOps, automation, and AI proficiency to modernize reliability engineering through observability, intelligent incident response, and responsible AI-assisted operations. In this role, you will improve service reliability, reduce operational toil, accelerate incident response, and help build scalable, resilient platforms supporting traditional, cloud-native, and AI/ML-powered workloads. The role will also help operationalize AI-enabled SRE practices such as:

  • alert intelligence
  • assisted root-cause analysis
  • runbook automation
  • telemetry summarization
  • governed self-healing workflows with appropriate human approval and audit controls

You will be expected to implement practical AIOps capabilities across observability, incident response, automation, and operational knowledge workflows, turning AI/ML insights into production-ready reliability improvements.

Qualifications

  • 6+ years of experience in SRE, AIOps, or production engineering in large-scale, cloud environments.
  • Strong expertise in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP).
  • Expert in ELK/OpenSearch, including:
    • Log ingestion (Logstash / Beats)
    • Elasticsearch index design, scaling, and tuning
    • Advanced Kibana querying and debugging
    • Dashboards, alerts, and observability patterns for production systems
  • Hands-on experience in logs, metrics, and tracing.
  • Solid understanding of incident management, RCA, SLOs, and operational best practices.
  • Good understanding of AIOps: anomaly detection, alert correlation, and intelligent alerting.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar.

Requirements

  • Demonstrated production use of LLMs/AI agents for SRE workflows β€” e.g., automated log triage, RCA drafting, runbook generation, or incident summarization.
  • Experience building or integrating AIOps capabilities: anomaly detection, alert correlation/clustering, predictive capacity signals.
  • Working knowledge of prompt engineering for operational use cases.
  • Comfort evaluating AI output critically.
  • Familiarity with agentic frameworks or MCP-style tool integration is a strong plus.
  • Understanding of the risk surface of AI-in-production.
  • Provide on-call production support, ensuring rapid triage, escalation handling, and service restoration.
  • Investigate production and customer issues, lead incident troubleshooting, and drive rapid RCA with clear follow-ups.
  • Own and evolve the observability stack, with deep expertise in ELK/OpenSearch.
  • Design and maintain observability across logs, metrics, and traces.
  • Build and enhance workflows for alerting, anomaly detection, and incident enrichment.
  • Develop automation, runbooks, and controlled self-healing mechanisms.
  • Drive improvements in system reliability, performance, scalability, and resilience.
  • Partner with engineering teams to improve deployment safety, operational readiness, and production stability.
  • Maintain high-quality runbooks, documentation, and knowledge bases.
  • Support capacity planning, performance tuning, and SLO-based reliability practices.
  • Apply security, access control, and operational guardrails across systems and automation.
  • Own AIOps implementation from use-case definition through production rollout.

Benefits

  • Remote work flexibility.
  • Inclusive and diverse work environment.
  • Opportunities for professional growth and development.

Company Description

Granicus is driven by the excitement of building, implementing, and maintaining technology that is transforming the Govtech industry by bringing governments and its constituents together. We are on a mission to support our customers with meeting the needs of their communities and implementing our technology in ways that are equitable and inclusive.

  • Consistently appeared on the GovTech 100 list over the past 5 years.
  • Recognized as one of the best companies to work for on BuiltIn.
  • Served 5,500 federal, state, and local government agencies.
  • More than 300 million citizen subscribers using our digital solutions.
Before You Apply
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineer 3 @Granicus India
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,048+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later