Site Reliability Engineer @Arango
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2wks ago

[Hiring] Site Reliability Engineer @Arango

2wks ago - Arango is hiring a remote Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: India

Role Description

At ArangoDB, we are building a robust, cloud-native infrastructure to support our distributed database systems, which power mission-critical applications for a wide range of industries. We are searching for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our infrastructure and applications, with a focus on automation, monitoring, and optimizing cloud environments.

As a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams.

Your goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you're passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you!

Key Responsibilities

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
  • Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
  • Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.
  • Develop strategies for disaster recovery, high availability, and fault tolerance.
  • Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).
  • Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
  • Participate in on-call rotations to support critical production systems and respond to incidents.
  • Collaborate with cross-functional teams to improve overall system reliability and scalability.
  • Collaborate with the Customer Success team to resolve customer issues.

Qualifications

  • 4-7 years of experience: SRE or DevOps Engineer background in cloud-native environments.
  • Self-organized, autonomous remote team player with strong communication skills.
  • AWS and GCP; advanced Linux internals (processes, environment variables); containerization and orchestration (Docker, Kubernetes at scale).
  • CI/CD pipelines (Jenkins, CircleCI); monitoring, alerting, and logging (Prometheus, Grafana, ELK stack); Git version control.
  • Core networking, security best practices, and systematic troubleshooting of complex infrastructure issues.
  • Programming proficiency in Golang or Python.

Nice-to-Have

  • Experience managing distributed databases or large-scale data storage systems.
  • Knowledge of security best practices in cloud environments.
  • Experience with scripting languages like Python or Bash.
  • Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus.
  • Experience working with GitOps.
  • Strong programming skills in Golang, with experience in developing automation tools, scripts, or services.

What Makes Arango Special?

At Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI.

Working at Arango means:

  • Contributing to cutting-edge AI and data infrastructure.
  • Collaborating with experienced engineers, marketers, and product leaders.
  • Helping shape how enterprises build AI-powered applications.

If you're excited about the intersection of AI, data, and social media, we’d love to hear from you.

Before You Apply
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineer @Arango
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 125,728+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later