Role Description
We are looking for a talented and motivated DevOps Engineer / Site Reliability Engineer (SRE) to join our team. The ideal candidate will have a strong background in software engineering, systems administration, cloud infrastructure, and automation, with a passion for continuous improvement, reliability, and operational excellence.
-
Design, implement, and maintain CI/CD pipelines to support efficient and reliable software delivery.
-
Automate infrastructure provisioning, configuration, and deployment processes.
-
Manage and optimize cloud infrastructure and services in Microsoft Azure.
-
Monitor and improve system performance, reliability, scalability, and availability.
-
Collaborate closely with development teams to ensure smooth deployment and operation of services.
-
Implement and maintain security best practices across infrastructure and deployment processes.
-
Troubleshoot and resolve infrastructure, deployment, and system issues in a timely manner.
-
Participate in on-call rotations and incident response activities.
-
Identify opportunities for automation and continuous improvement to increase operational efficiency and reduce manual effort.
Qualifications
-
Advanced/Fluent English β strong communication skills, both written and verbal.
-
Professional experience in DevOps, Cloud Engineering, Site Reliability Engineering, or a related field.
-
Hands-on experience with CI/CD tools, such as Jenkins, GitLab CI, or GitHub Actions.
-
Strong experience with Microsoft Azure cloud services and infrastructure.
-
Experience with Docker and Kubernetes.
-
Proficiency in scripting/programming languages such as Python, Bash, and/or PowerShell.
-
Experience with monitoring and logging tools such as Prometheus, Grafana, and ELK Stack.
-
Strong knowledge of Git and Git-based platforms, including GitHub and/or GitLab.
-
Hands-on experience with Infrastructure as Code (IaC), particularly Terraform.
-
Strong understanding of networking and security principles.
-
Experience with networking concepts such as VNet, DNS, load balancing, and firewalls.
-
Strong troubleshooting and problem-solving skills.
-
Excellent communication and collaboration skills.
Preferred Skills
-
Knowledge or experience applying Site Reliability Engineering (SRE) principles to improve system reliability, availability, and uptime.
-
Experience with configuration management tools such as Ansible or Chef.
-
Experience with databases such as MongoDB, MySQL, or PostgreSQL.
-
Experience with Azure automation and automation scripting using Python and Shell.
-
Experience integrating Azure IoT solutions and Event Hubs for real-time data processing.
-
Experience working with Agile development methodologies.
Location
MEX Work-at-Home
Time Type
Full time