Role Description
You are the paved road for every AI product we ship β the CI/CD, IDP, and MLOps substrate that lets model teams deploy without opening a ticket. This is an entry-to-mid level role.
Bitdeer is building an AI-operated GPU cloud, and the AI Cloud team ships the products on top of it. As Cloud Senior DevOps Engineer you build the paved road those products travel down:
-
CI/CD, infrastructure-as-code, MLOps pipelines, and the Internal Developer Platform that lets model, data-science, and product teams deploy at speed without stepping around governance.
-
You are the bridge between research and production, and the operator of the platform substrate the AIOps team plugs into.
Qualifications
-
Bachelor's degree or above in Computer Science, Engineering, or a related technical field.
-
5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles.
-
Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs).
-
Deep mastery of Docker and Kubernetes orchestration.
-
Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud).
-
Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.).
-
Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), Observability paradigms, and Site Reliability Engineering (SRE) principles.
-
Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills.
Requirements
-
Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models.
-
Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production.
-
Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker.
-
Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing.
-
Take ownership of high-availability design in production environments.
-
Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs.
-
Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm.
-
Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack).
-
Build the paved road that lets product, model, and data-science teams deploy without opening a ticket.
-
Work closely with R&D, Data Science, Security, and Business teams to streamline workflows.
-
Establish and enforce system stability and security standards.
-
Act as the technical lead during complex system anomalies and major incidents.
Preferred Qualifications (Plus)
-
Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server).
-
Proven track record working with large-scale distributed systems or high-concurrency environments.
-
Hands-on experience in designing and building Internal Developer Platforms (IDP).
-
Deep familiarity with Zero Trust architecture and automated security testing (DevSecOps).
-
Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams.
-
You see incidents, tickets, and manual runbooks as source material for the next automation.
-
You think in golden paths and defaults, not policies.
Benefits
-
Equal employment opportunities in accordance with country, state, and local laws.
-
No discrimination against employees or applicants based on various conditions.