Role Description
As a Senior DevOps Platform Engineer, you will own the deployment backbone that makes Salt’s multi-environment AI platform reliable, fast, and secure. You’ll design and operate infrastructure that runs across SaaS, VPC, and on-prem—supporting customers who depend on Salt for mission-critical life sciences workloads.
-
Drive velocity by building and maintaining high-speed, reliable CI/CD pipelines.
-
Own multi-cloud deployment consistency across SaaS, VPC, and on-prem environments (GCP, AWS, Azure), ensuring parity and predictable behavior.
-
Define and enforce infrastructure and deployment standards to prevent environment drift.
-
Implement and evolve monitoring, logging, and alerting to ensure platform health, performance, and SLOs.
-
Automate provisioning and configuration of infrastructure using infrastructure-as-code.
-
Design and maintain application security, backup, and disaster recovery procedures.
-
Partner closely with developers to improve application performance, reliability, and operability.
-
Participate in incident response, including on-call rotations, root cause analysis, and remediation.
-
Maintain clear, discoverable documentation for infrastructure, deployment processes, and runbooks.
-
Support QA by enabling automated testing in pipelines and stable test environments.
-
Champion and spread best practices for code quality, security, and deployment standards.
Qualifications
-
5+ years of experience operating production cloud infrastructure (GCP and/or AWS), ideally in multi-cloud environments.
-
Deep experience with CI/CD systems and pipeline design (e.g., GitLab CI, GitHub Actions).
-
Strong experience with Kubernetes and container orchestration in production.
-
Hands-on experience with infrastructure-as-code tools (Terraform, CloudFormation, or similar).
-
Proficiency with Docker and core containerization concepts.
-
Experience implementing and operating observability stacks (metrics, logs, traces).
-
Solid understanding of web application deployment, scaling, and reliability fundamentals.
-
Scripting skills in Python, Bash, or similar for automation.
-
Experience integrating automated tests into deployment workflows.
-
Strong problem-solving and debugging skills, especially in distributed systems.
-
Clear, concise communication and a collaborative working style.
Requirements
-
Experience with database administration, backup, and recovery strategies.
-
Familiarity with security best practices for cloud-native and multi-tenant applications.
-
Experience with performance testing, capacity planning, and cost optimization.
-
Strong understanding of networking, load balancing, and traffic routing.
-
Experience with AI/ML, HPC, or other data- and compute-intensive workloads.
-
Contributions to internal DevOps/Platform tooling or open-source automation projects.
Benefits
-
Competitive salary based on location and experience.
-
Comprehensive health coverage, including medical, dental, mental health, and 401(k).
-
Fully remote role, collaborating primarily within U.S. time zones.