Role Description
You will lead the DevOps, infrastructure engineering, and operational readiness for an innovative, high-impact power management platform designed to optimize energy consumption in large-scale data centers and alleviate pressure on power grids.
-
Own CI/CD, deployment, release, and operational readiness across Azure cloud and edge Kubernetes environments for a mission-critical energy management platform.
-
Consolidate a fragmented CI/CD estate (currently split across GitHub Actions and Azure DevOps) into a unified, canonical standard, driving adoption across engineering teams.
-
Build golden paths and reusable pipeline templates to ensure newly developed services are production-deployable on day one.
-
Establish repeatable, auditable, versioned, and reversible deployment pipelines for disconnected or intermittently connected edge sites.
-
Define and enforce operational readiness criteria (observability, alerting, runbooks, on-call rotation). Own the incident management lifecycle, blameless postmortems, and remediation tracking.
-
Embed automated security gates into CI/CD pipelines (container scanning, dependency scanning, secret detection, SBOM, signed artifacts, policy-as-code) and manage credential/secret lifecycles across cloud and edge locations.
-
Line-manage DevOps, MLOps, and QA automation engineers—handling 1-on-1s, performance goals, growth, and hiring. Set technical standards, drive nearshore engineering outcomes, and track DORA metrics for senior leadership.
Qualifications
-
Minimum 6+ years of hands-on experience in a DevOps leadership or engineering management role.
-
Proven track record running Kubernetes in production across both cloud and edge infrastructure.
-
Deep technical knowledge of Azure services, including AKS, Azure Identity/Active Directory, and cloud networking.
-
Extensive experience consolidating and scaling enterprise pipelines using GitHub Actions and Azure DevOps.
-
Hands-on experience establishing production telemetry using tools like Datadog and Grafana.
-
Strong experience in automated security delivery, secret rotation (especially in low-connectivity environments), and policy-as-code.
-
Direct experience mentoring, hiring, and managing technical talent (DevOps, QA, MLOps).
-
Proven success thriving in distributed remote environments using Agile methodologies.
-
Professional working proficiency for daily oral and written communication in English.
Requirements
-
Exposure to AI coding agents and automated code review workflows.
-
Familiarity with AWS infrastructure and MLOps practices (model deployment and monitoring).
-
Industry background in energy, industrial, utility, or data center domain engineering.
-
Knowledge of Microsoft AI Foundry or frameworks like LangGraph.
Time Zone & Collaboration
This role requires a minimum overlap of 6 hours with US Central (CT) and Pacific (PT) business hours during a standard 8-hour workday to collaborate effectively with distributed engineering teams.
Language
All interviews, technical documentation, daily team syncs, and operational communications will be conducted in English.