Platform Engineer, AI/ML Infrastructure @OpenTeams
Artificial Intelligence
Salary usd 120,000 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3d ago

[Hiring] Platform Engineer, AI/ML Infrastructure @OpenTeams

3d ago - OpenTeams is hiring a remote Platform Engineer, AI/ML Infrastructure. πŸ’Έ Salary: usd 120,000 - 250,000 per year πŸ“Location: USA

Role Description

OpenTeams builds AI platforms that governments and enterprises own outright: the infrastructure, the data, the models, and the evidence that the whole thing does what it claims. We're hiring several engineers to build and run that infrastructure.

The work spans the full depth of an AI platform:

  • Kubernetes clusters that schedule GPU workloads, move large datasets, and keep tenants isolated from one another.
  • Services that make it a platform rather than a cluster: workflow orchestration, data ingest, model serving, policy enforcement, audit logging.
  • The delivery path that gets released into production reliably and can prove what it shipped.
  • The cloud infrastructure that ties it all together, and the operational practices that keep a distributed system resilient.

Much of this has to run where you can't assume normal cloud resources, or even an internet connection. Portability, reproducibility, and operability are design inputs from the first commit rather than problems handed downstream.

We build on open source and contribute back. Kubernetes, Terraform and OpenTofu, Argo, Prometheus, Nebari, and others. Upstream work is part of the job, not something you do on weekends.

This posting covers multiple roles, spanning mid-level through senior. We understand nobody spans every area above, so tell us where you fit. We assign level-based roles based on what you've actually done rather than a year count.

Key Responsibilities

  • Build and operate Kubernetes-based infrastructure for demanding AI/ML workloads, including GPU scheduling, resource management, and multi-tenant isolation.
  • Design and implement platform services for orchestration, data ingest, model serving, and results management behind documented APIs.
  • Write infrastructure as code and build GitOps pipelines so environments are reproducible from source.
  • Build and operate CI/CD pipelines that produce versioned, signed, scanned release artifacts along with the documentation needed to deploy them.
  • Own reliability: capacity planning, upgrade paths, failure-mode analysis, backup and recovery, incident response, and postmortems.
  • Implement monitoring, logging, tracing, and alerting, and define the service level objectives they're measured against.
  • Deploy and validate the platform in restricted, disconnected, or limited-connectivity environments, and verify parity after each release.
  • Keep the platform portable by constraining dependencies to what's confirmed available in target environments.
  • Write runbooks and operational documentation that other engineers can execute without you in the room.
  • Contribute to Nebari and other open-source infrastructure, Kubernetes, and MLOps projects.
  • Work with security engineers, government stakeholders, and other engineers to turn requirements into systems that hold up.
  • Collaborate asynchronously across a distributed team.

Qualifications

  • U.S. citizenship, and the ability to obtain and maintain a U.S. security clearance.
  • Four or more years of hands-on experience building or operating production infrastructure, platforms, or distributed systems.
  • Production experience with Kubernetes and containerized workloads.
  • Experience with at least one major cloud platform: AWS, Azure, or Google Cloud.
  • Experience with infrastructure as code and CI/CD, using tools such as Terraform, OpenTofu, Pulumi, or Helm.
  • Working proficiency in Python, Go, Bash, or a comparable language.
  • Experience implementing or operating production monitoring and observability.
  • Ability to write documentation, runbooks, and deployment procedures that other people can actually follow.
  • Ability to work independently and collaborate well in a remote, distributed team.

Nice to Have

  • An active U.S. security clearance, particularly TS/SCI with CI polygraph.
  • Experience deploying or operating software in air-gapped, disconnected, or otherwise restricted environments.
  • Experience with Department of Defense, Intelligence Community, or comparably regulated programs.
  • Familiarity with the Risk Management Framework, NIST 800-53 or 800-171, or similar frameworks, and with producing the evidence they require.
  • Experience supporting an Authorization to Operate, or with continuous ATO models.
  • Supply chain security work: hardened images, artifact signing, SBOM generation, dependency and container scanning, policy enforcement.
  • Familiarity with cross-domain solutions, guards, data diodes, or similar transfer mechanisms.
  • Experience with classified cloud environments, including AWS Secret or Top Secret regions.
  • A DoD 8140/8570 qualifying certification such as Security+, CISSP, CASP+, or CISM, or willingness to obtain one after joining.
  • Experience building MLOps pipelines or infrastructure for AI/ML workloads.
  • Experience with GPU scheduling, distributed inference, or large-scale data and evaluation pipelines.
  • Experience with model-serving or gateway frameworks such as KServe, vLLM, or LLM-D.
  • Experience designing API-first services and vendor-agnostic platforms that run across multiple environments.
  • Experience with agentic workflow frameworks or multi-step AI pipeline orchestration.
  • Contributions to open-source Kubernetes, infrastructure, MLOps, or observability projects, and experience with Nebari specifically.
  • Familiarity with data sovereignty and privacy requirements for enterprise or government AI systems.
  • Experience leading technical initiatives, setting engineering standards, or mentoring other engineers.
  • Experience supporting rapid prototyping programs or defense innovation initiatives.

Benefits

  • 100% employer paid medical premiums for employees.
  • Self-managed PTO with a minimum time off requirement.
  • Opportunities for growth and collaboration with global experts.
  • Commitment to diversity, equity, inclusion, and belonging.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Platform Engineer, AI/ML Infrastructure @OpenTeams
Artificial Intelligence
Salary usd 120,000 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3d ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 120,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 123,989+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later