Senior Infrastructure Operations Engineer @KLDiscovery
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago

[Hiring] Senior Infrastructure Operations Engineer @KLDiscovery

1mth ago - KLDiscovery is hiring a remote Senior Infrastructure Operations Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

The Senior Infrastructure Operations Engineer owns the design, administration, and continuous improvement of KLDiscovery's compute, storage, and cloud infrastructure globally. This role takes end-to-end technical ownership of the physical server, virtualization, block/file/object storage, enterprise backup, and Azure IaaS environments underpinning KLDiscovery's production client systems, and serves as the primary technical escalation point and design authority within the Compute & Storage team. The Senior Infrastructure Operations Engineer leads infrastructure design decisions, drives standards development, partners with Enterprise Architecture and IT Security on architecture and compliance, and provides mentoring and technical direction to Infrastructure Operations Engineers. This role participates in 24x7x365 on-call rotation.

Key Responsibilities

  • Compute, Virtualization & Storage Architecture:
    • Hold end-to-end technical ownership of KLDiscovery’s physical server and virtualized compute environments.
    • Define and enforce configuration standards, capacity thresholds, and operational procedures.
    • Lead the design of significant compute changes, cluster expansions, and platform upgrades.
    • Review and approve significant configuration changes before implementation.
  • Backup, Recovery & Azure IaaS:
    • Own the enterprise backup platform and recovery strategy β€” job standards, recovery validation procedures, and RPO/RTO alignment.
    • Ensure recovery procedures are documented, tested on a defined schedule, and executable by any team member.
    • Manage the formal backup coverage request process for consuming teams.
    • Own the administration and governance of Azure IaaS infrastructure.
    • Define Azure infrastructure standards in conjunction with Enterprise Architecture.
    • Monitor and govern cloud costs; surface anomalies and optimization opportunities proactively.
    • Partner with Enterprise Architecture on hybrid cloud architecture direction and roadmap input.
  • OS Standards, Patching & Provisioning:
    • Own Windows Server and Linux configuration standards, patching cadence, and hardening baselines across all infrastructure globally.
    • Own the server patching function for all server infrastructure.
    • Define and own the server build runbook, provisioning standards, sizing guidelines, and handoff checklists in conjunction with Enterprise Architecture.
  • Security, Observability & Escalation:
    • Embed security controls into infrastructure design from inception.
    • Support audit and compliance reviews as needed.
    • Maintain working familiarity with applicable security and compliance frameworks (ISO 27001, CIS Controls).
    • Partner with the Automation & Observability team to ensure all owned infrastructure is covered by monitoring and alerting.
    • Serve as the primary escalation point within the Compute & Storage team for complex or time-critical infrastructure incidents.
    • Conduct root cause analysis for significant incidents, lead post-incident reviews and blameless retrospectives, and drive systemic remediation.
  • Continual Service Improvement & Automation:
    • Evaluate existing infrastructure for improvement, consolidation, and modernization opportunities.
    • Partner with the Automation & Observability team to identify, prioritize, and define requirements for infrastructure automation.
    • Own and govern infrastructure performance KPIs β€” availability, capacity utilization, incident response SLAs, and patching compliance.
  • Strategy, Cross-Team Coordination, Projects & Documentation:
    • Serve as the primary technical liaison to Database, Enterprise Platforms, Automation & Observability, Networking, and DC Operations.
    • Translate architectural designs from Enterprise Architecture into operational infrastructure standards and procedures.
    • Lead technical delivery of infrastructure projects from inception to completion with measurable milestones and managed scope.
    • Own the documentation standard for the Compute & Storage team.
  • Decision Scope & Accountability:
    • Authorized to make independent infrastructure configuration and architectural decisions within defined scope and standards.
    • Exercises discretion on decisions with broader impact.
  • Budgetary Awareness:
    • Weighs cost into all infrastructure recommendations.
    • Surfaces cost considerations and optimization opportunities proactively.
  • Mentoring & Development:
    • Provides significant mentoring and coaching to Infrastructure Operations Engineers.
    • Actively contributes to cross-training with Database, Automation & Observability, Enterprise Platforms, and DC Operations teams.

Qualifications

  • 6+ years in infrastructure engineering with hands-on ownership across compute, storage, and cloud in a production enterprise environment; prior senior or lead individual contributor experience preferred.
  • Expert-level administration of VMware vSphere and Nutanix in enterprise production environments.
  • Deep experience with block and file storage technologies β€” RAID, SAN (Fibre Channel or iSCSI), NAS protocols, and enterprise storage array administration.
  • Expert-level administration of Veeam or equivalent enterprise backup platform.
  • Strong working knowledge of Azure IaaS architecture.
  • Expert-level Windows Server administration.
  • Strong Linux (Ubuntu) administration.
  • Advanced PowerShell and/or Bash scripting.
  • Working knowledge of IaC concepts and tooling (Ansible, Terraform, or equivalent).
  • Working familiarity with security and compliance frameworks (ISO 27001, CIS Controls).
  • Familiarity with container technologies (Docker, Kubernetes).
  • Proficient in ITSM processes and ITIL-based Incident, Problem, Change, and Capacity Management.
  • Experience delivering and owning technical projects end-to-end.
  • Effective communication with peers, management, vendors, and internal customers at all levels.
  • Experience supporting 24x7 global production environments; on-call availability required.

Education

  • Bachelor’s degree in computer science, Information Technology, or equivalent experience.

Preferred

  • ITIL V3/4; VMware VCP, Nutanix NCP, or Azure Administrator/Solutions Architect (AZ-104/AZ-305); Microsoft Certified: Azure Administrator Associate; Veeam certification; CompTIA Security+ or CISSP; prior experience in eDiscovery, legal technology, or a similarly regulated environment.

Benefits

  • Paid time off, that offers various time off options to help employees maintain a work-life balance.
  • Ongoing learning and development through various training and education reimbursement programs.
  • A diverse and inclusive workplace where we all learn, grow, and achieve the greatest heights together.
  • A surrounding team of mission-driven individuals who genuinely love what they do.
  • Free, fun, interactive and incentivized global wellness program that promotes the wellbeing of our employees.
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Infrastructure Operations Engineer @KLDiscovery
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,860+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later