Platform & SRE Engineer @MWDN
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted Today

[Hiring] Platform & SRE Engineer @MWDN

Today - MWDN is hiring a remote Platform & SRE Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Ukraine

Role Description

A fast-growing product company building next-generation infrastructure for production AI systems. The team is focused on high-performance, low-latency distributed technologies that operate directly within real-time data flows and AI workloads. The product addresses complex engineering challenges related to runtime reliability, distributed processing, networking, observability, and large-scale system performance.

The client is looking for a Platform & SRE Engineer to join the engineering team and take ownership of the infrastructure and operational foundation behind our distributed runtime environments. This is a hands-on engineering role combining:

  • Linux systems engineering
  • AWS infrastructure
  • Machine-image management
  • Deployment automation
  • Observability
  • Security hardening
  • SRE practices

A key part of the role is owning the lifecycle of Linux-based runtime hosts β€” from AMI creation and machine provisioning through operating-system configuration, system services, runtime deployment, host tuning, upgrades, security hardening, monitoring, and operational troubleshooting.

You will maintain and evolve our internal host management and deployment tooling, ensuring runtime environments can be deployed, configured, upgraded, diagnosed, and operated reliably across cloud and, in the future, customer-hosted environments.

In parallel, you will help establish and own our SRE and observability practices, ensuring we can measure platform health and availability, identify failures quickly, and operate against clearly defined reliability objectives.

Qualifications

  • 8+ years of experience in DevOps, SRE, Platform Engineering, Linux Systems Engineering, or Infrastructure Engineering.
  • Strong Linux systems administration and troubleshooting skills, including systemd, processes, networking, filesystems, permissions, kernel configuration, and system-level debugging.
  • Strong hands-on AWS experience, particularly with EC2, networking, IAM, storage, and production Linux workloads.
  • Strong experience with AWS AMIs, including building, configuring, hardening, validating, versioning, and maintaining machine images for production environments.
  • Strong experience with Linux server hardening and production security practices, including least privilege, service isolation, user and permission management, secure system configuration, and attack-surface reduction.
  • Hands-on experience with Grafana, OpenTelemetry, Prometheus, Mimir, ClickHouse, Fluent Bit, or comparable observability technologies.
  • Experience building production monitoring, dashboards, metrics, logging, and alerting solutions.
  • Understanding of SRE principles, including SLIs, SLOs, availability, incident response, and root-cause analysis.
  • Experience writing automation and operational tooling using Python and Bash.
  • Strong understanding of networking fundamentals including TCP/IP, routing, DNS, TLS, and secure connectivity.
  • Experience automating machine provisioning, configuration, deployment, and upgrades.
  • Ability to independently troubleshoot complex issues spanning infrastructure, operating systems, networking, and application runtime.
  • Strong ownership mindset and ability to take responsibility for systems from deployment through production operation.

Requirements

  • Experience with Linux performance tuning including CPU pinning, NUMA, hugepages, IRQ affinity, and NIC tuning.
  • Experience with high-performance or low-latency networking environments.
  • Experience with DPDK or other user-space networking technologies.
  • Experience operationally integrating FPGA or other hardware accelerators into Linux environments.
  • Experience supporting software deployed in customer-managed or on-prem environments.
  • Experience with PKI, certificate management, TLS/mTLS, and machine identity.
  • Experience with Infrastructure as Code such as Terraform.
  • Experience with GitHub Actions or similar CI/CD systems.
  • Experience with vulnerability management and security/compliance initiatives such as SOC 2.
  • Experience working in startup or high-growth engineering environments.

Benefits

  • People-first management with minimal bureaucracy.
  • A friendly company culture, proven by employees who choose to return.
  • Flexible working hours.
  • 29 days of PTO (18 working days per year plus all national holidays).
  • 10 paid recovery days.
  • Full financial and legal support for independent contractors.
  • Free English classes, with native speakers or Ukrainian teachers.
  • Dedicated HR support.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Ukraine
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Platform & SRE Engineer @MWDN
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted Today
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Ukraine
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,868+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later