Infrastructure Engineer @Betby
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted YDay

[Hiring] Infrastructure Engineer @Betby

YDay - Betby is hiring a remote Infrastructure Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

At BETBY, our team is responsible for building, managing and maintaining secure, robust and scalable infrastructure to run our prod and beta/test/dev environments and core services using DevOps principles. This includes:

  • Designing, deploying, configuring, and maintaining scalable monitoring, logging, and alerting platforms for production and beta/test/dev environments;
  • Operating Prometheus, Alertmanager, Grafana, Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards, including upgrades, reliability, availability, capacity, and retention planning;
  • Building and maintaining reliable metrics and log collection pipelines for infrastructure and business-critical services;
  • Creating dashboards that provide clear, useful visibility into service health, performance, capacity, and operational risks;
  • Designing, tuning, and maintaining actionable alert rules and notification routing; reducing alert noise and improving incident response;
  • Monitoring infrastructure and application metrics and logs, troubleshooting issues, and improving stability and performance under heavy loads;
  • Managing metric cardinality, log volume, retention, storage consumption, and query performance to keep observability platforms scalable and cost-effective;
  • Establishing high-availability and recovery approaches for observability services and validating operational readiness;
  • Automating configuration management and standardizing observability configuration through Ansible, Terraform, Python, and bash;
  • Developing self-service observability patterns, reusable dashboards, alert templates, and documentation for engineering teams;
  • Supporting production incidents, investigating root causes with telemetry, and improving dashboards, alerts, and runbooks after incidents;
  • Evaluating new technologies and their implementation in existing infrastructure;
  • Working with Kubernetes, Linux systems, networking, databases, and message brokers to ensure meaningful observability coverage;
  • Maintaining and writing documentation of observability architecture, configurations, standards, and operational procedures.

Qualifications

  • Minimum 3 years of experience with administering Linux systems and operating monitoring, logging, or observability systems;
  • Experience with Debian-based systems;
  • Experience with Docker and Kubernetes;
  • Hands-on experience with Prometheus, Alertmanager, and Grafana, including metric collection, alert rules, routing, and dashboards; experience with VictoriaMetrics would be a plus;
  • Experience with Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards for log collection, transport, processing, storage, search, and visualization;
  • Understanding of metrics and log pipeline design, including reliability, scalability, data retention, capacity planning, and cardinality management;
  • Experience designing actionable alerts, reducing alert noise, and troubleshooting infrastructure and application issues using metrics and logs;
  • Proficiency in shell command line usage, scripting, and automation tools like Ansible/Terraform;
  • Python and bash scripting skills;
  • Understanding of networking concepts, including TCP/IP, DNS, VPN, Firewalls and the ability to configure and troubleshoot network settings;
  • Experience operating highly available services and planning capacity for production workloads;
  • Experience with configuration-as-code, Git-based workflows, and enabling self-service observability for engineering teams;

Requirements

  • OS: Debian;
  • Virtualization: KVM, Proxmox;
  • Storages: Ceph, S3;
  • Networking: IPsec, Open vSwitch, iptables, VRRP, OpenVPN;
  • DBMS: MongoDB, PostgreSQL, ClickHouse;
  • Message brokers: Kafka, RabbitMQ;
  • Service orchestration: Kubernetes;
  • Monitoring systems: VictoriaMetrics, Prometheus, Alertmanager, Grafana;
  • Logging pipeline: Fluent Bit, Kafka, Fluentd, OpenSearch, OpenSearch Dashboards;
  • Revision control and CI/CD tools: GitLab;
  • Cloud services: Amazon Cloudfront/WAF/S3/EC2/EKS/ELB;
  • Web services: Nginx, HAproxy;
  • Configuration management: Ansible, Terraform;
  • Scripting: Python, bash.

Benefits

  • Comprehensive health insurance with coverage for your well-being;
  • Paid sick leave up to 10 days without medical certificate;
  • 20 days of paid vacation plus additional leave for important life events;
  • Learning and growth opportunities with support for professional development;
  • Language learning support for multilingual collaboration;
  • Modern hardware provided for your work;
  • International team environment across multiple countries;
  • Corporate events and team activities;
  • Welfare support program for critical situations;
  • Gifts and support for major life milestones.
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Infrastructure Engineer @Betby
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted YDay
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs