Role Description
As a Platform Engineer at ThinkOn, you'll build and maintain the monitoring, observability, and infrastructure automation that keeps our cloud platform reliable. Your day-to-day will span:
-
Zabbix and Prometheus/Grafana for monitoring and dashboards
-
Opsgenie for alert routing and on-call management
-
Infrastructure-as-code tooling (Ansible, Terraform, GitLab CI/CD) to deploy and manage it all
You'll work within a VMware Cloud Foundation (VCF) and Kubernetes environment, collaborating with infrastructure, network, security, and DevOps teams.
Qualifications
-
Diploma or degree in Computer Science, IT, or a related field (or equivalent practical experience)
-
Eligible to obtain Secret Level Clearance within your first 3 months of employment
-
Relevant certifications are a plus but not required (e.g., Zabbix Certified Professional, CKA, CompTIA Linux+, ITIL v4 Foundations)
-
Hands-on experience with Zabbix (or comparable: Nagios, Icinga, Checkmk)
-
Working knowledge of Prometheus and Grafana—writing exporters, building dashboards, PromQL
-
Experience with alert management and on-call tooling (Opsgenie, PagerDuty, or similar)
-
Comfort with Linux systems administration
-
Proficiency in scripting and automation—Bash and Python at minimum
-
Experience with at least one IaC tool (Ansible, Terraform)
-
Familiarity with CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, or similar)
-
Understanding networking fundamentals—TCP/IP, DNS, SNMP, bandwidth/latency concepts
-
Basic database administration (PostgreSQL or MySQL) for monitoring tool backends
-
Strong diagnostic and troubleshooting skills
-
Clear written and verbal communication
-
Attention to detail
-
Comfort working independently in a remote environment
-
Ability to stay composed during incidents and work under time pressure
Requirements
-
Act as a first responder for monitoring-related incidents and alerts
-
Investigate and resolve performance issues, outages, or anomalies detected by monitoring systems
-
Escalate to the appropriate teams (network, security, infrastructure) when needed
-
Document incidents, root causes, and resolutions
-
Contribute to post-incident reviews
-
Provide technical support to internal teams on monitoring tools and dashboards
-
Work within VMware vCenter / VCF and Kubernetes environments to support monitoring and infrastructure needs
-
Manage notification infrastructure (SMTP relay configuration, delivery troubleshooting)
-
Support compliance requirements (ISO 27001, SOC 2)
Benefits
-
A remote-first culture
-
Competitive compensation package
-
Flexible time-off for vacation, plus illness & personal days
-
Comprehensive health and dental benefits, including GRSP & 401k Matching Program