Senior Automation & Observability Engineer @Ensono
Devops
Salary usd 113,000 - 1..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2wks ago

[Hiring] Senior Automation & Observability Engineer @Ensono

2wks ago - Ensono is hiring a remote Senior Automation & Observability Engineer. πŸ’Έ Salary: usd 113,000 - 147,000 per year πŸ“Location: USA

Role Description

We are seeking an experienced IoT / Observability Engineer responsible for monitoring, managing, automating, and optimizing enterprise infrastructure, applications, IoT platforms, and enterprise telemetry ecosystems. The ideal candidate will possess strong expertise in observability platforms, monitoring technologies, automation frameworks, and operational support processes to ensure high availability, reliability, and performance of business-critical systems.

The engineer will support enterprise monitoring operations, FOAK services, ELT platforms, incident management, and automation initiatives while collaborating with Infrastructure, Cloud, Network, Application, and Service Delivery teams.

Key Responsibilities

  • Monitoring & Observability
    • Design, implement, and maintain enterprise monitoring and observability solutions.
    • Develop and maintain dashboards, alerts, and visualizations using Grafana.
    • Monitor infrastructure, applications, middleware, and IoT services using IBM Instana, Grafana, SolarWinds, and related observability tools.
    • Configure and manage data collection using Telegraf, Prometheus, and monitoring agents.
    • Analyze metrics, logs, traces, events, and telemetry data to identify performance bottlenecks and service degradation.
    • Support SLO, SLA, and operational health monitoring initiatives.
    • Perform Root Cause Analysis (RCA) and troubleshooting for infrastructure and application issues.
  • Foak & Enterprise Logging/Telemetry
    • Support onboarding, monitoring, and operational management of FOAK (First Office Application Kit) services and enterprise applications.
    • Configure, validate, and troubleshoot Enterprise Logging & Telemetry (ELT) integrations across infrastructure, middleware, applications, and cloud platforms.
    • Monitor telemetry pipelines, log ingestion, event correlation, and data quality to ensure complete observability coverage.
    • Collaborate with engineering teams to improve telemetry standards, monitoring effectiveness, and proactive incident detection through ELT and observability frameworks.
    • Support FOAK application integrations with Grafana, Instana, Prometheus, and enterprise monitoring platforms.
  • Infrastructure & Platform Monitoring
    • Monitor and support Linux Servers.
    • Monitor and support Windows Servers.
    • Monitor and support VMware Infrastructure.
    • Monitor and support Citrix VDI Platforms.
    • Monitor and support DNS Services.
    • Monitor and support Proxy Services.
    • Monitor and support Middleware Platforms.
    • Monitor and support Integration Services.
    • Monitor and support Enterprise Applications.
    • Monitor and support IoT Platforms.
  • Additional Responsibilities
    • Investigate performance issues, recurring alerts, and infrastructure anomalies.
    • Validate monitoring platform health and monitoring coverage.
    • Monitor capacity, availability, CPU, memory, storage, and service health metrics.
    • Support platform upgrades, maintenance, and operational readiness reviews.
  • Database & Data Management
    • Configure and maintain InfluxDB time-series databases.
    • Manage data retention policies, performance tuning, and capacity planning.
    • Develop operational dashboards and reports for infrastructure and application performance insights.
    • Support telemetry data ingestion, storage optimization, and historical trend analysis.
  • Event & Incident Management
    • Monitor operational alerts, events, notifications, and incidents from enterprise monitoring platforms.
    • Acknowledge, investigate, troubleshoot, and resolve assigned incidents.
    • Coordinate with Infrastructure, Network, Cloud, Security, Application, and Service Delivery teams during incident resolution.
    • Participate in major incident bridges, DR exercises, and 24x7 operations support activities.
    • Follow escalation procedures, SOPs, operational runbooks, and ITIL processes.
    • Support Problem Management activities and contribute to RCA documentation.
  • Instana & APM Operations
    • Administer and support IBM Instana monitoring environments.
    • Monitor application, API, middleware, and microservices performance using Instana.
    • Validate Instana agent health following server patching and maintenance activities.
    • Support application onboarding and APM configuration standards.
    • Configure alerts, baselines, and performance thresholds.
    • Raise and track incidents related to Instana platform availability and performance.
  • Automation & Scripting
    • Develop automation solutions using Python, PowerShell, Shell Scripting (Bash), and VBScript.
    • Automate operational tasks, monitoring deployments, and remediation workflows.
    • Build reusable automation tools to improve operational efficiency.
    • Integrate monitoring platforms with enterprise automation frameworks.
    • Support webhook-based automation and event-driven operational workflows.
  • Configuration Management & Infrastructure Automation
    • Implement Infrastructure as Code (IaC) and automation using Ansible and Puppet.
    • Automate server provisioning and configuration management.
    • Automate monitoring agent deployment and onboarding.
    • Maintain automation playbooks and deployment pipelines.
    • Improve operational consistency and reduce manual efforts across environments.
  • Knowledge Management
    • Maintain SOPs, runbooks, monitoring procedures, and escalation matrices.
    • Participate in KT sessions, service onboarding, and operational readiness reviews.
    • Support service transition, migration, and continuous improvement initiatives.
    • Maintain observability standards and monitoring documentation.

Qualifications

  • 7 to 10+ years of experience in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering.
  • Experience supporting large-scale enterprise environments and 24x7 operations.
  • Hands-on experience with Grafana, Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB.
  • Experience supporting VMware, Citrix, middleware, enterprise applications, and cloud monitoring platforms.
  • Experience working with FOAK applications and Enterprise Logging & Telemetry (ELT) platforms.
  • Strong troubleshooting, RCA, incident management, and operational support skills.
  • Experience with automation frameworks and Infrastructure as Code (Ansible preferred).
  • Experience integrating observability platforms with enterprise automation solutions.

Requirements

  • Monitoring & Observability: Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, InfluxDB, OpenTelemetry, Grafana Alloy, APM Monitoring, Event Management, Alert Management, Observability Concepts, SLO/SLA Monitoring.
  • FOAK & Enterprise Telemetry: FOAK (First Office Application Kit) Support, Enterprise Logging & Telemetry (ELT), Log Aggregation & Correlation, Telemetry Data Analysis, Event Correlation, Application Onboarding, Monitoring Standards & Observability Frameworks.
  • Infrastructure: VMware, Linux Administration, Windows Server, Citrix VDI, DNS Services, Proxy Services, Middleware Technologies, Infrastructure Performance Monitoring.
  • Scripting & Programming: Python, PowerShell, Shell Scripting (Linux), VBScript.
  • Automation Tools: Ansible, Puppet, Webhooks, Infrastructure as Code (IaC).
  • ITSM & Operations: ServiceNow, Incident Management, Problem Management, Change Management, ITIL Framework, Major Incident Management.

Benefits

  • Unlimited Paid Days Off.
  • Three health plan options.
  • 401k with company match.
  • Eligibility for dental, vision, short and long-term disability, life and AD&D coverage, and flexible spending accounts.
  • Family Forming Benefit including fertility coverage and adoption/surrogacy reimbursement.
  • Paid childbearing and paternal leave.
  • Education Reimbursement, Student Loan Assistance or 529 College Funding.
  • Sabbatical leave.
  • Wellness program.
  • Flexible work schedule.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Automation & Observability Engineer @Ensono
Devops
Salary usd 113,000 - 1..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,991+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later