Role Description
We're building the financial data platform at Accelerant β the premium, claims, and paid data products that underpin financial processing, reserving analysis, and the monthly close β and it needs to stay fast, resilient, and observable as we scale. You'll drive the reliability and observability strategy across the platform and the enterprise systems it depends on: Velocity, MuleSoft, D365, Snowflake, Fabric, and the streaming and integration layers that move data through it. You are a key decider about what gets measured, how we define reliability, and where engineering needs to invest to keep production healthy.
We need someone who can prove a repeatable, define-to-alert observability pipeline, harden it, and scale it into systems that have never had real SLOs β and build modern, AI-assisted operational tooling that lets a small team punch far above its weight.
This Is a High-Autonomy, High-Impact Role for Someone Who:
-
Sees a recurring alert or a fragile deploy path and cannot leave it alone. Excels at shipping the right fix and the right automation, not the perfect one.
-
Has run real production systems at scale β not just written runbooks for them.
-
Has built with Datadog, OpenTelemetry, incident tooling, and AI coding assistants long enough to have strong opinions about what fits our needs.
-
Can prototype an operational agent in Cursor and iterate as they go.
-
Operates with autonomy, and can carry a technical discussion on system architecture, failure modes, and tradeoffs.
-
Is genuinely curious about applying emerging AI to reliability and operations.
Qualifications
-
Proven experience designing, operating, and scaling reliable production systems.
-
Deep hands-on expertise with modern observability tooling β Datadog, Prometheus/Grafana, and OpenTelemetry β including both push and pull ingestion patterns.
-
Strong background defining SLIs, SLOs, and error budgets β and translating them into business-level KPIs, not just infrastructure metrics.
-
Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) alongside enterprise SaaS such as D365 (F&O and/or Power Apps).
-
Hands-on incident management experience with tools like Incident.io and ServiceNow, and a track record of running effective on-call and postmortem practices.
-
Hands-on experience building with LLMs and AI coding assistants β Cursor in particular. Bonus if you've built and deployed agents.
-
Ability to define reliability strategy, reliability targets, and operational metrics β and defend them to engineering leadership and the business.
-
Strong communication skills β you can explain a root cause to a junior engineer and a reliability risk to a product lead.
-
Demonstrated bias for action and ability to operate autonomously in ambiguous, fast-changing environments.
Requirements
-
Drive the reliability and observability initiative.
-
Own the reliability roadmap end to end. Prove a repeatable define β emit β ingest β dashboard β alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution.
-
Harden the foundational platform.
-
Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability.
-
Expand observability breadth and depth.
-
Extend instrumentation across the six target systems β Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow).
-
Implement a scalable incident and review process.
-
Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow.
-
Scale automation, auditability, and reduce toil.
-
Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise.
-
Build specialized SRE agents using Cursor AI.
-
Design and ship AI agents for incident triage, log analysis, and root-cause investigation.
-
Host SRE agents on the AI fabric.
-
Partner with the AI platform team to deploy your agents on the org's AI fabric.
Company Description