Staff Platform Engineer @Alkami Technology
Software Development
Salary usd 140,000 - 1..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago

[Hiring] Staff Platform Engineer @Alkami Technology

2mths ago - Alkami Technology is hiring a remote Staff Platform Engineer. πŸ’Έ Salary: usd 140,000 - 175,000 per year πŸ“Location: USA

Role Description

The Staff Platform Engineer leads the effort to locate, make visible, and remediate sources of unreliability in the MANTL platform, including correctness problems that surface under failure conditions. This role works directly in the platform's application codebase (TypeScript) and its container-native deployment environment, combining application-engineering skill with reliability-engineering practice at a level of scope and independence beyond the Sr Platform Engineer. In addition to resolving reactive issues as they surface, this role owns and prioritizes a standing roadmap of known reliability risks, driving that work forward on its own timeline. This role partners with, but is organizationally and functionally distinct from, both Cloud Infrastructure Engineering and product application engineering teams, and is expected to set technical direction and standards for platform reliability work.

Essential Duties & Responsibilities

  • Lead investigation, troubleshooting, and resolution of the most complex reliability issues within MANTL platform application code, including microservice communication failures and correctness issues that emerge under failure conditions.
  • Set direction for how the platform identifies and addresses failure modes across its third-party and internal system integrations, establishing resilience strategies suited to each integration's specific behavior.
  • Own the design and evolution of monitoring, dashboards, and alerting strategy (Datadog preferred) across the platform, ensuring proactive visibility into emerging risks.
  • Establish standards and lead implementation of distributed tracing across microservices to accelerate root-cause identification organization-wide.
  • Lead diagnosis and remediation of significant application performance issues, including caching strategy, inefficient code paths, and query performance.
  • Set direction for platform and application hardening practices, including fault-injection and resilience testing, and drive adoption of defensive design patterns to prevent recurrence.
  • Own and improve CI/CD build pipeline architecture (GitHub Actions) supporting deployment of the platform.
  • Guide deployment and troubleshooting of container-native (Kubernetes) workloads as part of resolving complex platform reliability issues.
  • Define, prioritize, and drive execution of a roadmap of known reliability risks, independent of feature-delivery timelines.
  • Establish and maintain documentation and runbook standards covering platform reliability issues, root causes, and remediations.
  • Serve as the senior escalation point for complex platform reliability issues, partnering with Cloud Infrastructure Engineering and application engineering leadership on issues that cross domain boundaries.
  • Define reliability targets (SLOs/SLIs) for key platform services and advise leadership on reliability risk and tradeoffs.
  • Provide technical mentorship and guidance to other Platform Engineers.

Qualifications

  • 7 to 10 years of experience in software engineering, platform engineering, or a hybrid development/reliability engineering role.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience.

Requirements

  • Deep proficiency in TypeScript, with significant production experience in a Node.js/TypeScript service environment.
  • Strong hands-on experience deploying, operating, and troubleshooting containerized workloads in Kubernetes.
  • Proven experience designing, building, and maintaining CI/CD pipelines, GitHub Actions preferred.
  • Proven experience designing monitoring, dashboard, and alerting strategy in an APM/observability tool, Datadog preferred.
  • Deep experience troubleshooting distributed system and microservice communication failures across a variety of integration points and protocols.
  • Strong experience with distributed tracing tools and practices at scale.
  • Working familiarity with relational databases, sufficient to diagnose and resolve complex query and schema-level performance issues.
  • Proven track record resolving significant application performance issues (caching, inefficient code paths, slow queries).
  • Experience designing and leading failure-mode testing (fault injection, resilience/chaos-style testing) and driving adoption of defensive patterns such as idempotency, retries, and circuit breakers.
  • Demonstrated ability to work independently on ambiguous, high-impact reliability problems and to set technical direction for others.
  • Excellent communication skills, with the ability to present technical root cause, risk, and remediation plans to engineering and non-technical leadership.
  • Experience mentoring other engineers.

Preferred

  • Experience with message brokers or event-streaming platforms such as Kafka, particularly at scale.
  • Experience with OpenTelemetry or comparable distributed tracing frameworks at scale.
  • Experience in a regulated or compliance-driven environment (fintech, banking, or similar).
  • Familiarity with infrastructure-as-code tooling (Terraform or similar).
  • Experience partnering with infrastructure/SRE teams on issues that cross application and infrastructure boundaries.
  • Experience influencing engineering roadmap prioritization to secure time for platform stability work.

Benefits

  • Remote-first environment.
  • Unlimited paid time off.
  • 401(k) with employer match.
  • Diverse and inclusive culture.

Salary

The salary range for this position is: $140,000 - $175,000.

Work Authorization

We cannot offer employment sponsorship at this time. Candidates must be eligible to work in the US for full-time employment.

Company Description

Alkami Technology is an Equal Opportunity Employer and Prohibits Discrimination and Harassment of Any Kind. Alkami is committed to the principle of equal employment opportunity for all employees and to providing employees with a work environment free of discrimination and harassment.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Staff Platform Engineer @Alkami Technology
Software Development
Salary usd 140,000 - 1..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,920+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later