Principal AI Infrastructure & Accelerator Architect @EPAM Systems
Artificial Intelligence
Salary unspecified
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted 1wk ago

[Hiring] Principal AI Infrastructure & Accelerator Architect @EPAM Systems

1wk ago - EPAM Systems is hiring a remote Principal AI Infrastructure & Accelerator Architect. 💸 Salary: unspecified 📍Location: USA

Role Description

Are you a visionary technical leader passionate about defining the future of AI infrastructure and squeezing every drop of performance out of advanced hardware accelerators at scale? We are seeking a Principal Architect to lead, shape, and execute our technical strategy for AI performance, optimization, and hardware-software co-design.

In this elite, highly visible role, you will define the architectural vision for both AI training and serving infrastructure, delivering massive industry-wide impact. You will spearhead our Center of Excellence (CoE), scaling our practice and guiding the technical roadmap across next-generation Tensor Processing Units (TPUs), Graphics Processing Unit (GPU) fleets, state-of-the-art ML models, and advanced compiler toolchains.

Your architectural decisions will directly enable cutting-edge AI research and large-scale production deployments across Google Cloud, major enterprise customers, and the broader open-source ecosystem. If you thrive on solving intractable performance bottlenecks and redefining what is physically and computationally possible in AI infrastructure, this is your platform.

Responsibilities

  • Define and drive the multi-year technical roadmap for high-performance AI kernels, custom operations, and hardware-software co-design targeting TPU and GPU architectures.
  • Scale and mentor a world-class technical practice, establishing architectural governance, engineering standards, and best practices across the organization.
  • Act as the principal technical liaison partnering with ML researchers, core framework architects (JAX, PyTorch), and compiler engineering teams (XLA, MLIR) to eliminate systemic bottlenecks and shape future hardware/software requirements.
  • Architect foundational infrastructure—including enterprise-grade benchmarking suites, automated autotuning frameworks, regression analysis pipelines, and comprehensive documentation—empowering the global developer community.
  • Anticipate industry shifts by tracking advancements in hardware architectures, emerging model topologies, and compiler innovations to unlock step-changes in AI training and inference efficiency.

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or equivalent practical experience (Master's or Ph.D. preferred).
  • 15+ years of software engineering experience, with 8+ years focused on distributed systems, AI infrastructure, or high-performance computing (HPC) architecture.
  • 7+ years of experience designing and developing complex software systems in C++ or Python.
  • 5+ years of experience leading the architecture, design, and delivery of large-scale software products, frameworks, or developer ecosystems from inception to production.
  • Proven track record of architecting performance-critical systems at the kernel level, bridging hardware accelerators and high-level software frameworks.

Requirements

  • Deep expertise in optimizing TPU/GPU execution, leveraging low-level kernel languages/abstractions such as Pallas, Mosaic, Triton, or CUDA.
  • Comprehensive knowledge of modern ML frameworks (JAX, PyTorch), attention mechanisms, Mixture of Experts (MoEs), model quantization, and low-precision arithmetic.
  • Advanced understanding of modern accelerator architectures, including heterogeneous compute, complex memory hierarchies, data movement optimization, and multi-node scale-out fabrics.
  • Deep familiarity with compiler principles, code generation, and modern toolchains such as MLIR, OpenXLA, and LLVM.
  • Demonstrated leadership in building and scaling developer infrastructure, widely adopted Open-Source Software (OSS) libraries, and extensible high-performance APIs.
  • Exceptional strategic communication and stakeholder management skills, with a history of influencing cross-functional engineering teams, researchers, and executive leadership.
Before You Apply
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Principal AI Infrastructure & Accelerator Architect @EPAM Systems
Artificial Intelligence
Salary unspecified
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
×
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 ★★★★★ from 500+ reviews

⚡ 129,404+ remote jobs, refreshed hourly

🔔 Real-time alerts: Apply first

🛡️ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later