Principal AI Infrastructure & Accelerator Architect @EPAM
Artificial Intelligence
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1wk ago

[Hiring] Principal AI Infrastructure & Accelerator Architect @EPAM

1wk ago - EPAM is hiring a remote Principal AI Infrastructure & Accelerator Architect. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

Are you a visionary technical leader passionate about defining the future of AI infrastructure and squeezing every drop of performance out of advanced hardware accelerators at scale? We are seeking a Principal Architect to lead, shape, and execute our technical strategy for AI performance, optimization, and hardware-software co-design. In this elite, highly visible role, you will define the architectural vision for both AI training and serving infrastructure, delivering massive industry-wide impact.

You will spearhead our Center of Excellence (CoE), scaling our practice and guiding the technical roadmap across next-generation Tensor Processing Units (TPUs), Graphics Processing Unit (GPU) fleets, state-of-the-art ML models, and advanced compiler toolchains. Your architectural decisions will directly enable cutting-edge AI research and large-scale production deployments across Google Cloud, major enterprise customers, and the broader open-source ecosystem. If you thrive on solving intractable performance bottlenecks and redefining what is physically and computationally possible in AI infrastructure, this is your platform.

Responsibilities

  • Define and drive the multi-year technical roadmap for high-performance AI kernels, custom operations, and hardware-software co-design targeting TPU and GPU architectures.
  • Scale and mentor a world-class technical practice, establishing architectural governance, engineering standards, and best practices across the organization.
  • Act as the principal technical liaison partnering with ML researchers, core framework architects (JAX, PyTorch), and compiler engineering teams (XLA, MLIR) to eliminate systemic bottlenecks and shape future hardware/software requirements.
  • Architect foundational infrastructure, including enterprise-grade benchmarking suites, automated autotuning frameworks, regression analysis pipelines, and comprehensive documentation, empowering the global developer community.
  • Anticipate industry shifts by tracking advancements in hardware architectures, emerging model topologies, and compiler innovations to unlock step-changes in AI training and inference efficiency.

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or equivalent practical experience (Master's or Ph.D. preferred).
  • 15+ years of software engineering experience, with 8+ years focused on distributed systems, AI infrastructure, or high-performance computing (HPC) architecture.
  • 7+ years of experience designing and developing complex software systems in C++ or Python.
  • 5+ years of experience leading the architecture, design, and delivery of large-scale software products, frameworks, or developer ecosystems from inception to production.
  • Proven track record of architecting performance-critical systems at the kernel level, bridging hardware accelerators and high-level software frameworks.

Requirements

  • Deep expertise in optimizing TPU/GPU execution, leveraging low-level kernel languages/abstractions such as Pallas, Mosaic, Triton, or CUDA.
  • Comprehensive knowledge of modern ML frameworks (JAX, PyTorch), attention mechanisms, Mixture of Experts (MoEs), model quantization, and low-precision arithmetic.
  • Advanced understanding of modern accelerator architectures, including heterogeneous compute, complex memory hierarchies, data movement optimization, and multi-node scale-out fabrics.
  • Deep familiarity with compiler principles, code generation, and modern toolchains such as MLIR, OpenXLA, and LLVM.
  • Demonstrated leadership in building and scaling developer infrastructure, widely adopted Open-Source Software (OSS) libraries, and extensible high-performance APIs.
  • Exceptional strategic communication and stakeholder management skills, with a history of influencing cross-functional engineering teams, researchers, and executive leadership.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Principal AI Infrastructure & Accelerator Architect @EPAM
Artificial Intelligence
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,024+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later