Senior Solutions Architect, Generative AI @NVIDIA
Artificial Intelligence
Salary usd 184,000 - 3..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago

[Hiring] Senior Solutions Architect, Generative AI @NVIDIA

2mths ago - NVIDIA is hiring a remote Senior Solutions Architect, Generative AI. πŸ’Έ Salary: usd 184,000 - 356,500 per year πŸ“Location: USA

Role Description

NVIDIA is looking for an AI Solutions Architect with deep, hands-on experience in large-scale GPU systems. This role involves working with some of the world’s leading consumer internet companies and frontier labs building foundation models. Primary responsibilities include:

  • Accelerating customer workloads.
  • Designing high-performance AI infrastructure.
  • Leading technical engagements around NVIDIA technologies.
  • Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs.
  • Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability.
  • Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack.
  • Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.
  • Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed.
  • Partnering with NVIDIA’s engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback.

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience.
  • 6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role.
  • Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters.
  • Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using technologies such as InfiniBand, RoCE, or GPUDirect RDMA.
  • Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues.
  • Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers.
  • Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems.
  • Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.

Requirements

  • Experience architecting and operating large-scale production GPU clusters for distributed training or inference.
  • Deep expertise with NVIDIA infrastructure technologies such as DGX/HGX systems, NVLink, NVSwitch, NCCL, InfiniBand, and Spectrum-X.
  • Hands-on experience using tools and telemetry such as NCCL tests, DCGM, Nsight Systems, fabric counters, and host- or switch-level diagnostics to isolate performance and reliability issues.
  • Understanding of network topology, congestion control, collective communication patterns, and their impact on distributed AI workload performance.
  • Experience optimizing storage and data pipelines to sustain high-throughput training and inference workloads.

Benefits

  • Competitive salaries.
  • Generous benefits package.
  • Equity eligibility.

Company Description

NVIDIA is recognized as one of the technology world’s most sought-after employers. This role offers a chance to make a broad impact at NVIDIA by advancing innovation with our consumer internet & frontier labs partners. We are committed to fostering an inclusive work environment and proud to be an equal opportunity employer.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Solutions Architect, Generative AI @NVIDIA
Artificial Intelligence
Salary usd 184,000 - 3..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,809+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later