AI Solution Architect @Uvation
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 2wks ago

[Hiring] AI Solution Architect @Uvation

2wks ago - Uvation is hiring a remote AI Solution Architect. πŸ’Έ Salary: unspecified πŸ“Location: India

Role Description

We are seeking an experienced AI Solution Architect to design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise in NVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning.

Key Responsibilities

  • Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
  • Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.
  • Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.
  • Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
  • Design AI storage and data architectures using object storage, parallel file systems like Ceph, WEKA, or equivalent platforms.
  • Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
  • Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
  • Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.
  • Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.
  • Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.
  • Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.
  • Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.

Required Technical Skills

  • AI / ML Architecture
    • NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.
    • PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.
    • GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.
    • LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.
  • GPU & AI Factory Infrastructure
    • NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.
    • NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.
    • DGX/HGX/OEM GPU server architecture and lifecycle management.
    • AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.
  • High-Performance Networking
    • 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris.
    • NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.
    • BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.
    • GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.
  • AI Storage & Data Architecture
    • Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.
    • Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.
    • Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.
    • GPUDirect Storage and storage/network performance optimization.
  • AI Platform & Orchestration
    • Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.
    • HPC or other equivalent workload schedulers.
    • Model serving/inference platforms and MLOps platform architecture.
    • API gateways, service discovery, secrets management and platform integration.
  • Cloud & Hybrid Architecture
    • AWS and/or Azure AI infrastructure and security services.
    • Hybrid cloud connectivity, IAM, private networking, cloud storage and workload placement.
    • Cloud cost optimization, capacity planning and FinOps considerations for GPU workloads.
  • Security & Governance
    • Zero Trust, network segmentation, IAM/RBAC, PAM and workload identity.
    • GPU, DPU, container, Kubernetes, firmware and supply-chain security.
    • Encryption at rest/in transit, secrets management, audit logging and compliance controls.
    • AI-specific risks including data/model protection, tenant isolation and secure model access.
  • Observability & Reliability
    • Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.
    • Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.
    • High availability, backup/restore, disaster recovery, business continuity and failure-domain design.
    • Performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.

Architecture Deliverables

  • AI Factory reference architecture and solution blueprints.
  • High-Level Design (HLD) and Low-Level Design (LLD).
  • Network, compute, GPU and storage architecture diagrams.
  • Capacity, performance and scalability models.
  • Technology evaluation and vendor comparison documents.
  • Security architecture and threat-model inputs.
  • Bill of Materials (BOM) and infrastructure sizing.
  • Migration/deployment strategy and implementation roadmap.
  • Operational readiness checklist, runbooks and acceptance criteria.

Experience & Qualifications

  • 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.
  • Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.
  • Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.
  • Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.

Preferred Certifications

  • NVIDIA certifications or equivalent GPU/AI infrastructure credentials.
  • AWS Solutions Architect / Azure Solutions Architect.
  • TOGAF or equivalent enterprise architecture certification.
  • CCNP/CCIE or equivalent networking certification.
  • CISSP or equivalent security certification.
  • Kubernetes certifications such as CKA/CKAD.
  • Red Hat / Linux certifications.
Before You Apply
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
AI Solution Architect @Uvation
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,845+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later