Research Analyst - AI System Performance Modeling @SemiAnalysis
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted Today

[Hiring] Research Analyst - AI System Performance Modeling @SemiAnalysis

Today - SemiAnalysis is hiring a remote Research Analyst - AI System Performance Modeling. 💸 Salary: unspecified 📍Location: South Korea

Role Description

We are seeking an AI System Performance Analyst to model the inference and training performance of AI accelerators and rack-scale systems across real-world model workloads. This role sits at the intersection of computer architecture, machine-learning systems, and market analysis. Your work will directly support the development of our Inference Simulator , InferenceX , Tokenomics Model , and AI Cloud TCO research.

The central question you will answer repeatedly and rigorously is:

  • For a given model, context length, latency target, and parallelism strategy, how many tokens per second can each accelerator and system actually deliver—and what does each token ultimately cost?

You will translate chip- and system-level performance characteristics into defensible technical and economic conclusions for institutional investors, hyperscalers, semiconductor companies, and other industry participants. This role is location-flexible. Candidates based in Seoul or the broader APAC region are preferred, but location is not a requirement.

Responsibilities

  • Inference Performance Modeling
    • Build and extend first-principles performance models for large language model inference.
    • Model the differences between prefill and decode workloads.
    • Analyze arithmetic intensity, compute utilization, memory traffic, and roofline performance limits.
    • Model KV-cache capacity, memory-bandwidth requirements, and context-length scaling.
    • Evaluate batching behavior and the trade-offs between throughput, latency, and interactivity.
    • Develop performance curves across different service-level objectives and deployment configurations.
  • Accelerator and Hardware Analysis
    • Model performance across NVIDIA and AMD GPUs, Google TPUs, AWS Trainium, and emerging AI accelerators.
    • Compare accelerator architectures based on compute throughput, memory bandwidth, memory capacity, on-chip SRAM, interconnect, and system topology.
    • Assess how hardware design choices affect real-world inference and training performance.
    • Evaluate scale-up and scale-out limitations across chips, nodes, racks, and datacenter clusters.
    • Benchmark the relative strengths and weaknesses of heterogeneous accelerator platforms.
  • Model Architecture and Workload Analysis
    • Model dense transformer and mixture-of-experts architectures.
    • Analyze long-context, reasoning, multimodal, and other computationally demanding workloads.
    • Evaluate speculative decoding, quantization formats, sparsity, and other inference-optimization techniques.
    • Model disaggregated prefill and decode serving architectures.
    • Assess how model architecture, parameter count, active parameter count, sequence length, and precision affect system performance.
    • Track changes in model architecture that materially influence hardware requirements and deployment economics.
  • Parallelism and Rack-Scale Systems
    • Analyze tensor, pipeline, expert, and data-parallel strategies.
    • Evaluate how parallelism schemes map onto different accelerator and rack-scale architectures.
    • Model collective communication overheads, synchronization costs, and scaling efficiency.
    • Assess rack-scale systems such as NVL72-class platforms and comparable architectures.
    • Evaluate the effects of scale-up fabrics, network topology, link bandwidth, and congestion on delivered performance.
    • Identify system bottlenecks that prevent theoretical accelerator performance from being achieved in production.
  • Benchmarking and Validation
    • Validate performance models against published and independently gathered benchmarks.
    • Analyze benchmark data from frameworks and deployments using vLLM, SGLang, TensorRT-LLM, PyTorch, JAX, and similar platforms.
    • Reconcile differences between theoretical performance, vendor claims, benchmark results, and production deployments.
    • Identify methodological weaknesses, hidden assumptions, and configuration differences across benchmark datasets.
    • Develop reproducible benchmarking and analytical workflows.
    • Continuously refine model assumptions using new hardware disclosures, software improvements, and real-world performance data.
  • Performance Economics and TCO
    • Translate technical performance into economic metrics.
    • Model tokens per second per accelerator, server, rack, megawatt, and dollar of capital expenditure.
    • Evaluate tokens per watt and the impact of utilization on operating costs.
    • Connect hardware performance to datacenter power, cooling, networking, and infrastructure requirements.
    • Contribute performance inputs to AI cloud TCO and token-cost models.
    • Assess the economic implications of hardware selection, model architecture, latency targets, and deployment strategy.
    • Help institutional clients understand the cost curves of AI training and inference.
  • Research and Collaboration
    • Publish technical deep dives on accelerator, inference, training, and system performance.
    • Serve as a technical authority during client calls, briefings, and research discussions.
    • Communicate complex performance findings clearly to both engineering and investment audiences.
    • Collaborate with SemiAnalysis’ accelerator, networking, memory, datacenter, and market analysts.
    • Connect chip-level performance analysis to system-level, financial, and industry conclusions.
    • Contribute to major newsletters, research reports, client projects, and proprietary analytical products.

Requirements

  • 2–5+ years of experience in ML systems engineering, accelerator or GPU performance engineering, computer architecture, or performance-focused technical analysis.
  • Strong quantitative understanding of transformer inference and training workloads.
  • Ability to calculate and model:
    • FLOPs per token.
    • Memory traffic per token.
    • KV-cache capacity and bandwidth requirements.
    • Arithmetic intensity.
    • Batching effects.
    • Context-length scaling.
    • Model-architecture and hardware interactions.
  • Working knowledge of modern accelerator architectures and memory systems.
  • Understanding of HBM bandwidth and capacity trade-offs, on-chip SRAM, memory hierarchy, and data movement.
  • Familiarity with scale-up interconnects such as NVLink, UALink, or comparable technologies.
  • Understanding of scale-out networking and distributed-system performance.
  • Proficiency in Python for performance modeling, data analysis, simulation, and reproducible analytical tooling.
  • Hands-on familiarity with at least one serving or training framework, such as:
    • vLLM.
    • SGLang.
    • TensorRT-LLM.
    • PyTorch.
    • JAX.
  • Ability to independently define a modeling problem, identify the necessary data, build the analysis, validate the findings, and produce a defensible conclusion.
  • Strong written and verbal communication skills.
  • Ability to explain highly technical concepts clearly to both engineering and investor audiences.
  • Self-driven working style and the ability to operate effectively with minimal oversight.

Preferred Skills

  • Experience writing or optimizing GPU kernels using CUDA, Triton, HIP, or similar programming environments.
  • Experience profiling and optimizing production inference or training deployments.
  • Familiarity with kernel-level bottlenecks, operator fusion, memory access patterns, and hardware utilization.
  • Experience with training-performance modeling, including:
    • Model FLOPs utilization.
    • Parallelism-scaling efficiency.
    • Gradient and activation checkpointing.
    • Communication overhead.
    • Pipeline bubbles.
    • Failure-recovery and checkpointing overhead.
  • Experience benchmarking across heterogeneous hardware platforms.
  • Familiarity with non-NVIDIA accelerators, including TPUs, Trainium, AMD GPUs, or emerging custom silicon.
  • Understanding of model-serving infrastructure, schedulers, orchestration, and distributed inference systems.
  • Experience analyzing power consumption, datacenter infrastructure, and cost of ownership.
  • Prior published technical writing, academic research, open-source contributions, or conference presentations related to ML systems, accelerators, or computer architecture.
  • Familiarity with cloud accelerator pricing, utilization economics, and infrastructure-capacity planning.
  • Experience translating engineering performance into financial, market, or investment implications.

Growth Areas

  • This role offers the opportunity to become a leading authority on the performance and economics of AI computing systems.
  • Potential growth areas include:
    • Taking broader ownership of SemiAnalysis’ Inference Simulator, InferenceX, Tokenomics Model, and AI Cloud TCO research.
    • Developing proprietary methodologies for evaluating real-world accelerator and rack-scale performance.
    • Expanding coverage from inference into large-scale training, fine-tuning, reinforcement learning, and multimodal workloads.
    • Building deeper expertise in GPU kernels, compilers, serving frameworks, distributed systems, and performance optimization.
    • Developing industry-leading analysis of emerging accelerators and heterogeneous AI infrastructure.
    • Leading independent benchmarking projects across chips, servers, racks, and cloud platforms.
    • Becoming a recognized external expert on AI accelerator performance, inference economics, and token-cost modeling.
    • Publishing major research reports and presenting findings to investors, hyperscalers, semiconductor companies, and AI infrastructure providers.
    • Working closely with leading engineers, researchers, executives, and infrastructure decision-makers across the AI ecosystem.
    • Mentoring junior analysts and helping expand the firm’s ML systems and performance-analysis capabilities.
    • Progressing into a senior analyst, technical research lead, model owner, or broader AI infrastructure research leadership position.
Before You Apply
remote Be aware of the location restriction for this remote position: South Korea
Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Research Analyst - AI System Performance Modeling @SemiAnalysis
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted Today
Apply for this position
Did not apply
Applied
Sent Follow-Up
Interview Scheduled
Interview Completed
Offer Accepted
Offer Declined
Application Denied
Unlock 125,000+ Remote Jobs
remote Be aware of the location restriction for this remote position: South Korea
Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply
Applied
Sent Follow-Up
Interview Scheduled
Interview Completed
Offer Accepted
Offer Declined
Application Denied
Unlock 125,000+ Remote Jobs
×

Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 ★★★★★ from 500+ reviews
Unlock All Jobs Now

Maybe later