Form cover
Page 1 of 1

Research Fellow - GPU & TPU @Deep Variance

Deep Variance Inc. | India

Deep Variance is looking for a Research Fellow – AI Infrastructure to work on performance

engineering for modern AI workloads across GPUs and TPUs.

This is a hands-on research role focused on accelerator benchmarking, kernel optimization, systems evaluation, and inference performance.

What You’ll Work On

Optimize AI workloads on NVIDIA GPUs and Google TPUs.

Profile and improve latency, throughput, memory usage, and accelerator utilization.

Build and optimize kernels using CUDA and Pallas.

Benchmark GPU vs. TPU performance across different models and workloads.

Identify bottlenecks in PyTorch and JAX execution.

Explore kernel fusion, quantization, attention optimization, KV-cache optimization, and memory efficiency.

Reproduce promising AI systems research and validate it on real hardware.

Build reliable and reproducible benchmarking workflows.

Turn successful experiments into production-ready improvements.

Core Skills

Must Have

CUDA

PyTorch

Python

Strong understanding of GPU performance and profiling

Ability to reason about memory, compute, kernel execution, and system bottlenecks

Strong coding and debugging skills

Highly Valued

Pallas / JAX / TPU programming

Triton, CUTLASS, XLA

vLLM, SGLang, TensorRT-LLM

LLM inference and model serving

Experience writing or optimizing custom kernels

AI-Native Engineering

We expect strong use of modern coding agents such as Claude Code and OpenAI Codex for rapid prototyping, codebase exploration, benchmarking, debugging, and research implementation. You should still be able to independently reason about correctness, hardware behavior, and performance trade-offs.

Who Should Apply

We are particularly interested in students, researchers, and recent graduates from IITs, NITs, IISc, and other top-tier universities with strong backgrounds in:

Computer Science

AI / Machine Learning

Systems and Infrastructure

High-Performance Computing

GPU / Accelerator Programming

Compilers or Distributed Systems

Strong technical work matters more than titles. CUDA/Pallas projects, research, open-source contributions, benchmarking work, and strong GitHub profiles will stand out.

Fellowship Details

Duration: 3 months

Stipend: ₹50,000 to ₹80,000 per month depending on experience

Extension: Can be extended beyond 3 months based on performance and project requirements

Location: Remote, India

Deep Variance Inc.

www.deepvariance.com

Application

Full Name

Choose one if you had been to these colleges

A
B
C
D
E
F
G
H
I
J

Email

Resume

Earliest Start Date