Optimize AI workloads on NVIDIA GPUs and Google TPUs.
Profile and improve latency, throughput, memory usage, and accelerator utilization.
Build and optimize kernels using CUDA and Pallas.
Benchmark GPU vs. TPU performance across different models and workloads.
Identify bottlenecks in PyTorch and JAX execution.
Explore kernel fusion, quantization, attention optimization, KV-cache optimization, and memory efficiency.
Reproduce promising AI systems research and validate it on real hardware.
Build reliable and reproducible benchmarking workflows.
Turn successful experiments into production-ready improvements.