Hooman Ramezani

I'm Hooman, an ML Specialist Solutions Architect at Nebius. I work with clients to get the best performance when migrating large-scale ML workloads onto next-generation GPUs, profiling and optimizing training runs and tuning inference engines like vLLM and SGLang on NVIDIA B300 and GB300 clusters.

My expertise spans GPU performance optimization across large-scale ML workloads, including CUDA-level profiling, kernel and memory optimization, low-precision compute, and multi-node communication. Recently I co-authored a joint blog post with Meta's PyTorch team on 41% faster DeepSeek-V3 pre-training, presented at NVIDIA GTC 2026.

Before Nebius, I did my BASc in Systems Design Engineering at the University of Waterloo and my MASc at the University of Toronto. My master's thesis research focused on training small transformers for clinical segmentation and detection, with first-author work published at CVPR.

In my free time you can find me playing guitar. See more here.

Email  /  CV  /  LinkedIn  /  GitHub

profile photo

Research and Projects

I am interested in large-scale ML systems, GPU performance optimization, and AI for healthcare applications. My research experience spans my current work at Nebius on training and inference optimization, previous work with the UW VIP Lab, and my master's thesis research at UofT.

Time-to-first-token chart, shared KV cache vs HBM-only
Distributed KV Cache for AI Agents: 2.4x More Requests on NVIDIA B300
Hooman Ramezani, Alexey Averin, Ford Shaper, Anton Bykov, 2026
WEKA Blog / Nebius Blog

Led the Nebius side of bringing a distributed KV cache to market as a product on Nebius AI Cloud, giving every GPU server a shared memory of past context for AI agents. On NVIDIA B300 GPUs it served 2.4x more requests on the same hardware, with a 5.4x faster first token.

Sparse mixture-of-experts diagram
NVSHMEM and DeepEP for Wide Expert Parallelism
Hooman Ramezani, Nebius Team, 2026
Blog / Open Solution / DeepEP

Authored Nebius's deep dive on why mixture-of-experts training gets bottlenecked by GPU-to-GPU communication, and how GPU-initiated RDMA with NVSHMEM and DeepEP fixes it. Includes an open-source recipe to run it on Nebius.

DeepSeek-V3 throughput: 651 to 918 tokens/sec
PyTorch Blog: Enabling Up to 41% Faster Pre-training: MXFP8 and DeepEP for DeepSeek-V3 on B200 with TorchTitan
Hooman Ramezani, Meta PyTorch and Nebius Teams, 2026
Blog / Open Solution / TorchTitan

Co-authored a joint blog post with Meta's PyTorch team, presented at NVIDIA GTC 2026. We accelerated pre-training of DeepSeek-V3, a 671B-parameter frontier model, by 41% on 256 NVIDIA B200 GPUs with no loss in model quality, combining Blackwell-native MXFP8 training with DeepEP's optimized kernels for mixture-of-experts communication.

LN-Transformer architecture
LN-Transformer: Lung Nodule Transformer for Sparse CT Segmentation
Hooman Ramezani, Charlotte Vedrines, Dionne Aleman, Daniel Létourneau, CVPR, 2025
Paper / CVF

Published at CVPR 2025, a novel two-stage transformer for lung nodule segmentation. Strongest model on benchmark dataset with Dice 91.4%, F1 94.2%, combining Meta SAM and DETR architectures.

Lung CT scan with detected nodule
Lung-DETR: Deformable Detection Transformer for Sparse Lung Nodule Anomaly Detection
Hooman Ramezani, Dionne Aleman, Daniel Létourneau, arXiv, 2024
arXiv

A novel architecture to detect lung tumors, designed to handle extreme class imbalance and find tumors among vast amounts of healthy tissue.

Information bottleneck attribution heatmap
Enhancing DL Interpretability: IBA for Transformer Attribution
Hooman Ramezani, University of Toronto, 2024
Paper / Presentation

Information Bottleneck Attribution (IBA) leverages principles from information theory to identify critical information in neural networks for decision-making attribution. In this work IBA is successfully applied to CNN and Transformer models, enabling a detailed analysis of model decision-making.

Gait sensor signals
Parkinson's Freezing of Gait Detection
Hooman Ramezani, Medical Time Series Deep Learning, 2023
Paper / GitHub

A deep learning network for time-series analysis designed to identify gait freezing in patients with Parkinson's disease, utilizing biometric signals for the prevention of falls.

Maze pathfinding animation
Rat-Brain-Inspired Reinforcement Learning for Optimal Pathfinding in Mazes
Hooman Ramezani, Computational Neuroscience, 2023
Paper / GitHub

A deep reinforcement learning model inspired by the basal ganglia of mouse brains, designed to master maze navigation using Q-learning. It showcases the intricacies of decision-making and learning as the model identifies optimal paths through mazes.

Predicted grasp points on household objects
Grasp-Proposition-Net: Robotic Vision For Grasping Everyday Objects
Hooman Ramezani, UW VIP Lab, 2022
GitHub

Developed a 3D computer vision model with VIP-Lab and Festo for a robotic arm, designed to determine optimal grasp points using LiDAR camera data.

Applied Brain Research logo
Drone-Aided Surface Defect Detection
Hooman Ramezani, Applied Brain Research, Vision Model with Temporal Context, 2021
GitHub

A highly accurate embedded model for classifying surface defects via drones, utilizing a convolutional-RNN architecture and synthetic data generation. Model is optimized for on-device execution in real-world applications.