Ashutosh Sharma

Ashutosh Sharma

Research Engineer

MIT-IBM Watson AI Lab, IBM Research

I am a Research Engineer at IBM Research (MIT-IBM Watson AI Lab), where I develop and train the Granite family of LLMs at up to 200B scale. I lead the RL infrastructure for Agentic RL of Granite LLMs, enabling multi-scaffold, multi-task, and multi-turn reinforcement learning.

I did my M.S. in Computer Science at UIUC (Siebel Scholar) and my B.Tech (Honors) at IIT Bombay with dual minors in CS and AI. My interests lie at the intersection of mathematics, low-level systems engineering, and machine learning.

Outside of work, I enjoy singing, listening to music, and travelling.

Research Interests

  • Large Language Models & Post-Training
  • Reinforcement Learning & Agentic Systems
  • GPU Kernels & Systems Optimization
  • Distributed Systems
  • Systems Programming
  • Theoretical Machine Learning
  • Information Retrieval & NLP

Education

University of Illinois Urbana-Champaign

M.S. in Computer Science

Aug 2023 – May 2025  |  GPA: 3.79 / 4.0

  • Siebel Scholar — Awarded for academic excellence & leadership among graduate students worldwide
  • Thesis: Temporal Generalization for Cloud Computing
  • Coursework: Distributed Systems, Operating Systems, Communication Networks, Compilers, Databases

Indian Institute of Technology Bombay

B.Tech (Honors) in Mechanical Engineering

Jul 2019 – May 2023  |  GPA: 9.1 / 10.0

  • Dual Minors in Computer Science and Artificial Intelligence
  • Research Award — Recognized for outstanding achievement in Bachelor's thesis research
  • Coursework: Machine Learning, Cryptography, NLP, Computer Vision, Data Structures, Algorithms

Professional Experience

Research Engineer

IBM Research — MIT-IBM Watson AI Lab

June 2025 – Present

Researching & developing Granite LLMs (up to 200B scale) & Agentic RL systems

  • Built distributed async RL trainer with Megatron & FSDP for long-horizon RL; led SFT & RL post-training achieving SOTA reasoning, math, code, & tool-use performance
  • Training Block Diffusion Language Models with Multi-Token Prediction, Multi-Latent Attention, and custom noise schedules & auxiliary loss functions
  • Developed Fused Triton Kernels for efficient dLLM inference & compiled model forward operations achieving 2× faster inference than SOTA baselines

Applied Scientist Intern

Amazon

Summer 2024

Fine-tuned LLMs & developed end-to-end pipeline for product catalog matching using LoRA

  • Designed novel Multi-Task Learning objective & accelerated training via data parallelism & model quantization
  • Fine-tuned a Vision Language Model using Dual Encoder architecture for improving efficiency
  • Achieved 2% improvement in recall@90Precision and 1% in ROC AUC over production model

Software Engineer Intern

Wells Fargo & Co.

Summer 2022

Developed a Trading Platform in Flask & ReactJS; received Pre-Placement Offer

  • Integrated Apache Kafka with MongoDB for persistent storage of stock prices from external APIs
  • Developed a Trade Matching Engine for efficiently queuing and settling trades & an interactive GUI
  • Encrypted data in transit and performed unit testing in Pytest ensuring high code coverage

Selected Publications

Ultrasonics 2025

Towards Improving Breast Cancer Detection through Multi-Modal Image Generation

Sahar Almahfouz Nasser, Ashutosh Sharma, Anmol Saraf, Amruta Parulekar, Pranjal Haria, Amit Sethi

ICLR 2024

DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines

Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Saiful Haq, Ashutosh Sharma, et al.

ACL 2024

IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages

Saiful Haq, Ashutosh Sharma, Omar Khattab, Niyati Chhaya, Pushpak Bhattacharyya

EMNLP 2023

ANGEL: Enterprise Search System for the Non-Profit Industry

Saiful Haq*, Ashutosh Sharma*, Pushpak Bhattacharyya

Key Projects

Fused Triton kernels for ColBERT MaxSim with dimension tiling (d>128) and fused PQ scoring. Reached 80% peak HBM bandwidth on H100 via roofline-guided tiling, 1.9× over PLAID's GPU kernel, 469× exact-scoring throughput over WARP (SIGIR'25 Best Paper) and 8.5× over torch.compile.

Triton CUDA GPU Kernels Information Retrieval

GPU-native inverted index and fused scatter-add kernel scaling exact SPLADE retrieval to 8.8M docs. 235× speedup over Pyserini CPU and 23–270× over SPARe's GPU scatter-add kernel. Characterized the work- vs. bandwidth-efficiency tradeoff governing GPU sparse retrieval kernel design.

Triton CUDA Sparse Retrieval GPU

Causal Ordering in Distributed Shared Logs

Designed a metric for quantifying ordering anomalies in batch-ordered logs. Implemented Windowed Temporal Reordering using Lamport timestamps, reducing causal anomalies by 80% while preserving throughput.

Go Distributed Systems Lamport Timestamps

Totally Ordered Reliable Multicast

Implemented a decentralized total ordering algorithm for transaction processing with fault-tolerance via heartbeat-based failure detection and reliable multicast.

Go Distributed Systems Fault Tolerance

Linux Kernel Enhancements

Developed a Rate Monotonic Scheduler for real-time task scheduling. Optimized memory allocation with a Slab Allocator and implemented a kernel module for profiling page fault rates.

C Linux Kernel Systems Programming

Neural IR for Enterprise Search

Modified attention mechanisms using dependency parsing to encode syntax. Deployed a dual-encoder neural retriever with Elasticsearch, achieving 44.5% MAP@10 and 49.7% MRR@10 gains over ColBERTv2.

Transformers Elasticsearch NLP

TCP over UDP Implementation

Implemented a custom TCP protocol over UDP with ACK-based retransmission and AIMD congestion control including Slow Start, Congestion Avoidance, and Fast Recovery.

C++ Networking Congestion Control

Parallelizing Audio Analysis with FFT

Developed a music analyzer extracting nearest frequencies using Fast Fourier Transform. Parallelized with OpenMP and CUDA for CPU and GPU acceleration.

C++ CUDA OpenMP

Skills

Programming Languages

Python Go C++ C Rust JavaScript MATLAB SQL Java HTML/CSS

Frameworks & Technologies

PyTorch TensorFlow CUDA AWS GCP BigQuery Git Apache Kafka MongoDB Triton vLLM SGLang Megatron FSDP Docker Flask React