My current work at Adobe Research is directed towards audio language modeling and evaluations for full-duplex speech systems, reinforcement learning for LLM agents, and multimodal omni models. Previously, I have worked on agentic AI and multimodal language understanding — spanning audio and speech, documents, and language.

During my Ph.D. at the University of Maryland, my work centered on multimodal document understanding, information extraction, and long-context modeling. Google Scholar

Audio & Speech: Full-Duplex Language Modeling

I am currently building audio language models for full-duplex speech interaction — systems that can listen and generate speech simultaneously, handle interruptions, backchannels, and turn-taking natively, rather than the rigid turn-based latency of today’s half-duplex voice assistants. This includes benchmarking document grounding, hallucination, and instruction-following in full-duplex voice agents, reasoning-and-retrieval-while-speaking architectures, and efficient parametric memory for omni language models. Papers and project pages will be linked here as they become public.

This builds on earlier work of mine at the intersection of speech, retrieval, and language modeling:

Agentic AI, LLM Reasoning & Reinforcement Learning

  • ARGUS: Structured Uncertainty guided Clarification for LLM Agents Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi, Dinesh Manocha ACL 2026
  • Partial Policy Gradients for RL in LLMs Puneet Mathur, Branislav Kveton, Subhojyoti Mukherjee, Viet Dac Lai In submission, 2026
  • Cluster-R1: Large Reasoning Models Are Instruction-following Clustering Agents Peijun Qing, Puneet Mathur, Nedim Lipka, Varun Manjunatha, Ryan A. Rossi, Franck Dernoncourt, Saeed Hassanpour, Soroush Vosoughi In submission, 2026

Multimodal Document Intelligence, Attribution & RAG

Multimodal AI, Affective Computing & Video Understanding

Early Research: Vision, NLP & Social Media (2018–2020)

Workshop Papers