My research focuses on agentic AI and multimodal language understanding — spanning audio and speech, documents, and language. I work on reinforcement learning for LLM agents, audio/speech language modeling (including full-duplex spoken dialogue systems), multimodal document intelligence, and retrieval-augmented generation over visually-rich content (tables, charts, and flowcharts).

During my Ph.D. at the University of Maryland, my work centered on multimodal document understanding, information extraction, and long-context modeling. My current work at Adobe Research is directed toward agentic memory and continual learning, RL/GRPO-based reasoning, and audio language modeling for full-duplex speech systems.

Google Scholar

Audio & Speech: Full-Duplex Language Modeling

I am currently building audio language models for full-duplex speech interaction — systems that can listen and generate speech simultaneously, handle interruptions, backchannels, and turn-taking natively, rather than the rigid turn-based latency of today’s half-duplex voice assistants. This includes benchmarking document grounding, hallucination, and instruction-following in full-duplex voice agents, reasoning-and-retrieval-while-speaking architectures, and efficient parametric memory for omni language models. Papers and project pages will be linked here as they become public.

  • Duplex-R1: Full-Duplex Audio LMs that Reason, Retrieve, and Speak While Searching Puneet Mathur, Nedim Lipka, Zeyu Jin, Dinesh Manocha In submission, 2026 · (Coming soon)
  • DuplexSpeechBench–Document Grounding: Benchmarking Document Grounding and Hallucinations in Full-Duplex Voice Agents Puneet Mathur, Nedim Lipka, Zeyu Jin, Dinesh Manocha In submission, 2026 · (Coming soon)
  • DuplexSpeechBench–IFEval: Instruction Following Evaluation for Full-Duplex Voice Agents Puneet Mathur, Manan Suri, Dinesh Manocha In submission, 2026 · (Coming soon)
  • Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models Puneet Mathur, Manan Suri, Dinesh Manocha In submission, 2026 · (Coming soon)

This builds on earlier work of mine at the intersection of speech, retrieval, and language modeling:

Agentic AI, LLM Reasoning & Reinforcement Learning

  • ARGUS: Structured Uncertainty guided Clarification for LLM Agents Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi, Dinesh Manocha ACL 2026
  • Partial Policy Gradients for RL in LLMs Puneet Mathur, Branislav Kveton, Subhojyoti Mukherjee, Viet Dac Lai In submission, 2026
  • Cluster-R1: Large Reasoning Models Are Instruction-following Clustering Agents Peijun Qing, Puneet Mathur, Nedim Lipka, Varun Manjunatha, Ryan A. Rossi, Franck Dernoncourt, Saeed Hassanpour, Soroush Vosoughi In submission, 2026
  • SETA: Self Evolving Tutor Agents via Conversational Prompt Optimization Puneet Mathur, Jack Wang, Nedim Lipka In submission, 2025

Multimodal Document Intelligence, Attribution & RAG

Multimodal Finance, Affective Computing & Video Understanding

Early Research: Vision, NLP & Social Media (2018–2020)

Workshop Papers