Puneet Mathur

I am a Research Scientist at Adobe Research (AIR Lab) in San Jose, CA, working on agentic AI, reinforcement learning for large language models, and multimodal document intelligence. I completed my Ph.D. in Computer Science at the University of Maryland, College Park, advised by Dr. Dinesh Manocha, and hold an M.S. in Computer Science from UMD (2021) and a B.E. in Computer Engineering from Netaji Subhas Institute of Technology, Delhi (2018).

My current research centers on audio language modeling and full-duplex speech systems — models that can listen and speak at the same time, reason over acoustic and spoken context in real time, and handle natural conversational phenomena like interruptions, backchannels, and turn-taking, moving beyond the rigid turn-based latency of today’s voice assistants. This builds on my longer-standing interest in speech and audio (retrieval-augmented language modeling for ASR personalization, layout-informed neural speech synthesis) and sits alongside my broader work on agentic AI systems: memory and continual learning for LLM agents, RL/GRPO-based reasoning, and multi-agent orchestration for document and role-play tasks.

During my Ph.D., I worked on multimodal document understanding, information extraction, and long-context modeling across documents, language, video, and speech. I’ve also spent time as a Research Scientist Intern at Meta AI (NLP & Speech), Microsoft Research, and Adobe Research, at Verisk AI as an AI/ML Research Intern, and at Dataminr as an AI Research Intern — see the Experience page for details.

I maintain a full list of my publications on the Research page. I’m always happy to hear about potential collaborations — feel free to reach out by email!

Email  /  CV  /  Google Scholar  /  LinkedIn  /  GitHub

Updates

2026:Started new research on audio language modeling and full-duplex speech systems.
2026:ARGUS: Structured Uncertainty Guided Clarification for LLM Agents accepted to ACL 2026.
Nov 2025:Co-organized the RARA Workshop (Grounding Documents with Reasoning, Agents, Retrieval, and Attribution) at ICDM 2025.
2025:ChartEval accepted to AACL 2025.
2025:Follow the Flow and DELOC accepted to EMNLP 2025.
2025:SQLSpace accepted to EMNLP 2025 (Findings).
2025:ChartLens accepted to ACL 2025.
2025:DocVoyager accepted to CHI 2025.
2025:VisDoM and MoDS accepted to NAACL 2025.
2024:MATSA and DocEdit-v2 accepted to EMNLP 2024.
2024:Agent-DocEdit accepted to COLM 2024.
2024:DocPilot accepted to ACL 2024.
2024:DOC-RAG, Saliency-Aware Interpolative Augmentation, and DocScript accepted to COLING 2024.
Sept 2023:Joined Adobe Research as a Research Scientist in the Agentic Intelligence Research (AIR) Lab.
Dec 2023:Completed my Ph.D. in Computer Science at the University of Maryland, College Park.
2023:Research Scientist Intern, Visual Document Understanding, Microsoft Research, Redmond.
2023:PersonaLM accepted to EMNLP 2023 (Findings).
2023:DocEdit accepted to AAAI 2023; LayerDoc accepted to WACV 2023.
2022–23:Research Scientist Intern, NLP & Speech Team, Meta AI (Facebook AI), Menlo Park.
2022:Research Scientist Intern, Document Intelligence Lab, Adobe Research, San Jose.
2022:MONOPOLY accepted to ACM Multimedia 2022; DocTime accepted to NAACL 2022; DocLayoutTTS accepted to Interspeech 2022.
2022:AI/ML Research Intern, Verisk AI, New Jersey.
2022:DocInfer accepted to EMNLP 2022; DocFin accepted to EMNLP 2022 (Findings).
2021:TIMERS accepted to ACL 2021; Affect2MM accepted to CVPR 2021.
2021:Research Scientist Intern, Document Intelligence Lab, Adobe Research (remote).
2020:Multimodal Multitask Financial Risk Forecasting accepted to ACM Multimedia 2020 (Oral).
2020:AI Research Intern, Dataminr Inc., New York.