
I am a Research Scientist at Adobe Research (AIR Lab) in San Jose, CA, working on agentic AI, reinforcement learning for large language models, and multimodal document intelligence. I completed my Ph.D. in Computer Science at the University of Maryland, College Park, advised by Dr. Dinesh Manocha, and hold an M.S. in Computer Science from UMD (2021) and a B.E. in Computer Engineering from Netaji Subhas Institute of Technology, Delhi (2018).
My current research centers on audio language modeling and full-duplex speech systems — models that can listen and speak at the same time, reason over acoustic and spoken context in real time, and handle natural conversational phenomena like interruptions, backchannels, and turn-taking, moving beyond the rigid turn-based latency of today’s voice assistants. This builds on my longer-standing interest in speech and audio (retrieval-augmented language modeling for ASR personalization, layout-informed neural speech synthesis) and sits alongside my broader work on agentic AI systems: memory and continual learning for LLM agents, RL/GRPO-based reasoning, and multi-agent orchestration for document and role-play tasks.
During my Ph.D., I worked on multimodal document understanding, information extraction, and long-context modeling across documents, language, video, and speech. I’ve also spent time as a Research Scientist Intern at Meta AI (NLP & Speech), Microsoft Research, and Adobe Research, at Verisk AI as an AI/ML Research Intern, and at Dataminr as an AI Research Intern — see the Experience page for details.
I maintain a full list of my publications on the Research page. I’m always happy to hear about potential collaborations — feel free to reach out by email!
Email / CV / Google Scholar / LinkedIn / GitHub
Updates
| 2026: | Started new research on audio language modeling and full-duplex speech systems. |
| 2026: | ARGUS: Structured Uncertainty Guided Clarification for LLM Agents accepted to ACL 2026. |
| Nov 2025: | Co-organized the RARA Workshop (Grounding Documents with Reasoning, Agents, Retrieval, and Attribution) at ICDM 2025. |
| 2025: | ChartEval accepted to AACL 2025. |
| 2025: | Follow the Flow and DELOC accepted to EMNLP 2025. |
| 2025: | SQLSpace accepted to EMNLP 2025 (Findings). |
| 2025: | ChartLens accepted to ACL 2025. |
| 2025: | DocVoyager accepted to CHI 2025. |
| 2025: | VisDoM and MoDS accepted to NAACL 2025. |
| 2024: | MATSA and DocEdit-v2 accepted to EMNLP 2024. |
| 2024: | Agent-DocEdit accepted to COLM 2024. |
| 2024: | DocPilot accepted to ACL 2024. |
| 2024: | DOC-RAG, Saliency-Aware Interpolative Augmentation, and DocScript accepted to COLING 2024. |
| Sept 2023: | Joined Adobe Research as a Research Scientist in the Agentic Intelligence Research (AIR) Lab. |
| Dec 2023: | Completed my Ph.D. in Computer Science at the University of Maryland, College Park. |
| 2023: | Research Scientist Intern, Visual Document Understanding, Microsoft Research, Redmond. |
| 2023: | PersonaLM accepted to EMNLP 2023 (Findings). |
| 2023: | DocEdit accepted to AAAI 2023; LayerDoc accepted to WACV 2023. |
| 2022–23: | Research Scientist Intern, NLP & Speech Team, Meta AI (Facebook AI), Menlo Park. |
| 2022: | Research Scientist Intern, Document Intelligence Lab, Adobe Research, San Jose. |
| 2022: | MONOPOLY accepted to ACM Multimedia 2022; DocTime accepted to NAACL 2022; DocLayoutTTS accepted to Interspeech 2022. |
| 2022: | AI/ML Research Intern, Verisk AI, New Jersey. |
| 2022: | DocInfer accepted to EMNLP 2022; DocFin accepted to EMNLP 2022 (Findings). |
| 2021: | TIMERS accepted to ACL 2021; Affect2MM accepted to CVPR 2021. |
| 2021: | Research Scientist Intern, Document Intelligence Lab, Adobe Research (remote). |
| 2020: | Multimodal Multitask Financial Risk Forecasting accepted to ACM Multimedia 2020 (Oral). |
| 2020: | AI Research Intern, Dataminr Inc., New York. |