Research
My research focuses on agentic AI and multimodal language understanding — spanning audio and speech, documents, and language. I work on reinforcement learning for LLM agents, audio/speech language modeling (including full-duplex spoken dialogue systems), multimodal document intelligence, and retrieval-augmented generation over visually-rich content (tables, charts, and flowcharts).
During my Ph.D. at the University of Maryland, my work centered on multimodal document understanding, information extraction, and long-context modeling. My current work at Adobe Research is directed toward agentic memory and continual learning, RL/GRPO-based reasoning, and audio language modeling for full-duplex speech systems.
Audio & Speech: Full-Duplex Language Modeling
I am currently building audio language models for full-duplex speech interaction — systems that can listen and generate speech simultaneously, handle interruptions, backchannels, and turn-taking natively, rather than the rigid turn-based latency of today’s half-duplex voice assistants. This includes benchmarking document grounding, hallucination, and instruction-following in full-duplex voice agents, reasoning-and-retrieval-while-speaking architectures, and efficient parametric memory for omni language models. Papers and project pages will be linked here as they become public.
- Duplex-R1: Full-Duplex Audio LMs that Reason, Retrieve, and Speak While Searching In submission, 2026 · (Coming soon)
- DuplexSpeechBench–Document Grounding: Benchmarking Document Grounding and Hallucinations in Full-Duplex Voice Agents In submission, 2026 · (Coming soon)
- DuplexSpeechBench–IFEval: Instruction Following Evaluation for Full-Duplex Voice Agents In submission, 2026 · (Coming soon)
- Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models In submission, 2026 · (Coming soon)
This builds on earlier work of mine at the intersection of speech, retrieval, and language modeling:
- PersonaLM: Language Model Personalization via Domain-distributed Span Aggregated K-Nearest N-gram Retrieval Augmentation EMNLP 2023 (Findings)
- DOC-RAG: ASR Language Model Personalization with Domain-Distributed Co-occurrence Retrieval Augmentation COLING 2024
- DocLayoutTTS: Dataset and Baselines for Layout-informed Document-level Neural Speech Synthesis Interspeech 2022
- Risk Forecasting from Earnings Calls Acoustics and Network Correlations Interspeech 2020
- VolTAGE: Volatility Forecasting via Text Audio Fusion with Graph Convolution Networks for Earnings Calls EMNLP 2020
- Meta-learning for Low-Resource Speech Emotion Recognition ICASSP 2021
Agentic AI, LLM Reasoning & Reinforcement Learning
- ARGUS: Structured Uncertainty guided Clarification for LLM Agents ACL 2026
- Partial Policy Gradients for RL in LLMs In submission, 2026
- Cluster-R1: Large Reasoning Models Are Instruction-following Clustering Agents In submission, 2026
- SETA: Self Evolving Tutor Agents via Conversational Prompt Optimization In submission, 2025
Multimodal Document Intelligence, Attribution & RAG
- ChartEval: LLM-Driven Chart Generation Evaluation Using Scene Graph Parsing AACL 2025
- ChartLens: Fine-grained Visual Attribution in Charts ACL 2025
- Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents EMNLP 2025
- DELOC: Document Element Localizer EMNLP 2025
- SQLSpace: A Representation Space for Text-to-SQL to Discover and Mitigate Robustness Gaps EMNLP 2025 (Findings)
- DocVoyager: Anticipating Users' Information Needs and Guiding Document Reading through Question Answering CHI 2025
- VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation NAACL 2025
- MoDS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections NAACL 2025
- MATSA: Multi-Agent Table Structure Attribution EMNLP 2024
- DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding EMNLP 2024
- Agent-DocEdit: Language-Instructed LLM Agent for Content-Rich Document Editing COLM 2024
- DocPilot: Copilot for Automating PDF Edit Workflows in Documents ACL 2024
- DocScript: Document-level Script Event Prediction COLING 2024
- DocEdit: Language-guided Document Editing AAAI 2023
- LayerDoc: Layer-wise Extraction of Spatial Hierarchical Structure in Visually-Rich Documents WACV 2023
- DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection EMNLP 2022
- DocFin: Multimodal Financial Prediction and Bias Mitigation using Semi-structured Documents EMNLP 2022 (Findings)
- DocTime: A Document-level Temporal Dependency Graph Parser NAACL 2022
- TIMERS: Document-level Temporal Relation Extraction ACL 2021
Multimodal Finance, Affective Computing & Video Understanding
- Saliency-Aware Interpolative Augmentation for Multimodal Financial Prediction COLING 2024
- MONOPOLY: Financial Prediction from MONetary POLicY Conference Videos Using Multimodal Cues ACM Multimedia 2022
- 3MASSIV: Multilingual, Multimodal and Multi-Aspect Dataset of Social Media Short Videos Dataset CVPR 2022
- Multimodal Multi-Speaker Merger and Acquisition (M3A) Financial Forecasting: A New Task, Dataset, and Neural Baselines ACL 2021
- Affect2MM: Affective Analysis of Multimedia Content Using Emotion Causality CVPR 2021
- Multitask Learning for Emotionally Analyzing Sexual Abuse Disclosures NAACL 2021
- Dynamic Graph Modeling of Simultaneous EEG and Eye-tracking Data For Reading Task Identification ICASSP 2021
- Multimodal Multitask Financial Risk Forecasting ACM Multimedia 2020 (Oral)
Early Research: Vision, NLP & Social Media (2018–2020)
- Mixup Multi-Attention Multi-Tasking Model for Early-Stage Leukemia Identification ICASSP 2020
- Rethinking Retinal Landmark Localization as Pose Estimation: Naive Single Stacked Network for Optic Disk and Fovea Detection ICASSP 2020
- Utilizing Temporal Psycholinguistic Cues for Suicidal Intent Estimation ECIR 2020 (Short)
- #MeTooMA: Multi-Aspect Annotations of Tweets Related to the MeToo Movement Dataset ICWSM 2020
- Hindi-English Hate Speech Detection: Author Profiling, Debiasing, and Practical Perspectives AAAI 2020 (Oral)
- Exploring Classification of Histological Disease Biomarkers from Renal Biopsy Images Poster · Video WACV 2019
Workshop Papers
- Suicide Risk Assessment via Temporal Psycholinguistic Modeling AAAI Student Abstract and Poster 2020
- An Iterative Approach for Identifying Complaint Based Tweets in Social Media Platforms AAAI Student Abstract and Poster 2020
- SNAP-BATNET: Cascading Author Profiling and Social Network Graphs for Suicide Ideation Detection on Social Media NAACL Student Research Workshop 2019
- Speak Up, Fight Back! Detection of Social Media Disclosures of Sexual Harassment NAACL Student Research Workshop 2019
- Identification of Emergency Blood Donation Request on Twitter Poster SMM4H Workshop, EMNLP 2018
- Did You Offend Me? Classification of Offensive Tweets in Hinglish Language Poster Abusive Language Workshop (ALW2), EMNLP 2018
- Exploring and Learning Suicidal Ideation Connotations on Social Media with Deep Learning Poster WASSA Workshop, EMNLP 2018
- Detecting Offensive Tweets in Hindi-English Code-Switched Language Video Social NLP Workshop, ACL 2018