Audio Processing in the Age of Large Language Models

Loading...
Thumbnail Image

Files

Publication or External Link

External Link to Data Files

Date

Advisor

Manocha, Dinesh
Duraiswami, Ramani

Citation

Abstract

Understanding spoken language has long been a primary goal in computer language processing. The ability to comprehend audio--including speech, non-speech sounds, and music--is crucial for both AI agents and humans to interact effectively with the world. However, audio processing research has lagged behind other modalities like language and vision. This disparity is due to several factors, including the scarcity of large-scale datasets, advanced architectures, and the need for more effective training methods to address the inherent complexities of audio processing.

Amid these challenges, Large Language Models (LLMs) have emerged as a promising solution, demonstrating impressive capabilities in understanding and reasoning about the world through language. These models have already also shown potential in advancing foundational audio processing tasks such as Automatic Speech Recognition (ASR), cross-modal retrieval, and audio captioning and generation, as well as new and emerging tasks like open-ended question answering.

To this end, we explore innovative methods for enhancing audio understanding, perception, and reasoning in AI agents, with a specific focus on LLMs. Our contributions span six thrusts, each introducing new architectures, algorithms, datasets, or benchmarks that advance the state of the art:

  1. The Flamingo Family of Large Audio-Language Models – We develop the Flamingo series (Audio Flamingo 2, Audio Flamingo 3, Audio Flamingo Next, and Music Flamingo), a family of fully open LALMs capable of fine-grained audio perception, long-form audio understanding, and expert-level reasoning across sound, music, and speech. These capabilities are enabled by new audio-language encoders trained with novel objectives for improved fine-grained perception, training paradigms that scale context to long-form continuous audio, architectural mechanisms for faithful long-horizon temporal reasoning, and methods to enable strong audio reasoning for complex problems—including reinforcement learning with reasoning supervision for expert-level understanding. Together, these designs yield state-of-the-art results across a broad range of audio understanding and reasoning benchmarks, surpassing prior open and closed models on long-form and expert audio reasoning tasks.

  2. Data Generation at Scale – We develop methods to curate real data and generate synthetic audio and QA pairs that scale training corpora by one to two orders of magnitude beyond prior open datasets. Our contributions include AudioSkills, a 4M-instance expert audio-reasoning dataset, and LongAudio, the first long-form audio understanding dataset spanning audio up to 30 minutes--each substantially larger than any prior open resource. In data-scarce regimes, Synthio introduces preference-aligned text-to-audio generation coupled with MixCap-based caption diversification for synthetic augmentation, yielding absolute gains of 0.1%–39% on ten audio classification benchmarks over strong augmentation baselines.

  3. Improved Audio Perception – We propose robust audio encoders that enhance representation quality for LLMs, evaluated against standard audio-perception metrics such as linear-probe and fine-tuning accuracy on audio classification benchmarks, retrieval Recall on audio-caption benchmarks, and zero-shot classification accuracy on held-out sets. Our methods MAST, SLICER, and EH-MAM introduce novel self-supervised learning algorithms for extracting rich features from unlabeled audio, together achieving state-of-the-art performance across a wide range of downstream audio understanding tasks. ReCLAP and AF-CLAP advance audio-language encoders by incorporating both synthetic and real data along with compositionally-aware objectives, yielding consistent improvements over prior work on both retrieval and compositional reasoning. Each model in the Flamingo series is also powered by novel audio encoders trained at scale that unify all forms of audio understanding.

  4. Long-Form Audio Understanding & Reasoning}– We propose datasets and novel training paradigms that enable LLMs to reason over long-form audio, including environmental soundscapes, music, and spoken dialogues. Our Flamingo series of models is one of the first families of open models capable of long audio understanding and reasoning capabilities, backed by our novel curated datasets and training curricula--advancing LongAudioBench accuracy from 31.1% to 64.2% (+33.1%) and surpassing Gemini 2.5 Pro (68.6 vs. 60.4 GPT-4o score).

  5. Benchmarking Expert Audio Reasoning – We propose new benchmarks like MMAU and MMAU-pro that assess expert-level audio reasoning in LLMs, moving beyond traditional classification tasks to more complex auditory comprehension challenges. These benchmarks expose systematic gaps in current LALMs: even top proprietary models reach only 75.76% on MMAU and 58.7% on MMAU-Pro, against human performance in the mid-80s--establishing clear targets for the field.

  6. Advancing Audio Processing for Omni-Modal Understanding and Reasoning – Finally, we investigate methods for advancing audio processing in omni-modal systems that jointly reason over audio, visual, and textual information. The new capability of this work is long-form, fine-grained audio-visual reasoning: we introduce MMOU, the first omni-modal benchmark with an average video duration of 711.6 seconds—far exceeding prior benchmarks—and AV-Flamingo, a fully open audio-visual LALM purpose-built for long and complex real-world videos. AV-Flamingo is enabled by a large-scale audio-visual dataset explicitly curated for cross-modal learning rather than relying on separate unimodal corpora, a multi-stage curriculum that progressively extends context from short-range perception to long-horizon multi-event reasoning, and a temporally grounded chain-of-thought framework that anchors intermediate reasoning steps to timestamps in the audio and visual streams. Together, these contributions move omni-modal evaluation and modeling from short-clip recognition toward temporally extended, audio-grounded reasoning.

Notes

Rights