Audio

Meta SAM Audio: Segment Audio Content

sam-audio
Table of Contents

Meta’s Segment Anything vision began with a simple but powerful idea: make it easy to isolate anything a user cares about using intuitive prompts. With the introduction of SAM Audio, Meta extends that philosophy beyond images and video into one of the most challenging domains in media creation—sound.

SAM Audio is a unified AI model designed to segment sound from complex audio mixtures using natural prompts. Instead of relying on specialized tools or rigid workflows, users can isolate audio sources using text prompts, visual cues from video, or time span selection. This mirrors how people naturally think about audio problems and makes advanced audio separation more accessible than ever.

In this article, we review what SAM Audio is, how it works, the technology behind it, and why it represents a meaningful shift for audio editing, video workflows, and AI-powered creative tools.

 

What Is SAM Audio?

SAM Audio is Meta’s first unified model for open-domain audio separation. It enables users to isolate specific sounds—such as voices, instruments, or environmental noise—from complex audio recordings using intuitive, multimodal prompts.

Traditional audio editing tools are often designed for narrow tasks, such as speech enhancement or music stem separation. SAM Audio takes a different approach. It supports a wide range of real-world audio scenarios within a single model and interaction paradigm.

Examples of what SAM Audio can do include:

  • Isolating vocals or instruments from music recordings
  • Removing background noise from videos filmed outdoors
  • Cleaning up long podcast recordings with recurring noise
  • Extracting sound sources tied to visible objects in video

By unifying these capabilities, SAM Audio reduces fragmentation and lowers the barrier to professional-quality audio editing.

sam-audio-guitar

 

Multimodal Prompting: How Users Interact with Sound

A defining feature of SAM Audio is its support for multiple prompt types that reflect how people naturally describe and locate sound.

Text Prompting

Text prompting allows users to type natural language descriptions such as “dog barking”, “singing voice”, or “traffic noise”. The model uses this description to extract the corresponding sound from the audio mixture.

This approach makes audio separation more intuitive, especially for users who are not audio professionals.

Visual Prompting

Visual prompting leverages video context. Users can click on a person or object in a video that is producing sound—such as a speaker or a musical instrument—and SAM Audio isolates the associated audio.

This is particularly valuable in video editing workflows, where sound sources are often visually identifiable but difficult to separate using audio-only tools.

sam-audio

 

Span Prompting (Time-Based Prompts)

Span prompting allows users to mark time ranges where a target sound occurs. For example, selecting moments where a dog barks throughout a podcast enables the model to remove that sound across the entire recording.

This method is especially useful for long-form content where unwanted sounds repeat over time.

All three prompting methods can be used independently or combined, giving users precise control over audio segmentation.

sam-audio-span-prompt

 

The Technology Behind SAM Audio

At its core, SAM Audio uses a generative modeling framework built on a flow-matching diffusion transformer. The system takes an audio mixture and one or more prompts, encodes them into a shared representation, and generates separated audio tracks.

To support real-world performance, Meta developed a large-scale data engine that combines:

  • Advanced audio mixing techniques
  • Automated multimodal prompt generation
  • Pseudo-labeling pipelines for realistic training data

The training dataset includes both real and synthetic audio mixtures spanning speech, music, and general sound events, ensuring strong generalization across environments.

 

Perception Encoder Audiovisual (PE-AV)

SAM Audio is powered by Perception Encoder Audiovisual (PE-AV), the technical backbone that aligns visual and audio information.

PE-AV extracts frame-level video features and synchronizes them with audio representations over time. This temporal alignment enables SAM Audio to accurately separate sounds that are visually grounded, such as on-screen speakers or instruments.

The model is trained on over 100 million videos using large-scale multimodal contrastive learning, allowing it to handle diverse, real-world scenarios where visual context matters.

 

SAM Audio Judge: Reference-Free Evaluation

Evaluating audio separation quality is challenging, especially when clean reference tracks are unavailable. To address this, Meta introduced SAM Audio Judge, an automatic evaluation model designed to align with human perception.

SAM Audio Judge evaluates outputs across multiple perceptual dimensions, including precision, recall, faithfulness, and overall quality. This enables consistent benchmarking without relying on reference signals.

 

SAM Audio-Bench: A Real-World Benchmark

Meta also released SAM Audio-Bench, a comprehensive benchmark covering speech, music, and general sound effects. Each sample includes rich prompts such as text descriptions, visual masks, and time markers.

Unlike earlier benchmarks based on synthetic mixtures, SAM Audio-Bench uses real-world audio and video, making evaluations more representative of practical use cases.

 

Performance and Known Limitations

Meta reports that SAM Audio achieves state-of-the-art performance across a wide range of audio separation tasks and operates faster than real time, making it suitable for interactive applications.

Some current limitations include:

  • Audio cannot be used as a prompt
  • Prompt-free audio separation is not supported
  • Separating highly similar sound sources remains challenging

Despite these constraints, SAM Audio establishes a new baseline for unified, prompt-based audio segmentation.

 

Key Use Cases

  • Music production: Isolating vocals or instruments
  • Podcasting: Removing recurring background noise
  • Video editing: Linking sound to visible sources
  • Accessibility: Enhancing clarity in complex acoustic environments

 

Quick Takeaways

  • SAM Audio is Meta’s first unified model for audio segmentation
  • Supports text, visual, and time span prompts
  • Powered by the Perception Encoder Audiovisual (PE-AV)
  • Includes reference-free evaluation with SAM Audio Judge
  • Introduces a real-world benchmark with SAM Audio-Bench

 

Conclusion

SAM Audio represents a significant step forward in how AI systems understand and manipulate sound. By applying the Segment Anything philosophy to audio, Meta enables creators and developers to work with sound using natural, intuitive prompts.

Rather than replacing existing tools, SAM Audio provides a flexible foundation that can power the next generation of audio and video editing experiences. As multimodal AI continues to evolve, SAM Audio sets a strong precedent for accessible, high-quality audio separation at scale. While this model is focussed on Audio, Segment Anything worked also on 3D with his model Sam 3D.

 

FAQs

What is SAM Audio?

SAM Audio is a unified AI model that segments sound from complex audio mixtures using text, visual, and time span prompts.

Can SAM Audio work with video?

Yes. Visual prompting allows SAM Audio to isolate sounds associated with visible objects or people in video.

Is SAM Audio open source?

Meta has released the model, benchmarks, and research artifacts for public use and exploration.

Where can I try SAM Audio?

SAM Audio is available through the Segment Anything Playground and for download via Meta’s official repositories.

References

  • Meta AI – Introducing SAM Audio (December 16, 2025)
  • GitHub – facebookresearch/sam-audio
  • Meta AI – Segment Anything Playground

What do you think about Meta SAM Audio: Segment Audio Content? Leave a comment below.

Find Your Perfect AI Tool