Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching
Generate 16 kHz audio scenes — blending speech, music, sound effects, and environmental acoustics — from structured text descriptions.