A label such as airport or park leaves out much of what a recording contains. AVSD-Scenes adds natural-language descriptions to urban audio-visual recordings so researchers can study the events and context that a scene category alone does not express.
The AVSD-Scenes paper describes 12,291 paired audio-visual scene descriptions derived from the TAU Urban Audio-Visual Scenes 2021 development dataset. The captions are generated by models, not a corpus of human-written ground truth. The authors evaluate whether those descriptions retain useful information and where their coverage falls short.
Two modalities, then a fused description
The underlying TAU recordings come from ten European cities and contain synchronized audio and visual sequences. The construction pipeline processes each modality separately before combining their descriptions.
Qwen2-Audio-7B-Instruct identifies sound events and produces an audio description. Qwen2.5-VL-7B-Instruct identifies visible objects and actions and produces a visual description. The visual instructions ask the model to avoid unobservable intentions and describe visible characteristics when it cannot identify something confidently.
The authors then compare Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506 and Gemma-3-27B-IT as fusion models. Each receives the events and descriptions from both modalities and produces a single paragraph of about 30 to 60 words. Instructions ask it to stay grounded in those inputs and avoid unsupported details.
Scene labels appear in the main pipeline's prompts. That choice is relevant to the classification results, so the authors also conduct an experiment that removes labels from the instructions.
Strong scene-level results, weaker exact matching
The evaluation uses embedding models including CLIP, CLAP and ImageBind to measure semantic alignment and retrieval. Fusing the descriptions improves text-to-audio scene retrieval in the reported tests, while text-to-visual retrieval remains high and similar to the modality-specific descriptions.
For scene classification, the authors follow the existing 70% training and 30% test split and train a support vector machine on embeddings. They report up to 94.5% accuracy with description embeddings alone and 95.4% when descriptions are combined with audio and visual embeddings. These are accuracies for the evaluated urban scene categories, not measures of whether every detail in a caption is true.
Exact-instance retrieval is much weaker. The paper reports that descriptions tend to capture scene-level semantics rather than clip-specific details, while recordings from the same scene category can resemble one another. A representation that helps identify a park therefore need not find one particular park recording.
The no-label experiment retains strong scene-discriminative information, though the authors report higher overall performance with label guidance. That supports a content-derived signal without removing the distinction between the two prompting conditions.
Human assessment leaves room for better grounding
Four participants rate Mistral-generated descriptions for 100 randomly selected samples balanced across classes. On the five-point scale, mean visual fidelity is 4.14 and fluency 4.26, while audio fidelity is 3.51 and overall quality 3.36. The hallucination criterion scores 3.22, with higher values indicating fewer unsupported details.
That small assessment complements the automatic model-judge evaluation rather than certifying the entire dataset. The authors identify reducing hallucinations and scaling the dataset as future work.
AVSD-Scenes gives researchers a defined pipeline and dataset for testing how text can connect sound with visual context. Its reported strengths concern broad scene representation; applications needing exact events or individual-clip identification must account for the remaining grounding and retrieval limits.
Comments