(01) Collection
Binaural and ambisonic recordings captured across Stuttgart courtyards, squares and street canyons at varied times of day.
A working corpus of urban soundscape excerpts collected for surrogate background sound generation. Each clip carries a location context, a coarse source label, and perceptual ratings used to condition text-to-audio retrieval and mixing experiments.
Open on Hugging FaceAu-Yeung, H. H. (2026). videssonus/mydataset [Data set]. Hugging Face.
| SC-0011 | Courtyard | Birds | 10 | 48.2 | 4.3 |
| SC-0042 | Street canyon | Traffic | 10 | 68.7 | 1.9 |
| SC-0067 | Square | Voices | 10 | 61.4 | 3.1 |
| SC-0104 | Courtyard | Voices | 10 | 55 | 3.6 |
| SC-0138 | Park | Birds | 10 | 44.1 | 4.6 |
| SC-0159 | Street canyon | Construction | 10 | 74.3 | 1.4 |
| SC-0201 | Square | Traffic | 10 | 66.2 | 2.2 |
| SC-0246 | Park | Water | 10 | 52.8 | 4.4 |
| SC-0288 | Courtyard | Traffic | 10 | 59.9 | 2.6 |
| SC-0311 | Square | Music | 10 | 63.5 | 3.9 |
Click a column header to sort the sample
Clips in sample · computed from the bundled sample
Binaural and ambisonic recordings captured across Stuttgart courtyards, squares and street canyons at varied times of day.
Each clip is tagged with a dominant source class, a context string, and listener ratings for pleasantness and eventfulness.
Conditioning and evaluation of surrogate background sound generators; not a calibrated noise-level reference.