Over one million QA pairs pairing binaural audio with panoramic depth images and room impulse responses, supporting geometry-aware spatial audio reasoning (from the OWL project).
A comprehensive multimodal dataset combining audio, video, and sensor data with question-answer pairs for cross-modal understanding research.
LLM-assisted softly-labelled IMU sensor data captioning dataset with feature summarizations and narrations of human activities.
Open-source sensor-aware question-answering or instruction-tuning dataset covering IMU data interpretations for human activities.
Large-scale dataset for multichannel sound source localization research with diverse acoustic scenarios and precise ground truth localization information.
A dataset consisting of STFTs derived from the UrbanSound8K Dataset, and ESC50 Dataset which are either aliasing or nonaliasing.
Audio large language model that infers where a sound is in 3D space — its direction and distance — with spatially grounded reasoning behind each prediction.
State-of-the-art multimodal reasoning model combining audio, video, and sensor data for comprehensive scene understanding and cross-modal reasoning.
Specialized language model fine-tuned for sensor data interpretation and natural language generation, bridging sensors and human-understandable descriptions.
For questions about datasets, models, or collaboration opportunities, please reach out to our team.