Multimodal Dataset of Brazilian Sign Language (Libras) on Health and Safety

Funded by the Mozilla Foundation, Data-Pop Alliance developed a multimodal dataset of Brazilian Sign Language (Libras), comprising a collection of video recordings and associated metadata, to address the critical underrepresentation of Latin American sign languages in AI training corpora and advance inclusive AI and accessibility research for Deaf communities in Brazil.

With a focus on healthcare communication, the dataset includes content related to medical terminology, symptoms, preventive care, mental health, and clinical interactions, while maintaining an intersectional and gender-diverse approach. The dataset aims to: (1) support the development and improvement of AI systems, including sign language recognition, machine translation, and multimodal alignment across video, text, and audio; (2) enable public-interest applications, particularly in the areas of health and safety; and (3) contribute to the development of more inclusive and representative technologies.