Skip to content

Add Mixat dataset (Emirati-English code-switched speech)#713

Open
kftadros wants to merge 1 commit into
ARBML:mainfrom
kftadros:add-mixat-dataset
Open

Add Mixat dataset (Emirati-English code-switched speech)#713
kftadros wants to merge 1 commit into
ARBML:mainfrom
kftadros:add-mixat-dataset

Conversation

@kftadros

Copy link
Copy Markdown
Contributor

Adds the Mixat corpus: 15 hours of Emirati Arabic–English code-switched speech collected from two public podcasts featuring native Emirati speakers, with manual transcriptions (Al Ali & Aldarmaki, LREC-COLING 2024 / SIGUL).

Validated locally against schema.json with validate_schema.py (passed). Recorded the venue as LREC-COLING since SIGUL isn't in the controlled venue list; happy to adjust any fields on request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant