| + | African American Vernacular | 8 kHz | Call-center | 211 |
| Short Description | African American Vernacular Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 612, Male: 1242, and Unknown: 12 |
|
| + | | 16 kHz | Podcast | 154 |
| Short Description | African American Vernacular Media data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 151, Male: 150, and Unknown: 10 |
|
| + | Afrikaans | 8 kHz | General Conversation | 368 |
| Short Description | Afrikaans General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Afrikaans spoken in Africa | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 502, Male: 390, and Unknown: 2 |
|
| + | | 16 kHz | Podcast | 658 |
| Short Description | Afrikaans Media Files | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 750, Male: 1278, and Unknown: 52 |
|
| + | Arabic | 8 kHz | General Conversation | 292 |
| Short Description | Arabic General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Arabic from Gulf countries | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 171, Male: 534, and Unknown: 1 |
|
| + | | 48 kHz | Scripted Monologue | 1,947 |
| Short Description | Arabic Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 838 Male 1209 Unknown 78 |
|
| + | Assamese (In Pipeline) | | Call-Center | 60 |
| Short Description | Assamese (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Assamese (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Assamese (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Bengali (In Pipeline) | | Call-Center | 60 |
| Short Description | Bengali (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Bengali (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Bengali (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Boston English | 8 kHz | Call-Center | 177 |
| Short Description | Boston Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 605, Male: 711, and Unknown: 0 |
|
| + | | 8 kHz | General Conversation | 32 |
| Short Description | Boston General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 53, Male: 83, and Unknown: 0 |
|
| + | | 16 kHz | Podcast | 93 |
| Short Description | Boston Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 43, Male: 181, and Unknown: 2 |
|
| + | Canadian French | 48 kHz | Scripted Monologue | 1,222 |
| Short Description | Canadian French | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 974 Male 631 Unknown 1 |
|
| + | Chinese English | 8 kHz | Call-Center | 169 |
| Short Description | Chinese Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 1790, Male: 523 and Unknown: 13 |
|
| + | | 16 kHz | Podcast | 249 |
| Short Description | Chinese Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 126, Male: 346 and Unknown: 6 |
|
| + | Chinese Simplified | 48 kHz | Scripted Monologue | 2,762 |
| Short Description | Chinese Simplified | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1920 Male 1535 Unknown 270 |
|
| + | Chinese Traditional | 48 kHz | Scripted Monologue | 1,028 |
| Short Description | Chinese Traditional | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1069 Male 262 Unknown 3 |
|
| + | Danish | 8 kHz | General Conversation | 372 |
| Short Description | Danish General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 311, Male: 417, Unknown: 0 |
|
| + | | 16 kHz | Podcast | 664 |
| Short Description | Danish Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female: 369, Male: 864, Unknown: 27 |
|
| + | | 48 kHz | Scripted Monologue | 2,579 |
| Short Description | Danish Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range, Danish from Denmark | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1551 Male 1233 Unknown 42 |
|
| + | English Deep South | 8 kHz | Call-Center | 151 |
| Short Description | English Deep South Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 221 , Male 1004 , Unknown 7 |
|
| + | | 8 kHz | General Conversation | 56 |
| Short Description | English Deep South General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 99, Male 31, Unknown 0 |
|
| + | | 16 kHz | Podcast | 266 |
| Short Description | English Deep South Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 204, Male 356, Unknown 21 |
|
| + | German | 8 kHz | Call-Center | 64 |
| Short Description | German Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Mono | | Recording Platform | Desktop | | WER (%) | | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 478 Male 1440 Unknown 0 |
|
| + | | 8 kHz | IVR | 200 |
| Short Description | German IVR data | | Dataset Description | Human to Machine. An IVR type of flow where there is a TTS prompt (e.g. ”How may I help you”) followed by a spontaneous human response | | Audio Channel | Mono | | Recording Platform | Desktop | | WER (%) | | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 10115 Male 8750 Unknown 0 |
|
| + | Gujarati (In Pipeline) | | Call-Center | 60 |
| Short Description | Gujarati (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Gujarati (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Gujarati (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Hebrew | 8 kHz | General Conversation | 399 |
| Short Description | Hebrew General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Hebrew in Israel | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 414 , Male 399 , Unknown 1 |
|
| + | | 16 kHz | Podcast | 427 |
| Short Description | Hebrew Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 361 , Male 513, Unknown 13 |
|
| + | Hindi | 16 kHz | Podcast | 219 |
| Short Description | Hindi Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 83 , Male 309, Unknown 0 |
|
| + | | 48 kHz | Scripted Monologue | 2,867 |
| Short Description | Hindi Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1977 Male 1864 Unknown 147 |
|
| + | Hinglish | 8 kHz | Call-Center | 208 |
| Short Description | HINGLISH Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 822, Male 1262 , Unknown 0 |
|
| + | | 16 kHz | Podcast | 216 |
| Short Description | HINGLISH Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 75, Male 380, Unknown 0 |
|
| + | Hispanic English | 8 kHz | Call-Center | 212 |
| Short Description | Hispanic Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 822, Male 1262, Unknown 0 |
|
| + | | 16 kHz | Podcast | 155 |
| Short Description | Hispanic Call Media audio | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 140, Male 219, Unknown 5 |
|
| + | Indonesian | 8 kHz | General Conversation | 496 |
| Short Description | Indonesian General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Bahasa Indonesian | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 524, Male 454, Unknown 2 |
|
| + | | 16 kHz | Podcast | 643 |
| Short Description | Indonesian Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 746, Male 1507, Unknown 129 |
|
| + | Irish | 8 kHz | General Conversation | 192 |
| Short Description | Irish General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 213 , Male 153 , Unknown 0 |
|
| + | Japanese | 48 kHz | Scripted Monologue | 2,335 |
| Short Description | Japanese Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1460 Male 1221 Unknown 194 |
|
| + | Kannada (In Pipeline) | | Call-Center | 60 |
| Short Description | Kannada (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Kannada (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Kannada (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Korean | 8 kHz | Call-Center | 107 |
| Short Description | Korean Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1086, Male 210 , Unknown 4 |
|
| + | | 16 kHz | Podcast | 204 |
| Short Description | Korean media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 70 Male 303, Unknown 25 |
|
| + | | 48 kHz | Scripted Monologue | 1,955 |
| Short Description | Korean Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1195 Male 1134 Unknown 122 |
|
| + | Malay | 8 kHz | General Conversation | 266 |
| Short Description | Malay General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Malay in Malaysia | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 316, Male 176 , Unknown 0 |
|
| + | | 16 kHz | Podcast | 344 |
| Short Description | Malay Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 236, Male 626, Unknown 47 |
|
| + | Malayalam (In Pipeline) | | Call-Center | 60 |
| Short Description | Malayalam (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Malayalam (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Malayalam (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Marathi (In Pipeline) | | Call-Center | 60 |
| Short Description | Marathi (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Marathi (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Marathi (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Spanish (Mexico) | 48 kHz | Scripted Monologue | 1,492 |
| Short Description | Mexican Spanish Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1016 Male 1069 Unknown 95 |
|
| + | Dutch | 48 kHz | Scripted Monologue | 1,205 |
| Short Description | Dutch Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1285 Male 531 Unknown 3 |
|
| + | New York English | 8 kHz | Call-Center | 103 |
| Short Description | New York English Call-center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 610, Male 532, Unknow 0 |
|
| + | | 8 kHz | General Conversation | 107 |
| Short Description | New York English General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 118, Male 114, Unknown 0 |
|
| + | | 16 kHz | Podcast | 140 |
| Short Description | New York English Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 66, Male 230, Unknown 11 |
|
| + | New Zealand English | 8 kHz | General Conversation | 148 |
| Short Description | New Zealand English General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 167, male 121, Unknown 4 |
|
| + | | 16 kHz | Podcast | 400 |
| Short Description | New Zealand English Media audio | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 367, male 678, Unknown 26 |
|
| + | Oriya (In Pipeline) | | Call-Center | 60 |
| Short Description | Oriya (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Oriya (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Oriya (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Polish | 16 kHz | Podcast | 269 |
| Short Description | Polish Media audio | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 173 Male 354 Unknown 6 |
|
| + | Polish (Poland) | 48 kHz | Scripted Monologue | 1,482 |
| Short Description | Polish Poland - Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1324 Male 701 Unknown 24 |
|
| + | Punjabi (In Pipeline) | | Call-Center | 60 |
| Short Description | Punjabi (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Punjabi (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Punjabi (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Russian | 48 kHz | Scripted Monologue | 2,398 |
| Short Description | Russian Scripted Monologue | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1689 Male 1937 Unknown 214 |
|
| + | Scottish (English Accent) | 8 kHz | General Conversation | 292 |
| Short Description | Scottish General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 285 , Male 260, Unknown 3 |
|
| + | Singapore English | 8 kHz | Call-Center | 218 |
| Short Description | Singapore Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 2139 , Male 884, Unknown 21 |
|
| + | | 16 kHz | Podcast | 247 |
| Short Description | Singapore Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 160, Male 455, Unknown 37 |
|
| + | South African English | 8 kHz | Call-Center | 261 |
| Short Description | South African English Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1274 , Male 935 , Unknown 1 |
|
| + | | 16 kHz | Podcast | 251 |
| Short Description | South African English Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 235, Male 432, Unknown 36 |
|
| + | Swahili | 8 kHz | Call-Center | 230 |
| Short Description | Swahili Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 611, Male 833, Unknown 0 |
|
| + | | 16 kHz | Podcast | 265 |
| Short Description | Swahili Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 118, Male 493, Unknown 25 |
|
| + | Swedish | 8 kHz | Call-Center | 250 |
| Short Description | Swedish Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1581, male 727, Unknown 2 |
|
| + | | 16 kHz | Podcast | 278 |
| Short Description | Swedish Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 195, male 500, Unknown 21 |
|
| + | Tamil (In Pipeline) | | Call-Center | 60 |
| Short Description | Tamil (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 100 |
| Short Description | Tamil (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 40 |
| Short Description | Tamil (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Telugu | 8 kHz | General Conversation | 553 |
| Short Description | Telugu General Conversation data | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 574 , Male 564, Unknown 0 |
|
| + | | 16 kHz | Podcast | 648 |
| Short Description | Telugu Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 207, Male 963, Unknown 2 |
|
| + | Telugu (In Pipeline) | | Call-Center | 30 |
| Short Description | Telugu (In Pipeline) Call-Center data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | General Conversation | 50 |
| Short Description | Telugu (In Pipeline) General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | | | Podcast | 20 |
| Short Description | Telugu (In Pipeline) Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | |
|
| + | Thai | 8 kHz | General Conversation | 183 |
| Short Description | Thai General Conversation | | Dataset Description | Unscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, An informal register used between friends | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 338, Male 96, Unknown 8 |
|
| + | | 16 kHz | Podcast | 173 |
| Short Description | Thai Media audio | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 143, Male 502, Unknown 26 |
|
| + | Turkish Turkey | 48 kHz | Scripted Monologue | 2,027 |
| Short Description | Turkish Turkey | | Dataset Description | Single-utterance recordings, which tend to fall in the 5 to 30 second range | | Audio Channel | Mono | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 1561 Male 1241 Unknown 31 |
|
| + | Vietnamese | 8 kHz | General Conversation | 295 |
| Short Description | Vietnamese General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, Northern (e.g.,Hanoi), Central, and Southern (e.g., Ho Chi Minh City). | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 400, male 380, Unknowns 2 |
|
| + | | 16 kHz | Podcast | 257 |
| Short Description | Vietnamese Media audio data | | Dataset Description | Licensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes | | Audio Channel | Mono | | Recording Platform | Web Sourcing | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 249, male 200, Unknowns 45 |
|
| + | Welsh (English Accent) | 8 kHz | General Conversation | 278 |
| Short Description | Welsh General Conversation data | | Dataset Description | Unscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, | | Audio Channel | Dual | | Recording Platform | Desktop | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Female 270, Male 324, Unknown 0 |
|
| + | UK English | 16 kHz | Wake Word | 200 Speakers |
| Short Description | Wake Word UK English | | Dataset Description | keyphrases collection of data
- 200 speakers
- 4 unique keyphrases per speaker
- 25-30 repeated keyphrases recordings per unique keyphrase
- 25-30 audio files per unique keyphrase
- 120 total recorded utterances per speaker
| | Audio Channel | 1 channel | | Recording Platform | Mobile App | | WER (%) | 5.0 | | Audio Format | .wav | | Transcription Format | .json | | Use Case | ASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling | | Number of Speakers | Gender: 50% male, 50% female, +/- 10%. |
|