Most Trusted Speech Data Collection Services for your AI

Train your NLP models, VAs, TTS prototypes, and more with quality conversational data, with our audio and speech data collection services

Speech data collection
Countries
0 +
Hours of
Speech Data
0 +
Projects
0 +
Languages (100+ Dialects)
0 +

8/16/44/48 kHz

Sampling rate

Professional Audio / Voice Data Collection Services

Any subject. Any scenario.

At Shaip, our expertise lies in creating high-quality speech datasets designed for varied AI/ML requirements. We offer an expansive range of languages and record in diverse settings making our datasets comprehensive and adaptable. Our focus is on feeding models with the highest volume of custom speech data, in the least possible time. With us on board, you can expect: 

Speech collection
  • Curated high-quality multilingual audio / voice data to improve accuracy
  • Highest possible level of domain specificity to target diverse scenario setup
  • Scale your ML model to suit diverse demographics and verticals
  • Recording Environments: Studio Quality, featuring crystal-clear audio with minimal background noise, & Natural Environments, where recordings incorporate ambient sounds to mimic real-world situations.

Our Expertise

Align Audio Data to for Smarter NLP Models

Shaip offers end-to-end speech/audio data collection services in over 100+ languages to enable voice-enabled technologies to cater to a diverse set of audiences across the globe. We can work on projects of any scope and size; from licensing existing off-the-shelf audio datasets, to managing custom audio data collection, to audio transcription and annotation. No matter how big is your speech data collection project, we can customize the audio collection services to suit your needs to build high-quality NLP datasets that target dialects, tones, and languages. Choose from our wide range of speech datasets and audio data collection resources, for voice-enabling intelligent setups.

Monologue speech

Monologue Scripted & Spontaneous Speech

It focuses on processing speech from a single speaker. Utilize scripted prompts to feed into single-channel audio files, ensuring the capture of unique speech patterns, tones, and nuances specific to that individual.

Dialogue speech

Dialogue Scripted & Spontaneous Speech

Two-person interaction, replicating real-world conversations and dialogues with multilingual exposure via dual-channel files and transcribed resources.

Multi-party conversations

Group / Muti-party
Conversations

Multi-person discussions, capturing group dynamics, overlaps, and varied tones so as to accurately train speech models.

Wake-word utterances collection

Wake-word / Key Phrase / Utterances Collection​

Train AIs to identify key phrases or wake words or utterances with similar meanings using diverse, rich, and authentic utterances for advanced natural language processing and understanding.

Acoustic speech

Acoustic Data
Collection

We can professionally record studio-quality audio data be it restaurants, offices, or homes or from various environments and languages, whilst covering a wider acoustic range (Comprehensive Sound Datasets).

Automatic speech recognition

Automatic Speech Recognition (ASR)

Improve accuracy of your automatic speech recognition (ASR) systems by having access to state-of-art diversified speech/audio datasets, from a wide array of demographics.

Natural language utterance

Multilingual Speech/Audio Training Data

Our skilled language professionals, across the globe offer multilingual audio/speech data in various languages and dialects. This effort fosters global communication and bridges language barriers, contributing to more inclusive and effective AI solutions.

Digital virtual assistants

Text-to-Speech
(TTS)

Build a text-to-speech (TTS) multilingual model with the help of our global workforce, who help you collect speech data in 150+ languages & dialects to enhance your AI models from in-car controls to chatbots and learning solutions with high-quality audio data.

Call center recordings

Call Center
Conversations

Genuine exchanges between agents and clients, supporting numerous languages such as Spanish, German, American English, Bengali, Japanese, Chinese, and Hindi.

Success Stories

Conversational AI datasets with over 3k hours of data across 8 languages

Looking to build a multilingual platform for Indian languages, the client partnered with Shaip to collect, segment and transcribe large datasets in multiple Indian languages. This would help develop effective speech models that could power the client’s innovative new platform.

Problem: Over 3,000 hours of audio data collected in 8 Indian languages, segmented and transcribed to develop automatic speech recognition.

Solution: We provided data collection, segmentation, transcription, and delivered JSON files with metadata. We collected 3000 hours of audio data in 8 Indian languages at scale for the client’s speech technology project.

Speech data collection case study

Reasons to choose Shaip as your Trustworthy Speech Data Collection Partner

People

People

Dedicated and trained teams:

  • 30,000+ collaborators for Data Creation, Labeling & QA
  • Credentialed Project Management Team
  • Experienced Product Development Team
  • Talent Pool Sourcing & Onboarding Team
Process

Process

Highest process efficiency is assured with:

  • Robust 6 Sigma Stage-Gate Process
  • A dedicated team of 6 Sigma black belts – Key process owners & Quality compliance
  • Continuous Improvement & Feedback Loop
Platform

Platform

The patented platform offers benefits:

  • Web-based end-to-end platform
  • Impeccable Quality
  • Faster TAT
  • Seamless Delivery

Off-the-Shelf Speech / Audio Datasets

Language DatasetSample RateDataset TypeTotal Audio Hours
+African American Vernacular8 kHzCall-center211
Short DescriptionAfrican American Vernacular Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 612, Male: 1242, and Unknown: 12
+16 kHzPodcast154
Short DescriptionAfrican American Vernacular Media data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 151, Male: 150, and Unknown: 10
+Afrikaans8 kHzGeneral Conversation368
Short DescriptionAfrikaans General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Afrikaans spoken in Africa
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 502, Male: 390, and Unknown: 2
+16 kHzPodcast658
Short DescriptionAfrikaans Media Files
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 750, Male: 1278, and Unknown: 52
+Arabic8 kHzGeneral Conversation292
Short DescriptionArabic General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Arabic from Gulf countries
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 171, Male: 534, and Unknown: 1
+48 kHzScripted Monologue1,947
Short DescriptionArabic Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 838 Male 1209 Unknown 78
+Assamese (In Pipeline) Call-Center60
Short DescriptionAssamese (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionAssamese (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionAssamese (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Bengali (In Pipeline) Call-Center60
Short DescriptionBengali (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionBengali (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionBengali (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Boston English8 kHzCall-Center177
Short DescriptionBoston Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 605, Male: 711, and Unknown: 0
+8 kHzGeneral Conversation32
Short DescriptionBoston General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 53, Male: 83, and Unknown: 0
+16 kHzPodcast93
Short DescriptionBoston Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 43, Male: 181, and Unknown: 2
+Canadian French48 kHzScripted Monologue1,222
Short DescriptionCanadian French
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 974 Male 631 Unknown 1
+Chinese English8 kHzCall-Center169
Short DescriptionChinese Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 1790, Male: 523 and Unknown: 13
+16 kHzPodcast249
Short DescriptionChinese Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 126, Male: 346 and Unknown: 6
+Chinese Simplified48 kHzScripted Monologue2,762
Short DescriptionChinese Simplified
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1920 Male 1535 Unknown 270
+Chinese Traditional48 kHzScripted Monologue1,028
Short DescriptionChinese Traditional
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1069 Male 262 Unknown 3
+Danish8 kHzGeneral Conversation372
Short DescriptionDanish General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 311, Male: 417, Unknown: 0
+16 kHzPodcast664
Short DescriptionDanish Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale: 369, Male: 864, Unknown: 27
+48 kHzScripted Monologue2,579
Short DescriptionDanish Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range, Danish from Denmark
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1551 Male 1233 Unknown 42
+English Deep South8 kHzCall-Center151
Short DescriptionEnglish Deep South Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 221 , Male 1004 , Unknown 7
+8 kHzGeneral Conversation56
Short DescriptionEnglish Deep South General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 99, Male 31, Unknown 0
+16 kHzPodcast266
Short DescriptionEnglish Deep South Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 204, Male 356, Unknown 21
+German8 kHzCall-Center64
Short DescriptionGerman Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelMono
Recording PlatformDesktop
WER (%)
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 478 Male 1440 Unknown 0
+8 kHz IVR200
Short DescriptionGerman IVR data
Dataset DescriptionHuman to Machine. An IVR type of flow where there is a TTS prompt (e.g. ”How may I help you”) followed by a spontaneous human response
Audio ChannelMono
Recording PlatformDesktop
WER (%)
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 10115 Male 8750 Unknown 0
+Gujarati (In Pipeline) Call-Center60
Short DescriptionGujarati (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionGujarati (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionGujarati (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Hebrew8 kHzGeneral Conversation399
Short DescriptionHebrew General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Hebrew in Israel
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 414 , Male 399 , Unknown 1
+16 kHzPodcast427
Short DescriptionHebrew Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 361 , Male 513, Unknown 13
+Hindi16 kHzPodcast219
Short DescriptionHindi Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 83 , Male 309, Unknown 0
+48 kHzScripted Monologue2,867
Short DescriptionHindi Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1977 Male 1864 Unknown 147
+Hinglish8 kHzCall-Center208
Short DescriptionHINGLISH Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 822, Male 1262 , Unknown 0
+16 kHzPodcast216
Short DescriptionHINGLISH Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 75, Male 380, Unknown 0
+Hispanic English8 kHzCall-Center212
Short DescriptionHispanic Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 822, Male 1262, Unknown 0
+16 kHzPodcast155
Short DescriptionHispanic Call Media audio
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 140, Male 219, Unknown 5
+Indonesian8 kHzGeneral Conversation496
Short DescriptionIndonesian General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Bahasa Indonesian
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 524, Male 454, Unknown 2
+16 kHzPodcast643
Short DescriptionIndonesian Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 746, Male 1507, Unknown 129
+Irish8 kHzGeneral Conversation192
Short DescriptionIrish General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 213 , Male 153 , Unknown 0
+Japanese48 kHzScripted Monologue2,335
Short DescriptionJapanese Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1460 Male 1221 Unknown 194
+Kannada (In Pipeline) Call-Center60
Short DescriptionKannada (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionKannada (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionKannada (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Korean8 kHzCall-Center107
Short DescriptionKorean Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1086, Male 210 , Unknown 4
+16 kHzPodcast204
Short DescriptionKorean media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 70 Male 303, Unknown 25
+48 kHzScripted Monologue1,955
Short DescriptionKorean Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1195 Male 1134 Unknown 122
+Malay8 kHzGeneral Conversation266
Short DescriptionMalay General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, Malay in Malaysia
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 316, Male 176 , Unknown 0
+16 kHzPodcast344
Short DescriptionMalay Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 236, Male 626, Unknown 47
+Malayalam (In Pipeline) Call-Center60
Short DescriptionMalayalam (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionMalayalam (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionMalayalam (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Marathi (In Pipeline) Call-Center60
Short DescriptionMarathi (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionMarathi (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionMarathi (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Spanish (Mexico)48 kHzScripted Monologue1,492
Short DescriptionMexican Spanish Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1016 Male 1069 Unknown 95
+Dutch48 kHzScripted Monologue1,205
Short DescriptionDutch Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1285 Male 531 Unknown 3
+New York English8 kHzCall-Center103
Short DescriptionNew York English Call-center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 610, Male 532, Unknow 0
+8 kHzGeneral Conversation107
Short DescriptionNew York English General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 118, Male 114, Unknown 0
+16 kHzPodcast140
Short DescriptionNew York English Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 66, Male 230, Unknown 11
+New Zealand English 8 kHzGeneral Conversation148
Short DescriptionNew Zealand English General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 167, male 121, Unknown 4
+16 kHzPodcast400
Short DescriptionNew Zealand English Media audio
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 367, male 678, Unknown 26
+Oriya (In Pipeline) Call-Center60
Short DescriptionOriya (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionOriya (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionOriya (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Polish16 kHzPodcast269
Short DescriptionPolish Media audio
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 173 Male 354 Unknown 6
+Polish (Poland)48 kHzScripted Monologue1,482
Short DescriptionPolish Poland - Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1324 Male 701 Unknown 24
+Punjabi (In Pipeline) Call-Center60
Short DescriptionPunjabi (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionPunjabi (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionPunjabi (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Russian48 kHzScripted Monologue2,398
Short DescriptionRussian Scripted Monologue
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1689 Male 1937 Unknown 214
+Scottish (English Accent)8 kHzGeneral Conversation292
Short DescriptionScottish General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 285 , Male 260, Unknown 3
+Singapore English8 kHzCall-Center218
Short DescriptionSingapore Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 2139 , Male 884, Unknown 21
+16 kHzPodcast247
Short DescriptionSingapore Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 160, Male 455, Unknown 37
+South African English8 kHzCall-Center261
Short DescriptionSouth African English Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1274 , Male 935 , Unknown 1
+16 kHzPodcast251
Short DescriptionSouth African English Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 235, Male 432, Unknown 36
+Swahili8 kHzCall-Center230
Short DescriptionSwahili Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 611, Male 833, Unknown 0
+16 kHzPodcast265
Short DescriptionSwahili Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 118, Male 493, Unknown 25
+Swedish8 kHzCall-Center250
Short DescriptionSwedish Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1581, male 727, Unknown 2
+16 kHzPodcast278
Short DescriptionSwedish Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 195, male 500, Unknown 21
+Tamil (In Pipeline) Call-Center60
Short DescriptionTamil (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation100
Short DescriptionTamil (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast40
Short DescriptionTamil (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Telugu8 kHzGeneral Conversation553
Short DescriptionTelugu General Conversation data
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 574 , Male 564, Unknown 0
+16 kHzPodcast648
Short DescriptionTelugu Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 207, Male 963, Unknown 2
+Telugu (In Pipeline) Call-Center30
Short DescriptionTelugu (In Pipeline) Call-Center data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+General Conversation50
Short DescriptionTelugu (In Pipeline) General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio Channel
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Podcast20
Short DescriptionTelugu (In Pipeline) Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio Channel
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of Speakers
+Thai8 kHzGeneral Conversation183
Short DescriptionThai General Conversation
Dataset DescriptionUnscripted telephonic conversation between two people. Approx. Audio Duration (Range) - 15-60 minutes, An informal register used between friends
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 338, Male 96, Unknown 8
+16 kHzPodcast173
Short DescriptionThai Media audio
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 143, Male 502, Unknown 26
+Turkish Turkey48 kHzScripted Monologue2,027
Short DescriptionTurkish Turkey
Dataset DescriptionSingle-utterance recordings, which tend to fall in the 5 to 30 second range
Audio ChannelMono
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 1561 Male 1241 Unknown 31
+Vietnamese8 kHzGeneral Conversation295
Short DescriptionVietnamese General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes, Northern (e.g.,Hanoi), Central, and Southern (e.g., Ho Chi Minh City).
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 400, male 380, Unknowns 2
+16 kHzPodcast257
Short DescriptionVietnamese Media audio data
Dataset DescriptionLicensable Public domain audio/video files such as interviews, podcasts etc - 1 to 5 people. Approx. Audio Duration (Range) 15-60 minutes
Audio ChannelMono
Recording PlatformWeb Sourcing
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 249, male 200, Unknowns 45
+Welsh (English Accent)8 kHzGeneral Conversation278
Short DescriptionWelsh General Conversation data
Dataset DescriptionUnscripted, synthetic telephonic conversation between "agent" and "customer", Approx. Audio Duration (Range) 5-15 Minutes,
Audio ChannelDual
Recording PlatformDesktop
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersFemale 270, Male 324, Unknown 0
+UK English16 kHzWake Word200 Speakers
Short DescriptionWake Word UK English
Dataset Descriptionkeyphrases collection of data
  • 200 speakers
  • 4 unique keyphrases per speaker
  • 25-30 repeated keyphrases recordings per unique keyphrase
  • 25-30 audio files per unique keyphrase
  • 120 total recorded utterances per speaker
Audio Channel1 channel
Recording PlatformMobile App
WER (%)5.0
Audio Format.wav
Transcription Format.json
Use CaseASR, Virtual Assistant, Chatbot, Conversational AI, Speech Analytics, TTS, Language Modelling
Number of SpeakersGender: 50% male, 50% female, +/- 10%.

Services Offered

Expert text data collection isn’t all-hands-on-deck for comprehensive AI setups. At Shaip, you can even consider the following services to make models way more widespread than usual:

Text data collection

Text Data Collection Services

The true value of Shaip cognitive data collection services is that it gives organizations the key to unlock critical information found within unstructured data

Image data collection

Image Data Collection Services

Make sure that your computer vision model identifies every image accurately, to seamlessly train next-gen AI models of the future

Video data collection

Video Data Collection Services

Now focus on computer vision along with NLP for training your models to identify objects, individuals, deterrents, and other visual elements to perfection

Featured Clients

Empowering teams to build world-leading AI products.

Shaip contact us

Want to build your own audio dataset?

Connect with our in-house speech data collection expert to set up an audio repository that best fits your requirement

  • This field is for validation purposes and should be left unchanged.
  • By registering, I agree with Shaip Privacy Policy and Terms of Service and provide my consent to receive B2B marketing communication from Shaip.

Speech Data Collection for an ML Model refers to the process of gathering audio recordings of spoken language. This collection aids in training and refining machine learning algorithms, particularly those centered on understanding and processing human voices.

When aiming to collect audio data for Automatic Speech Recognition (ASR), you should start by defining your project’s specific needs, including the desired language, accent, and type of speech. After setting these parameters, ensure you obtain all necessary permissions to respect user privacy. Then, use appropriate recording devices or software to capture clear audio samples. Each recording should be meticulously annotated with its transcription or other pertinent metadata and stored systematically for effortless access.

A speech dataset in machine learning is pivotal for training, testing, and validating models tailored to recognize, transcribe, or interpret spoken language. Such datasets pave the way for a myriad of applications, from voice assistants and transcription services to voice biometrics.

For collecting precise data from diverse languages and accents, collaboration with native speakers of the desired linguistic backgrounds is vital. Aim for a varied and representative sample to cover a broad spectrum of demographic nuances. Employ standardized recording equipment in uniform environments to ensure audio consistency. And importantly, annotate each data piece with detailed transcriptions and metadata, denoting the specific language and accent.