Speech Dataset

English Deep South Dataset

High-Quality English Deep South Call-Center, General Conversation, and Podcast Dataset for AI & Speech Models

English Deep South Dataset

Overview

  • LanguageEnglish Deep South
  • Dataset TypesCall-Center, General Conversation, Podcast
  • CountrySouthern United States

Description

Unscripted synthetic telephonic conversations between an agent and a customer are available with durations of 5 to 15 minutes, while unscripted telephonic conversations between two people typically range from 15 to 60 minutes. Additionally, licensable public domain audio or video files, such as interviews or podcasts with 1 to 5 participants, are also available within the 15 to 60 minute range.

Dataset Highlights

  • LanguageEnglish Deep South
  • Sample Rate16 kHz / 8 kHz
  • Total Hours725:30:27
  • Speakers2,689+
  • ChannelsDual / Mono
  • Duration Range5 - 60 mins

Use Cases

  • ASR (Automatic Speech Recognition)
  • Virtual Assistants & Chatbots
  • Conversational AI
  • Speech Analytics
  • TTS (Text-to-Speech)
  • Language Modelling

Sample Dataset Specs

Dataset TypeCall CenterGeneral ConversationMedia Data
Sampling Rate8 kHz8 kHz16 kHz
Speakers2 Speakers2 SpeakersMultiple Speakers
ChannelDualDualMono
Total Hours266:44:22197:25:07261:20:58
Total # of Speakers6341,490565

Key Benefits

  • Diverse & RepresentativeBroad speaker coverage for inclusive, real-world AI.
  • High-Quality AudioClean, balanced recordings for reliable model performance.
  • Ready for AI/MLStructured and annotated for seamless integration.
  • Scalable & FlexibleMultiple durations, speakers, and domains to fit your needs.

Technical Details

  • Recording PlatformDesktop
  • Audio Format.wav
  • Transcription Format.json
  • WER (%)5
  • Gender DistributionFemale 221, Male 1004, Unknown 7
  • Age Range18-50

Featured Clients

Trusted by leading AI teams and enterprises worldwide.

  • Microsoft
  • Amazon Web Services
  • Google

Can’t find what you are looking for?

New off-the-shelf datasets are being collected across all data types

Contact us now to let go of your audio/speech training data collection worries

  • This field is for validation purposes and should be left unchanged.
  • By registering, I agree with Shaip Privacy Policy and Terms of Service and provide my consent to receive B2B marketing communication from Shaip.