Speech Dataset

Indonesian Dataset

3 licensable sets · 3 dataset types · 1 country

3 licensable Indonesian speech sets: 2,300 hours across 3 dataset types, sampled at 8–16 kHz in mixed mono/stereo and mono channel.

Indonesian Dataset

Overview

  • LanguageIndonesian
  • Dataset TypesCall-Center, Scripted Monologue, Utterance
  • CountryIndonesia

Dataset Highlights

  • LanguageIndonesian
  • Sample Rate16 kHz / 8 kHz
  • Total Hours2,300 hrs
  • ChannelsMono & Stereo (mixed) / Mono
Available sets

What you can license today

Volumes are as reported by the sourcing partner; confirmed at scoping. Pricing is quoted per audio hour.

Language & Accent · CountryDataset TypeVolumeAudio ChannelAudio Frequency
IndonesianIndonesiaCall Centre (real)200 hoursMono & Stereo (mixed)8 kHz
IndonesianIndonesiaScripted Monologue1,000 hoursMono16 kHz+
IndonesianIndonesiaUtterance1,100 hoursMono16 kHz+

Use Cases

  • ASR (Automatic Speech Recognition)
  • Conversational AI
  • Call-Centre & Speech Analytics
  • Voice Assistants & Command Recognition
  • Language Modelling

Key Benefits

  • Diverse & RepresentativeBroad speaker coverage for inclusive, real-world AI.
  • High-Quality AudioClean, balanced recordings for reliable model performance.
  • Ready for AI/MLStructured and annotated for seamless integration.
  • Scalable & FlexibleMultiple durations, speakers, and domains to fit your needs.

Featured Clients

Trusted by leading AI teams and enterprises worldwide.

  • Microsoft
  • Amazon Web Services
  • Google

Can't find the Indonesian set you need?

Tell us the dataset type, volume and recording conditions — we scope custom collection in this language too.

  • This field is for validation purposes and should be left unchanged.
  • By registering, I agree with Shaip Privacy Policy and Terms of Service and provide my consent to receive B2B marketing communication from Shaip.