Set Data Melayu
High-Quality Malay Call Center, General Conversation, and Podcast Dataset for AI & Speech Models
Unscripted telephonic conversations between two people in Malay from Malaysia are available with durations of 15 to 60 minutes, along with licensable public domain audio or video files such as interviews and podcasts with 1 to 5 participants, also ranging from 15 to 60 minutes.
| Dataset Type | General Conversation | General Conversation | Media Data | Call Center |
|---|---|---|---|---|
| Sampling Rate | 8 kHz | 48 kHz | 16 kHz | 8 kHz |
| Speakers | 2 Speakers | 2 Speakers | Multipal Speakers | 2 Speakers |
| Channel | Dual | Mono | Mono | Mono |
| Total Hours | 239:49:43 | 90:19:23 | 343:57:16 | 2,000:00:00 |
| Total # of Speakers | 432 | 140 | 907 | On Request |
Trusted by leading AI teams and enterprises worldwide.
New off-the-shelf datasets are being collected across all data types
Contact us now to let go of your audio/speech training data collection worries
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
Google Tag Manager simplifies the management of marketing tags on your website without code changes.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
Marketing cookies are used to follow visitors to websites. The intention is to show ads that are relevant and engaging to the individual user.
Google Ads is an online advertising platform that enables businesses to create targeted ads displayed on Google search results and partner sites.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.