High-Quality Polish Media Data and Scripted Monologue for AI & Speech Models
This dataset includes licensable public domain audio or video files such as interviews and podcasts with 1 to 5 participants (15–60 minutes), along with scripted monologues where a single speaker delivers predefined content for training and evaluating speech and language models.
| Dataset Type | Media Data | Scripted Monologue |
|---|---|---|
| Sampling Rate | 16 kHz | 48 kHz |
| Speakers | Multipal Speakers | Single Speaker |
| Channel | Mono | Mono |
| Total Hours | 268:56:51 | 2,348:00:00 |
| Total # of Speakers | 532 | 2,699 |
Trusted by leading AI teams and enterprises worldwide.
New off-the-shelf datasets are being collected across all data types
Contact us now to let go of your audio/speech training data collection worries
We use cookies to improve your experience on our site. By using our site, you consent to cookies.
Manage your cookie preferences below:
Essential cookies enable basic functions and are necessary for the proper function of the website.
Google Tag Manager simplifies the management of marketing tags on your website without code changes.
Statistics cookies collect information anonymously. This information helps us understand how visitors use our website.
Google Analytics is a powerful tool that tracks and analyzes website traffic for informed marketing decisions.
Service URL: policies.google.com (opens in a new window)
Marketing cookies are used to follow visitors to websites. The intention is to show ads that are relevant and engaging to the individual user.
Google Ads is an online advertising platform that enables businesses to create targeted ads displayed on Google search results and partner sites.
Service URL: policies.google.com (opens in a new window)
You can find more information in our Cookie Policy and Privacy Policy.