High-quality Training Data to build Multi-lingual Conversational AI for Indian Languages
Offered high-quality audio data collection, transcription, and annotation services, to train their AI-powered Speech Processing multilingual Voice Suite.
Project Overview
The client builds technologies that are bridging the language divide in the digital world. Content from apps and portals can be delivered in multiple languages in real-time through their Language-as-a-Service platform. The Language Platform delivers static and dynamic content from apps and portals in multiple languages in real-time.
They deliver language localization across several customer such as mobile brands, chipset manufacturers, consumer Internet, banks and financial institutions, education governments, entertainment, and more.
Key Stats
Audio Data Collected, Transcribed and Annotated
300 Hrs
No. of
Languages
3 (100 hrs x 3 languages)
Languages Transcribed & Annotated
Indian English, Hindi, Telugu
Project
Timeline
12 Weeks
Challenges
Lack of access to quality linguists, expert data annotators, and stringent timelines has been a bottleneck in the progress and adoption of conversational AI. Sourcing linguists, transcribing and annotating datasets in large volumes with the required quality, sufficient enough to build AI capabilities, is time-consuming, expensive task that requires skilled resources from various domains.
The critical requirements of the client were:
- Access to Quality Linguists for Audio Collection: Collect 300 hours of audio for languages such as Indian English, Hindi & Telugu from expert linguists with demographic diversity.
- Complex Audio Transcription & Annotation with minimum. accuracy of 90%:
- Transcribed the entire audio files, capturing all words as spoken – including hesitations, filler words, false starts, and other verbal tics
- Provided error-free output despite stringent transcription guidelines associated with dialectal pronunciations, mispronounced words, non-standard usage, and punctuation.
- Data Delivery: Delivery of high-quality and diverse audio files along with corresponding transcripts (JSON format) for 3 languages within 12 weeks timeline.
Solution
With our deep understanding of conversational AI, we helped the client collect, transcribe and annotate the data by a team of expert linguists and annotators to train their AI-powered multilingual Voice Suite.
After evaluating many vendors, The client chose Shaip because of our expertise to complete conversational AI projects within stringent timelines, cost-effectively, and with the required quality. Their team was also impressed with Shaip’s robust and comprehensive project execution capability.
| Language | No. of Hrs | Unscripted General Conversation (50%) | Unscripted Call Center Conversation (30%) | Online Scrapped Media Content (20%) |
|---|---|---|---|---|
| Hindi | 100 | 50 | 30 | 20 |
| English | 100 | 50 | 30 | 20 |
| Telegu | 100 | 50 | 30 | 20 |
| Total | 300 | 150 | 90 | 60 |
- Additional Audio Collection Requirement
- Frequency: Sampling rate – 16 kHz
- Format: WAV
- Duration – General duration (5-30 mins) – conversations
- Speech Domain – BFSI, Telecom, and Retail
- Demographic Diversity – Gender: 1:1 Ratio, Age Range: 18-60, Capture Speaker ID to identify each unique Speaker
- Audio Segmentation, Transcription & Annotation Guidelines followed
2.1 Segmentation- Create segments of 15 seconds each (timestamped to milliseconds) for individual files larger than 30 seconds
- Segment Sound Types – speech, babble, music, noise, overlap
- Segment Labeling – start time, end time, segment ID, loudness level, primary sound type,
language code, speaker ID
2.2 Transcription & Annotation- Transcribed the entire audio files, capturing all words as spoken – including hesitations, filler words, false starts, and other verbal tics i.e. symbols, characters
- Provided error-free output despite stringent transcription guidelines associated with dialectal pronunciations, mispronounced words, non-standard usage, punctuations, capitalization, abbreviations, contractions, numbers, acronyms, disfluent & unintelligible speech, non-target
languages, and more.
- Data Delivery
All WAV audio and transcript files were delivered in JSON format in accordance with the unique speaker demographic included within 12 weeks.
The Outcome
The high-quality annotated audio data from expert linguists empowered the client to train their AI-powered Speech Processing multilingual Voice Suite accurately in 3 languages i.e. Indian English, Hindi, & Telugu in the stipulated time.
With the gold-standard training datasets, the client was able to offer intelligent and robust Speech-to-text, language translation, & multilingual keypads services to solve real-world problems.
We were impressed with Shaip’s robust project execution capability, their expertise to source, transcribe and annotate the required audio data from expert linguists. Multiple quality checks, adherence to strict timelines, and cost value leadership that they offer is second to none.