Speech Dataset

Wake Word French Dataset

High-Quality French Wake Word Dataset for AI & Speech Models

Wake Word French Dataset

Overview

  • LanguageFrench
  • Dataset TypesWake Word / Keyphrase
  • CountryFrench

Description

Wake Words / Voice Command / Trigger Word / Keyphrase collection of data 200 speakers3 unique keyphrases per speaker25-30 repeated keyphrases recordings per unique keyphrase25-30 audio files per unique keyphrase90 total recorded utterances per speaker

Dataset Highlights

  • LanguageFrench
  • Sample Rate16 kHz
  • Total Hours200 speakers
  • Speakers200+
  • Channels1 channel

Use Cases

  • Wake Word Detection
  • Voice Command & Trigger
  • Keyword Spotting
  • Hands-free Activation
  • On-device Voice UI

Sample Dataset Specs

Dataset TypeWake Word
Sampling Rate16 kHz
Speakers200 Speakers
Channel1 channel
Total Hours200 speakers
Total # of Speakers200

Key Benefits

  • Diverse & RepresentativeBroad speaker coverage for inclusive, real-world AI.
  • High-Quality AudioClean, balanced recordings for reliable model performance.
  • Ready for AI/MLStructured and annotated for seamless integration.
  • Scalable & FlexibleMultiple durations, speakers, and domains to fit your needs.

Technical Details

  • Recording PlatformMobile App
  • Audio Format.wav
  • Transcription Format.json
  • WER (%)5
  • Gender DistributionFemale 50%, Male 50%, Unknown 10%
  • Age Range18-50

Featured Clients

Trusted by leading AI teams and enterprises worldwide.

  • Microsoft
  • Amazon Web Services
  • Google

Can’t find what you are looking for?

New off-the-shelf datasets are being collected across all data types

Contact us now to let go of your audio/speech training data collection worries

  • This field is for validation purposes and should be left unchanged.
  • By registering, I agree with Shaip Privacy Policy and Terms of Service and provide my consent to receive B2B marketing communication from Shaip.