Large-volume raw text, prompt and script corpora across Modern Standard Arabic and regional dialects, for language model pretraining and for widening vocabulary and dialect coverage.

Arabic Prompt Corpus

Volume
65,000 prompts
Scale
~1M words · 2M tokens
Language
Arabic (MSA and dialects)
Region
Pan-Arab region
Format
.xlsx / JSON
Licensing
Non-exclusive

Sixty-five thousand Arabic prompts across Modern Standard Arabic and regional dialects. Distinct from the Q&A sets in that it supplies the instruction side without paired responses, which suits teams generating their own completions, running preference-tuning work, or building evaluation harnesses that need broad prompt coverage in Arabic.

Arabic Raw Text Corpus (MSA and Dialects)

Volume
20M words
Scale
~40M tokens
Language
Arabic (MSA and dialects)
Region
Pan-Arab region
Format
TXT / JSON
Licensing
Non-exclusive

Twenty million words of raw Arabic text spanning Modern Standard Arabic and regional dialects, the largest single Arabic corpus in this catalogue. Intended for language model pretraining and for expanding vocabulary and dialect coverage in models that currently handle MSA well and dialect poorly. Delivered as plain text or JSON depending on your pipeline.

Saudi Arabic TTS Script Prompts

Volume
11,000 prompts
Scale
~223K words
Language
Arabic (MSA and Saudi)
Region
Saudi Arabia
Format
.xlsx / JSON
Licensing
Non-exclusive

Eleven thousand script prompts written for text-to-speech recording in Modern Standard Arabic and Saudi dialect. Phrasing is built for being spoken aloud rather than read, which matters for voice work and is the reason general text corpora make poor TTS scripts. Suited to speech synthesis pipelines and to voice-assistant response modelling.