Structured question-and-answer pairs and instruction-style prompts in Arabic and English, covering Gulf, Levantine and Egyptian dialects alongside Modern Standard Arabic, with domain sets for STEM, food and agriculture.

Arabic Agriculture Q&A Pairs (MSA)

Volume
10,000 Q&A pairs
Scale
~1M words · 2M tokens
Language
Arabic (MSA)
Region
Pan-Arab region
Format
.xlsx / JSON
Licensing
Non-exclusive

Agricultural question-and-answer pairs in Modern Standard Arabic covering crops, irrigation and farming practice relevant to arid and semi-arid regions. Built for agri-advisory assistants and domain adaptation work, where general-purpose corpora carry almost no coverage of regional growing conditions, local crop varieties or water-constrained cultivation.

Arabic STEM Q&A Pairs (MSA)

Volume
10,000 Q&A pairs
Scale
~1M words · 2M tokens
Language
Arabic (MSA)
Region
Pan-Arab region
Format
.xlsx / JSON
Licensing
Non-exclusive

Science, technology, engineering and mathematics question-and-answer pairs in Modern Standard Arabic. Technical terminology and reasoning are consistently the weakest areas in general-purpose Arabic training corpora, so this set targets vocabulary and problem structures that broad web-scraped data does not reach in any depth. Suited to education, research and technical-assistant applications.

Egyptian Arabic Q&A Pairs (MSA and Egyptian Dialect)

Volume
22,000 Q&A pairs
Scale
~2.2M words · 4.4M tokens
Language
Arabic (MSA and Egyptian)
Region
Egypt
Format
.xlsx / JSON
Licensing
Non-exclusive

Egyptian Arabic question-and-answer pairs spanning Modern Standard Arabic and Cairene dialect. Egypt is the largest Arabic-speaking market and its dialect is widely understood across the wider region, which makes this set useful both for country-specific deployments and as a broad-coverage addition to a pan-Arab training mix. Delivered in the same structure as the other dialect sets in this category.

Emirati Arabic Q&A Pairs (MSA and UAE Dialect)

Volume
20,000 Q&A pairs
Scale
~2M words · 4M tokens
Language
Arabic (MSA and UAE)
Region
United Arab Emirates
Format
.xlsx / JSON
Licensing
Non-exclusive

Question-and-answer pairs authored by UAE-based native speakers, blending Modern Standard Arabic with Emirati dialect usage. Suited to fine-tuning assistants for Gulf markets, where training only on MSA tends to produce responses that read as stilted or regionally mismatched to local users. Pairs well with the Saudi set for broader Gulf coverage across a single model.

Levantine Arabic Food and Cuisine Q&A Pairs

Volume
20,000 Q&A pairs
Scale
~2M words · 4M tokens
Language
Arabic (MSA and Levant)
Region
Levant
Format
.xlsx / JSON
Licensing
Non-exclusive

Food, cooking and cuisine question-and-answer pairs covering Levantine dishes, ingredients and preparation methods, written in Modern Standard Arabic and Levantine dialect. Aimed at consumer assistants, recipe and grocery applications, and any model that needs culturally grounded food knowledge for the region rather than translated Western equivalents.

Pan-Arab General Q&A Pairs (English)

Volume
30,000 Q&A pairs
Scale
~3M words · 6M tokens
Language
English
Region
Pan-Arab region
Format
.xlsx / JSON
Licensing
Non-exclusive

English-language question-and-answer pairs authored with Arab regional context and subject matter. Pairs directly with the Modern Standard Arabic general set for bilingual or cross-lingual instruction-tuning, and is useful for evaluating how reliably a model carries regional knowledge from one language into the other rather than losing it at the boundary.

Pan-Arab General Q&A Pairs (MSA)

Volume
30,000 Q&A pairs
Scale
~3M words · 6M tokens
Language
Arabic (MSA)
Region
Pan-Arab region
Format
.xlsx / JSON
Licensing
Non-exclusive

General-domain question-and-answer pairs in Modern Standard Arabic, written for use across the whole Arabic-speaking region rather than any single market. This is the broadest starting point in the category for teams beginning Arabic instruction-tuning, and the natural base layer to combine with the dialect-specific and domain-specific sets listed alongside it.

Saudi Arabic Q&A Pairs (MSA and Saudi Dialect)

Volume
30,000 Q&A pairs
Scale
~3M words · 6M tokens
Language
Arabic (MSA and Saudi)
Region
Saudi Arabia
Format
.xlsx / JSON
Licensing
Non-exclusive

Human-authored question-and-answer pairs written in Modern Standard Arabic alongside Saudi dialect, covering everyday, cultural and knowledge-seeking exchanges. Built for instruction-tuning and evaluating Arabic language models where Gulf dialect coverage matters. Each pair is reviewed before delivery and the set ships with a consistent field structure throughout, so it drops straight into an existing fine-tuning pipeline without reformatting.