Text datasets for AI training
Conversational Q&A, code, news, clinical and raw text corpora — ready to license and scoped to your project.
Licensed Text Datasets for LLM and NLP Training
Shaip’s text data catalog gives AI teams ready-to-license text datasets for training, fine-tuning and evaluating large language models (LLMs) and NLP systems. It covers eight categories. Each dataset is sourced from real-world material and can be scoped to your project.
The catalog includes conversational Q&A and instruction datasets in English and Arabic, covering Modern Standard Arabic and Saudi, Emirati, Egyptian and Levantine dialects. It also has code repository datasets in nine programming languages, several with full pull-request and commit history. For current-events models and RAG, there are news datasets in English, Spanish and Arabic, including a 975,000-article archive dating back to 1999. Healthcare AI teams can use clinical Q&A and case datasets, available as golden and silver sets, with de-identification included.
For pretraining, there are raw text and web-scale corpora, including a 20-million-word Arabic corpus. Scanned document corpora support document AI and information extraction. Academic and publisher text covers twelve science and engineering subjects, and scripted entertainment text covers film, TV, animation and microdramas.
Every dataset comes with clear licensing and quality review. Samples are available so you can evaluate a dataset before you buy. Browse a category below or contact us to discuss custom text data collection.
Conversational Q&A & Instruction Datasets
These datasets contain human-authored question-and-answer pairs and instruction-style prompts in Arabic and English, spanning Modern Standard Arabic alongside Saudi, Emirati, Egyptian and Levantine dialects. Domain-specific sets cover STEM, food and agriculture. They support instruction-tuning, fine-tuning and evaluation of language models where regional dialect coverage and culturally grounded knowledge matter.
Code & Software Repository Datasets
Production codebases drawn from real, shipped software across C#, React, TypeScript, Angular, Java, Go, PHP, Swift and BrightScript. Several entries include full pull-request and commit history, showing how code changed under review. They support code-completion, code-generation, code-understanding and code-review model training, with repositories available for evaluation prior to licensing.
News & Journalistic Text Datasets
Professionally authored and edited news articles from a major international news organisation, available as a recurring feed in English, Spanish and Arabic, and as an archival corpus of 975,000 articles published since 1999. Every article carries byline, date, headline, body and topic tags. They suit current-events model training, retrieval-augmented generation and temporal reasoning tasks.
Clinical Q&A & Case Datasets
Medical question-answer sets, clinical case summaries and full-text medical articles for healthcare-focused model training and evaluation. Golden and silver datasets are available, the former prepared to a higher review standard for benchmarking. Where source material contains personally identifiable or protected health information, redaction, de-identification and masking are scoped into the engagement.
Raw Text & Web-Scale Corpora
Large-volume raw text, prompt and text-to-speech script corpora spanning Modern Standard Arabic and regional dialects, including a 20-million-word raw corpus. They support language model pretraining and the expansion of vocabulary and dialect coverage in models that handle MSA well but perform poorly on spoken regional variants.
Document & Scanned Text Corpora
Real-world scanned documents and structured forms delivered as PDFs, in authentic layouts and scan quality rather than synthetic renders. They support document AI, OCR-adjacent natural language processing and information extraction, particularly for key-value and tabular document understanding where layout carries as much meaning as the text itself.
Academic & Publisher Text Datasets
Full-text books and journal articles licensed from an academic publisher and delivered via API, covering twelve subject areas across physical sciences and engineering — astronomy, biomedical, civil, electrical and electronics, energy, industrial, materials science, mechanical, nanotechnology, physics, polymer science and security management. They suit scientific and technical pretraining and retrieval over peer-reviewed source material.
Scripted Entertainment & Narrative Text
Production scripts, transcripts and narrative text drawn from a large scripted-entertainment catalogue spanning feature films, television, animation and vertical microdramas, across drama, comedy, science fiction, western, action, mystery, romance, thriller and faith genres. They support narrative and dialogue generation, screenplay-format modelling and script-to-scene alignment.
Frequently Asked Questions (FAQ)
1. What are text datasets for AI training?
Text datasets for AI training are curated collections of written language, such as question-and-answer pairs, instructions, articles, code, clinical records, books and scripts, used to pretrain, fine-tune and evaluate large language models (LLMs) and NLP systems. Shaip’s text data catalog offers ready-to-license datasets across eight categories, each sourced from real-world material.
2. What types of text datasets does Shaip offer?
The catalog covers eight categories: conversational Q&A and instruction datasets, code and software repository datasets, news and journalistic text, clinical Q&A and case datasets, raw text and web-scale corpora, scanned document corpora, academic and publisher text, and scripted entertainment and narrative text.
3. Which languages and dialects are covered?
Datasets are available in English, Arabic and Spanish. Arabic coverage includes Modern Standard Arabic along with Saudi, Emirati, Egyptian and Levantine dialects, which helps models that handle MSA well but struggle with spoken regional variants.
4. Which AI use cases do these datasets support?
Common use cases include LLM pretraining, instruction-tuning and fine-tuning, model evaluation and benchmarking, retrieval-augmented generation (RAG), code completion and code review, healthcare AI, document AI and information extraction, and dialogue and screenplay generation.
5. Can I evaluate a dataset before licensing it?
Yes. Samples or evaluation access can be arranged before you license a dataset, so your team can check fit, format and quality. Contact us with your use case to request a sample from the relevant category.
6. How is sensitive data such as PII and PHI handled?
Where source material contains personally identifiable or protected health information, as with clinical data, redaction, de-identification and masking are scoped into the engagement so the data meets your privacy and compliance requirements.
7. What is the difference between golden and silver datasets?
Golden datasets are prepared to a higher review standard and are best suited to benchmarking and evaluation. Silver datasets go through standard review and are well suited to larger-scale training where volume matters more.
8. Can the datasets be customized or extended?
Yes. Datasets can be scoped to your project by language, domain, volume or format. If the catalog does not cover what you need, Shaip can collect and annotate custom text data to your specifications.
9. What licensing options are available?
Licensing is scoped to each engagement based on the dataset, volume, intended use and duration. Some datasets, such as news content, are available both as a recurring feed and as a one-time archival corpus.
10. How are the datasets delivered?
Delivery depends on the dataset. Options include structured files, PDFs for scanned documents, recurring feeds for news, and API delivery for academic and publisher content. Delivery format can be aligned with your training pipeline.
11. How much do text datasets cost?
Pricing depends on the category, volume, exclusivity and licensing terms. Share your requirements with our team for a quote tailored to your project.
12. How do I get started?
Browse the category that matches your use case, then contact Shaip with your requirements. Our team will share dataset details, arrange samples and scope a license for your project.