Text datasets for AI training

Conversational Q&A, code, news, clinical and raw text corpora — ready to license and scoped to your project.

Text datasets

Licensed Text Datasets for LLM and NLP Training

Shaip’s text data catalog gives AI teams ready-to-license text datasets for training, fine-tuning and evaluating large language models (LLMs) and NLP systems. It covers eight categories. Each dataset is sourced from real-world material and can be scoped to your project.

The catalog includes conversational Q&A and instruction datasets in English and Arabic, covering Modern Standard Arabic and Saudi, Emirati, Egyptian and Levantine dialects. It also has code repository datasets in nine programming languages, several with full pull-request and commit history. For current-events models and RAG, there are news datasets in English, Spanish and Arabic, including a 975,000-article archive dating back to 1999. Healthcare AI teams can use clinical Q&A and case datasets, available as golden and silver sets, with de-identification included.

For pretraining, there are raw text and web-scale corpora, including a 20-million-word Arabic corpus. Scanned document corpora support document AI and information extraction. Academic and publisher text covers twelve science and engineering subjects, and scripted entertainment text covers film, TV, animation and microdramas.

Every dataset comes with clear licensing and quality review. Samples are available so you can evaluate a dataset before you buy. Browse a category below or contact us to discuss custom text data collection.

Conversational Q&A & Instruction Datasets

These datasets contain human-authored question-and-answer pairs and instruction-style prompts in Arabic and English, spanning Modern Standard Arabic alongside Saudi, Emirati, Egyptian and Levantine dialects. Domain-specific sets cover STEM, food and agriculture. They support instruction-tuning, fine-tuning and evaluation of language models where regional dialect coverage and culturally grounded knowledge matter.

Conversational q&a & instruction datasets

Code & Software Repository Datasets

Production codebases drawn from real, shipped software across C#, React, TypeScript, Angular, Java, Go, PHP, Swift and BrightScript. Several entries include full pull-request and commit history, showing how code changed under review. They support code-completion, code-generation, code-understanding and code-review model training, with repositories available for evaluation prior to licensing.

Code & software repository datasets

News & Journalistic Text Datasets

Professionally authored and edited news articles from a major international news organisation, available as a recurring feed in English, Spanish and Arabic, and as an archival corpus of 975,000 articles published since 1999. Every article carries byline, date, headline, body and topic tags. They suit current-events model training, retrieval-augmented generation and temporal reasoning tasks.

News & journalistic text datasets

Clinical Q&A & Case Datasets

Medical question-answer sets, clinical case summaries and full-text medical articles for healthcare-focused model training and evaluation. Golden and silver datasets are available, the former prepared to a higher review standard for benchmarking. Where source material contains personally identifiable or protected health information, redaction, de-identification and masking are scoped into the engagement.

Clinical q&a & case datasets

Raw Text & Web-Scale Corpora

Large-volume raw text, prompt and text-to-speech script corpora spanning Modern Standard Arabic and regional dialects, including a 20-million-word raw corpus. They support language model pretraining and the expansion of vocabulary and dialect coverage in models that handle MSA well but perform poorly on spoken regional variants.

Raw text & web-scale corpora

Document & Scanned Text Corpora

Real-world scanned documents and structured forms delivered as PDFs, in authentic layouts and scan quality rather than synthetic renders. They support document AI, OCR-adjacent natural language processing and information extraction, particularly for key-value and tabular document understanding where layout carries as much meaning as the text itself.

Document & scanned text corpora

Academic & Publisher Text Datasets

Full-text books and journal articles licensed from an academic publisher and delivered via API, covering twelve subject areas across physical sciences and engineering — astronomy, biomedical, civil, electrical and electronics, energy, industrial, materials science, mechanical, nanotechnology, physics, polymer science and security management. They suit scientific and technical pretraining and retrieval over peer-reviewed source material.

Academic & publisher text datasets

Scripted Entertainment & Narrative Text

Production scripts, transcripts and narrative text drawn from a large scripted-entertainment catalogue spanning feature films, television, animation and vertical microdramas, across drama, comedy, science fiction, western, action, mystery, romance, thriller and faith genres. They support narrative and dialogue generation, screenplay-format modelling and script-to-scene alignment.

Scripted entertainment & narrative text

Text datasets for AI training are curated collections of written language, such as question-and-answer pairs, instructions, articles, code, clinical records, books and scripts, used to pretrain, fine-tune and evaluate large language models (LLMs) and NLP systems. Shaip’s text data catalog offers ready-to-license datasets across eight categories, each sourced from real-world material.

The catalog covers eight categories: conversational Q&A and instruction datasets, code and software repository datasets, news and journalistic text, clinical Q&A and case datasets, raw text and web-scale corpora, scanned document corpora, academic and publisher text, and scripted entertainment and narrative text.

Datasets are available in English, Arabic and Spanish. Arabic coverage includes Modern Standard Arabic along with Saudi, Emirati, Egyptian and Levantine dialects, which helps models that handle MSA well but struggle with spoken regional variants.

Common use cases include LLM pretraining, instruction-tuning and fine-tuning, model evaluation and benchmarking, retrieval-augmented generation (RAG), code completion and code review, healthcare AI, document AI and information extraction, and dialogue and screenplay generation.

Yes. Samples or evaluation access can be arranged before you license a dataset, so your team can check fit, format and quality. Contact us with your use case to request a sample from the relevant category.

Where source material contains personally identifiable or protected health information, as with clinical data, redaction, de-identification and masking are scoped into the engagement so the data meets your privacy and compliance requirements.

Golden datasets are prepared to a higher review standard and are best suited to benchmarking and evaluation. Silver datasets go through standard review and are well suited to larger-scale training where volume matters more.

Yes. Datasets can be scoped to your project by language, domain, volume or format. If the catalog does not cover what you need, Shaip can collect and annotate custom text data to your specifications.

Licensing is scoped to each engagement based on the dataset, volume, intended use and duration. Some datasets, such as news content, are available both as a recurring feed and as a one-time archival corpus.

Delivery depends on the dataset. Options include structured files, PDFs for scanned documents, recurring feeds for news, and API delivery for academic and publisher content. Delivery format can be aligned with your training pipeline.

Pricing depends on the category, volume, exclusivity and licensing terms. Share your requirements with our team for a quote tailored to your project.

Browse the category that matches your use case, then contact Shaip with your requirements. Our team will share dataset details, arrange samples and scope a license for your project.