Computer Vision Datasets

Shaip Expands Its Computer Vision Data Catalog with New Image and Video Datasets

Most computer vision projects don’t stall at the model. They stall at the data. Teams have the architecture, the compute and the engineers, and then spend months trying to source images and video that actually match the conditions their model will meet in production. That sourcing problem is what our computer vision data catalog exists to solve, and we have just made it considerably bigger. New image and video datasets have been added across every category in the catalog, from face biometrics and anti-spoofing through to OCR, fashion, industrial inspection and long-form video.

Why teams license visual data instead of collecting it

Building a dataset in-house is rarely a two-week job. Recruiting participants across the right demographics, running the capture, handling consent, then annotating to a usable standard can take a full quarter before a single training run happens. Licensing an existing dataset compresses that into a conversation and a contract. It also makes it cheaper to test a hypothesis. If a model needs a particular kind of data to work, it is far better to find that out with licensed data in week one than after six months of your own collection.

What’s new, category by category

The update touches all nine dataset categories in the catalog. Here is what stands out in each.

Anti-spoofing

Anti-spoofing

Three video datasets now cover the attack types liveness systems actually face: 3D masks and makeup-based disguises across 4,836 videos, mask and replay attacks from 50 participants, and a larger real-versus-replay collection of 11,466 videos. All are supplied unannotated, which suits pretraining and benchmark evaluation, with expert labelling available on request.

Know more
Facial and body part recognition

Facial and body part recognition

This is the deepest part of the refresh. New additions include a 10-million-image face recognition dataset, a 44,000-image Asian facial collection, a 70,000-image selfie-and-official-ID dataset covering more than 6,000 identities, and a 24,000-image smartphone palm recognition set built from 2,000 participants worldwide. Two historical datasets capture genuine ageing across six months to more than ten years using real photographs rather than synthetic ageing, which matters for age-invariant recognition and KYC. A kids age progression dataset and a synthetic children's faces dataset address a well-documented gap: children are largely absent from age detection and deepfake prevention data, and the synthetic set widens demographic coverage without using any real-person data.

Know more
Environment and scene segmentation

Environment and scene segmentation

Autonomous driving and smart-city teams get the volume here: drivable area segmentation at 115,300 images, lane line segmentation at 135,300 images across 35 road-marking categories, and walkway segmentation at 87,800 images. Alongside those sit CCTV traffic scenes, panoptic road and building scenes, indoor object segmentation and a 10,000-clip home security camera collection for anomaly and activity detection.

Know more
Language and text (ocr)

Language and text (OCR)

The newest additions lean towards Arabic and Chinese. They include 25,000 Arabic calligraphy images gathered from public archives and working calligraphers, 15,000 handwritten and scanned Arabic documents, 10,000 street and storefront signs containing Arabic text, and a Chinese publication library spanning books, journals and public policy material. Useful for scene text detection, handwriting recognition and document digitisation in scripts that most open datasets underserve.

Know more
Clothing and fashion

Clothing and fashion

Retail and e-commerce models can draw on a two-million-image clothing classification set, a one-million-image keypoint set covering 80 clothing types, 500,000 segmentation images, fabric classification across 11 material categories, and an e-commerce product dataset spanning 16 categories and over 200,000 SKUs.

Know more
Gesture, pose and activity

Gesture, pose and activity

New entries include 10,000 smart-home activity images, 10,000 home-activity videos covering doorstep and indoor behaviour under mixed lighting, 21,000 gesture images, a 21-point hand skeleton keypoint set, and a posture classification dataset covering 14 distinct pose types.

Know more
Machine and industry

Machine and industry

Insurance and manufacturing use cases are well served: 48,000 360-degree damaged car walkaround videos, an annotated damaged car image set, licence plate collections for the US, Canada and Mexico, driver behaviour images captured from vehicle interiors, 120,000 machine part defect images, and 41,000 industrial smelting flame images classified across 10 conditions.

Know more
Object and contour segmentation

Object and contour segmentation

Two additions worth flagging: a 15,000-image food, recipe and cooking dataset created by chefs and aimed at multimodal and LLM training, and a grocery product image collection suited to shelf monitoring, SKU recognition and automated checkout.sified across 10 conditions.

Know more
Video at scale segmentation

Video at scale

For teams training video and multimodal models, the 'other datasets' category now carries serious volume: 80,000 hours of kids' content across 240,000 YouTube videos, 500 hours of short films, documentaries and wedding footage, 500 hours of historical documentaries featuring more than 20,000 unique faces, 3,000 hours from a documentary filmmaker collection spanning eight countries, and 1,000 hours of martial arts fights.

Know more

How these Computer Vision datasets are built

Every dataset in the catalog is ethically sourced with contributor consent and handled in line with GDPR and other global privacy standards, with de-identification available where a use case calls for it. Datasets are delivered in standard formats such as JSON, CSV or XML with metadata attached, so they drop into existing training pipelines without a conversion step. Where a dataset ships unannotated, our annotation teams can label it to your specification. And where the catalog doesn’t hold exactly what you need, we can collect it. 

Where to start

Browse the catalog by category and request a sample of anything that looks relevant. If your requirement is specific, whether that’s a demographic split, a lighting condition or a particular object class, tell us and we will either point you to the closest fit or scope a custom collection. New datasets are being added continuously, so it is worth checking back.

Enjoyed this article? Follow Shaip on LinkedIn for more updates.

Social Share