YouTube Datasets for AI Teams
Structured public YouTube data for training, evaluation and analysis. We collect it through our own residential proxy network and deliver it in JSON, CSV or Parquet — as a packaged dataset, a custom crawl or an API feed.
What Every Record Contains
Eight core fields per video, consistent across every format and delivery mode.
The full video title as published — ready for text pipelines and tokenization.
Channel name and channel-level metadata for grouping, deduplication and source tracking.
View count at collection time — a clean popularity signal for ranking and sampling.
Like count for engagement scoring, weak labels and quality filtering.
Video length in seconds — separate shorts from long-form content in one pass.
Publish timestamp in ISO format for time-window slicing and trend analysis.
Creator tags plus the video language — filter by topic and locale without extra preprocessing.
Comment count per video, with full comment text available as an optional add-on.
{
"video_id": "dQw4w9WgXcQ",
"title": "Never Gonna Give You Up",
"channel": "Rick Astley",
"views": 1620000000,
"likes": 18400000,
"duration_sec": 213,
"published_at": "2009-10-25T06:57:33Z",
"language": "en",
"tags": ["music", "80s", "pop"],
"comments_count": 2300000
}Three Ways to Get the Data
From a one-off download to a live feed into your pipeline — pick the delivery that fits your workflow.
Ready-Made Dataset
We slice the data to your spec, package it and hand over a one-time download. The fastest way to get training data into your hands.
Custom Collection
Define keywords, channels, regions and time windows. We run the crawl to your spec and can keep it refreshed on an agreed schedule.
API Access
Pull fresh data on demand through an API — built for teams wiring collection straight into their own data pipeline.
What Teams Build With It
Titles, tags and comments across 195+ countries make diverse, real-world corpora for pretraining and fine-tuning LLM and multimodal models.
Engagement fields like views and likes give you real-world signals for benchmarking ranking and retrieval systems.
Track what is being published and watched, sliced by region, language and time window — down to a single niche.
Train and audit moderation models on real public video metadata and comment data instead of synthetic samples.
Compliance Built In
Public data only, handled to enterprise standards from collection to delivery.
Frequently Asked Questions
Ready to Talk Data?
Tell us your scope on Telegram — keywords, regions, time window — and we will scope the collection and send you a quote.
