LAION releases Big Video Dataset with 10 million hours of open footage
LAION has released the Big Video Dataset, described as one of the largest open video datasets for AI research, spanning roughly 80 million videos, 10 million hours of runtime, and 55 million auto-described clips. According to The Decoder, models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks, offering researchers a new scale baseline for open video-text work. The announcement arrives as LAION can likely point to a 2024 Hamburg court ruling in its legal positioning, though available excerpts stop before the ruling's full implications are spelled out. Builders should treat the benchmark gap as an early signal while awaiting fuller documentation on clip generation, licensing boundaries, and how the legal argument applies to downstream training.
LAION releases Big Video Dataset with 10 million hours of open footage
LAION has released the Big Video Dataset (BVD), one of the largest open video datasets for AI research. models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks.
Key takeaway
Open video research gains a 10-million-hour LAION corpus with reported gains up to 2.1 points over InternVid on video-to-text tests.
What happened
LAION has released the Big Video Dataset (BVD), which The Decoder describes as one of the largest open video datasets for AI research, containing 80 million videos, 10 million hours of runtime, and 55 million auto-described clips.
Reporting states that models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks, and notes LAION can likely point to a 2024 Hamburg court ruling in its legal framing.
Evidence
LAION released the Big Video Dataset with 80 million videos and 10 million hours of footage.
The Decoder · attributed
LAION's Big Video Dataset (BVD) is one of the largest open video datasets for AI research, with 80 million videos, 10 million hours of runtime, and 55 million auto-described clips.
BVD-trained models beat InternVid-trained models by up to 2.1 percentage points on video-to-text benchmarks.
The Decoder · attributed
models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks.
LAION may cite a 2024 Hamburg court ruling in its legal positioning for the dataset.
The Decoder · attributed
Legally, LAION can likely point to a 2024 Hamburg court ruling that allows
Why it matters
Large open video corpora can reset training economics and competitive benchmarks for multimodal models outside closed provider pipelines.
Limits and uncertainties
Available reporting on the 2024 Hamburg court ruling is truncated before its full legal scope is described.
The packet does not detail quality controls, licensing terms, or provenance checks for the 55 million auto-described clips.
Practical implications
Benchmark teams may need to compare video-to-text models against BVD-trained baselines rather than InternVid-only references.
Compliance reviewers should hold production adoption until LAION publishes complete legal and licensing documentation beyond the cited court ruling excerpt.
What to watch
Full LAION disclosure on how the 2024 Hamburg court ruling applies to BVD distribution and downstream training.
Independent replication of the reported up to 2.1 percentage point advantage over InternVid on common video-to-text benchmarks.