Developer publishes metadata for 5.6 billion TikTok videos on Hugging Face
Creators

Developer publishes metadata for 5.6 billion TikTok videos on Hugging Face

TechNews Editorial
TechNews EditorialOct 7, 2026 · 1 min read
Share

Why it matters

AI developers want this data to train models on short-form video patterns, while TikTok's official research tools and terms of service strictly limit scraping.

The facts

  • A developer published metadata for 5.6 billion public TikTok videos on Hugging Face.
  • The 460-gigabyte dataset spans from July 2014 through October 2026 and is free for non-commercial use.
  • The collection was gathered using a private mobile API without logins.

A developer known as hashfunction published metadata for approximately 5.6 billion public TikTok videos on the open-source AI repository Hugging Face. The data is available to download for free.

The dataset was published under the datasocial account and covers a timeframe from July 2014 through October 2026. It totals 460 gigabytes of Parquet files split into monthly segments.

The dataset contains metadata and labels

The files contain metadata rather than video footage. Each row includes captions, hashtags, sound IDs, on-screen text, TikTok Shop product and seller IDs, and engagement metrics such as views, likes, comments, shares, saves, and downloads.

Certain fields feature TikTok labels indicating whether a video is flagged as artificial intelligence generated or restricted from the For You page. The artificial intelligence flag is frequently empty for older videos sourced from an archive.

An opened metadata file reveals two video records with engagement fields; one contains an artificial intelligence indicator while the other leaves it empty.
Illustration: AI & Tech News

The data was scraped via mobile API

The dataset remains ungated and open for anyone to download, with Hugging Face recording 1,181 downloads so far. It is also accessible online through datasocial.ai.

DataSocial explained in a write-up that the collection relied on reaching TikTok's private mobile API. The process utilized generated device identities mimicking Android phones, reverse-engineered request signatures, and a spoofed TLS handshake. The system gathered 3.23 billion creator profiles, 5.94 billion videos, and 2.8 billion comments over three weeks without any logins or accounts.

Read nextOpenAI Publishes 722 Manuscripts Containing Math Problem Solutions

Commercial uses require paid tiers

TikTok officially restricts data scraping through Section 3.4 of its United States terms of service, which bans automated extraction without written approval. Official access requires using TikTok Research Tools, which involve an application process limited to qualifying academic researchers and specific non-profit bodies in approved regions.

The Hugging Face dataset carries a CC BY-NC 4.0 license, making it free for research and personal use. Commercial use, creator profiles, and daily updates are directed to datasocial.ai, while the underlying scraper source code is sold separately for $1,699.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading