Someone Scraped 5.6 Billion TikTok Videos and Put the Data on Hugging Face for Free

A developer using the handle hashfunction published metadata for roughly 5.6 billion public TikTok videos on Hugging Face under the datasocial account, covering July 2014 through October 2026. The 460 GB Parquet dataset contains captions, hashtags, engagement counts and other metadata and is offered under a CC BY-NC 4.0 license while directing commercial users to datasocial.ai.

By AI Newsroom· Reviewed by Pranav, Founder & Editor-in-ChiefPublished 2 minutes agoUpdated 2 minutes ago0 views
Someone Scraped 5.6 Billion TikTok Videos and Put the Data on Hugging Face for Free

Why It Matters

Large-scale metadata like this is valuable training material for AI models that analyze short-form video language, engagement and commerce signals, and its availability outside TikTok's official research channels raises legal and policy questions given the platform's prohibition on automated scraping without written permission.

Key Facts

  • Publisher: Hashfunction (datasocial account) on Hugging Face
  • Dataset size: Approximately 5.6 billion videos; 460 GB of Parquet files
  • Date range: July 2014 through October 2026
  • License: CC BY-NC 4.0 (free for research and personal use)
  • Commercial access / updates: Routed to datasocial.ai (commercial use, creator profiles, daily updates)

A developer known as hashfunction uploaded metadata for roughly 5.6 billion public TikTok videos to Hugging Face, publishing the files under the datasocial account. The dataset spans July 2014 through October 2026 and is stored as monthly Parquet files totalling about 460 GB. Hugging Face shows the upload has been downloaded 1,181 times so far, and the data is also accessible through datasocial.ai.

The collection is metadata only, not video footage. Each record contains fields such as captions, hashtags, sound IDs, on-screen text, TikTok Shop product and seller IDs, and engagement metrics including views, likes, comments, shares, saves and downloads. Some TikTok-derived labels appear as well, for example whether a clip was flagged as AI-generated or excluded from the For You page; older entries often have empty AI flags, consistent with archive-derived records.

Datasocial's dataset card and write-up say the data was obtained from TikTok's private mobile API without logging in, using generated device identities that impersonate Android phones, reverse-engineered request signatures and a spoofed TLS handshake. The write-up claims the system gathered 3.23 billion creator profiles, 5.94 billion videos and 2.8 billion comments in a three-week run. TikTok's U.S. terms of service, specifically section 3.4, bar automated extraction of data from the platform without written approval.

The release is presented under a non-commercial Creative Commons license and the public Hugging Face card routes commercial inquiries, creator profile access and daily update requests to datasocial.ai. The scraper's source code is offered separately for sale at $1,699. The source also notes that other large TikTok metadata sets, including a 4.5-billion-video collection purportedly obtained the same way, are available on Hugging Face.

Keep Reading