TikTok Video Metadata Released on Hugging Face
A developer has published a massive dataset of 5.6 billion TikTok videos to a popular AI model repository.
Updated on Oct. 6, 2026 in Artificial Intelligence

Live Poll
Do you believe it is acceptable for developers to scrape public social media data for AI?
A developer known as hashfunction has released a 460 GB dataset containing metadata for 5.6 billion TikTok videos on Hugging Face. The collection covers content spanning from July 2014 through October 2026.
Why it matters
AI developers utilize large-scale video metadata to train machine learning models designed to predict viral trends and consumer behavior. This massive data dump raises significant concerns regarding automated scraping practices.
The dataset totals 460 GB in Parquet files and has already reached 1,181 downloads. The scraper utilized reverse-engineered request signatures and spoofed TLS handshakes to mimic Android device activity.
The players
Hugging Face
This American company operates a widely used platform where the machine learning community shares models, datasets, and demo applications.
TikTok
This popular short-form video hosting service is owned by ByteDance and faces ongoing scrutiny regarding its data privacy policies and API security.
DataSocial
This firm provides specialized software tools for data scraping, including code packages currently priced at $1,699.
The details
The developer acquired the data by accessing TikTok's private mobile API, bypassing standard barriers by spoofing device identities. While the dataset is licensed under CC BY-NC 4.0, TikTok's terms of service explicitly prohibit automated scraping without written authorization.
Timeline
The dataset contains information collected between July 2014 and October 2026.
Data collection intensified in a three-week period that included 5.94 billion videos.
The dataset was published on October 6, 2026.
The Tech Race
This incident highlights the escalating conflict between platforms protecting their proprietary data and developers seeking massive inputs for AI training. It follows a tightening legal environment, including the 2025 Reddit lawsuits against scraping firms, which established a precedent for platform defense.
Users should be aware that public metadata from their social media activity is increasingly being aggregated to train commercial AI models without individual consent. This practice may influence the future predictive algorithms that govern the content feeds individuals interact with daily.
The takeaway
The proliferation of massive datasets scraped from private mobile APIs underscores the fragility of digital privacy in the age of AI. Readers should consider adjusting their platform privacy settings to limit the visibility of their activity to third-party scrapers.
Further reading
Find more context on the ethics of AI model training on our Artificial Intelligence page.
More information
Review the technical capabilities and scraping software offerings on the DataSocial official website.
Source note: This article includes information reported by Decrypt.
Live Poll
Do you believe it is acceptable for developers to scrape public social media data for AI?







