Developer publishes metadata for 5.6 billion TikTok videos on Hugging Face
A developer has published metadata for 5.6 billion public TikTok videos on Hugging Face for free download.
Company
TikTok faces significant developments regarding data scraping, e-commerce integration, and legal precedents involving AI voice cloning. Recently, a developer published a 460-gigabyte dataset on Hugging Face containing metadata for 5.6 billion public TikTok videos spanning from July 2014 through October 2026. This collection, gathered via a private mobile API without logins, is free for non-commercial use and attracts AI developers looking to train models, despite TikTok's strict scraping limitations. In e-commerce, TikTok now offers a new conversational AI shopping assistant alongside an in-app one-click checkout feature. Developed with partners including Salesforce, Shopify, Shoplazza, and Stripe, these features allow users to discover products and make direct purchases from their For You feed. Meanwhile, a landmark legal ruling in Japan impacts content on the platform. A Tokyo court ruled that human voices are legally protected after anime voice actor Kenjiro Tsuda challenged an anonymous TikTok account that used artificial intelligence to clone his voice for narration. Although the court dismissed the removal request because the account was already deleted, the ruling establishes that unauthorized voice cloning infringes on publicity rights. These events highlight TikTok's evolving landscape across data privacy, social commerce, and intellectual property.
3 stories · Updated
A developer has published metadata for 5.6 billion public TikTok videos on Hugging Face for free download.
TikTok announced the rollout of an in-app AI shopping assistant and a one-click checkout feature for direct purchases from brands.
A Tokyo court ruled that the human voice is protected by law in a landmark case involving anime voice actor Kenjiro Tsuda and an anonymous TikTok account.