Cara Scrape Claim Exposes Artist Platform Limits

On Aug. 14, a Reddit user claimed to have scraped all of Cara, while creator Zhang Jingna said the platform held over 12 million images. The post was later deleted and the user banned, and Zhang wrote in a July blog post that broad protection against AI use is not technically possible in the sense many artists may infer from anti-AI claims.
A Reddit user claimed to have scraped all of Cara, an art platform built by Singapore photographer Zhang Jingna, as part of a "fun project," according to AsiaOne. Zhang disclosed the incident in an Aug. 14 Instagram post, while the user described the collection in the DefendingAIArt subreddit, AsiaOne reported.
The Reddit post was subsequently deleted and the account was banned, according to AsiaOne. Zhang said Cara uses anti-bot and anti-scraping measures and that the platform contained over 12 million images associated with more than one million users.
Zhang wrote on Instagram: "Artists just want a place online where platforms don't feed their work to AI. So I built Cara. But people just won't leave us alone." Her post also criticized the lack of legal and governmental protections for artists, according to AsiaOne.
Limits of technical protection
In a July 27 post on her personal blog, Zhang cautioned artists against platforms that claim complete protection from AI scraping or training. She wrote that such protection is not technically possible in the broad sense many users may infer from the term "protect."
Zhang contrasted broad platform assurances with tools such as Glaze, which she described as having more limited protections against style mimicry in certain cases, rather than protection against all AI uses. She also wrote that terms of service prohibiting AI training do not by themselves prevent unauthorized collection or use of images.
The reported scrape claim has not been independently verified in the material reviewed. Scraping is the automated collection of web-accessible data; collection alone does not establish that content entered a training dataset or was used to train a model.
Dataset governance questions
For teams developing models or operating creator platforms, comparable disputes illustrate the gap between access controls, contractual restrictions, and enforceable data provenance. Bot defenses can raise the cost of bulk collection, but publicly served media can remain accessible to determined collectors unless platforms combine technical controls with monitoring, rate limits, authentication, and clear evidence trails.
The episode also underscores a recurring dataset-governance issue: a platform rule against training does not itself demonstrate downstream compliance by third parties. Practitioners sourcing image data need records of origin, permissions, collection methods, and removal processes, especially where creator communities explicitly object to model training.
Key Points
- 1A now-deleted Reddit post claimed to have scraped all of Cara, while Zhang said the platform contained over 12 million images.
- 2Zhang's public comments distinguish anti-scraping defenses from guarantees against model training, a limitation relevant to dataset governance.
- 3Comparable incidents show that terms of service and bot controls alone cannot establish dataset provenance, consent, or downstream compliance.
Scoring Rationale
The unverified scrape claim concerns a large alleged collection of creator images and raises practical questions about web-data provenance and anti-bot controls. It is relevant to teams building image datasets and platforms, but it does not document a confirmed model-training event or a new technical vulnerability.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

