What happens when a digital vandal decides to build your defensive wall? That’s the core, unsettling question raised by the story of Cara. The art portfolio platform was attacked. The attacker felt remorse. And now, the scraper is working with its creator.
XOOMAR Intelligence
Analyst Take
According to Wired, the platform was subjected to three targeted scrapes in August. The first yielded a 12-terabyte archive of roughly 12 million works, posted publicly by a Reddit user who later turned into an ally. This fight reveals a raw truth: ideological battles over AI training are playing out in real-time on digital infrastructure, with artists caught in the crossfire.
As we previously reported in Artists Built A Site To Escape AI. Scrapers Are Coming For It Anyway, Cara's mission made it a target. The August attacks were not subtle data leaks. They were public performances designed to prove a point and inflict harm. The original post on the subreddit r/DefendingAIArt by the user MandarinDrawnPoppy994 framed the scrape as a "fun project" that cost less than $10.
A second scraper took about 8.5 million links plus metadata to Hugging Face. A third posted 123,000 images to Academic Torrents. For Cara founder Jingna Zhang, this was a coordinated assault. "This is now a targeted attack meant to cause artists pain," she wrote.
The trolls' goal was simple: dismantle the argument that a safe space online is possible. By scraping and publishing the data from a platform whose explicit purpose is to prohibit such use, they weaponized its existence. The message to Cara's 1.5 million artists was clear: your refuge is an illusion. Your consent, in the face of a determined adversary with a script, is optional.
What Does A Scraper See When They Look At An "Off-Limits" Site?
Heft (the scraper behind the first attack) provided a chillingly simple technical post-mortem. To him, a site like Cara isn't a fortress. It's a library with unlocked doors. He scraped its entire public library of images, not for commercial gain, but to answer a technical curiosity about its potential for AI training. He later concluded the dataset wasn't even particularly useful for that purpose.
Scaling is trivial: A single motivated individual with basic software skills and a nominal budget can vacuum up millions of works. Protections are porous: Heft explains that many proposed fixes, like certain login gates or obfuscation tools, can be bypassed by a determined party "in like a few minutes, literally." The real target is data density: Platforms like Cara, which aggregate high-quality, labeled artwork from creators explicitly opposed to AI training, become high-value trophies for ideological scrapers. The attack isn't just about acquiring the art; it's about proving the platform's core promise is technically bankrupt.
This demystification is the first step in Heft's contrition. He moved from seeing artists' work as rows in a dataset to understanding the personal violation. "I saw people sharing how they were having panic attacks over the scrape, how they deleted their entire portfolios from the internet," he told WIRED.
When Your Attacker Apologizes, Do You Teach Them Or Hunt Them?
The story's central tension isn't legal or technical. It's human. After being confronted, Heft didn't just delete his dataset. He apologized sincerely and offered his expertise.
“In retrospect, not only deliberately targeting Cara but presenting it the way I did in the post was cruel and thoughtless,” Heft says. “I missed the consequences that this would have beyond causing a bit of anger.”
Zhang faced a choice familiar in cybersecurity: do you prosecute the ethical hacker who exposed your weakness, or do you enlist them to help fix it? She chose the latter.
Heft joined Cara's Discord as a troubleshooter. His new role is to audit proposed defenses and explain, from an attacker's perspective, exactly how they would fail. This gives Cara's small team something invaluable: an insider's map of their own vulnerabilities. The collaboration is pragmatic, not philosophical. It accepts a grim premise: perfect defense is impossible. Better defense, however, is achievable with the right guide.
Can You Build A Better Alarm Instead Of A Higher Wall?
If you can't stop the scrape, what can you do? The answer Zhang and Heft are developing shifts the battleground from prevention to detection and response. Their new tool is called Lantern.
Lantern is an open-source project separate from Cara. Its function is straightforward but powerful:
How Lantern Works
| Component | Function |
|---|---|
| One-Way Fingerprint | Artists upload a unique hash of their image without storing the image itself on Lantern's servers. |
| Dataset Scanning | The tool regularly scans new, publicly available AI training datasets. |
| Artist Notification | If a fingerprint match is found, the artist gets an alert with a link to the infringing dataset. |
| Actionable Intel | The artist can then pursue a takedown notice or other legal action with specific evidence. |
This tool doesn't prevent the initial theft. It creates a persistent, searchable claim on the artwork that can be tracked across the data ecosystem. It turns a passive violation into an actionable event. "He’s of the belief that no site can be made truly ‘unscrapable,’" Zhang says of Heft. So they are building a tripwire for after the breach happens.
Is An Artist's New Core Skill Knowing How To Hide?
The Cara saga illuminates a brutal evolution for the creative class. The initial phase was shock, as covered in our analysis of The Great AI Art Heist. The second phase was organized protest and the creation of opt-out sanctuaries like Cara. We are now entering a third phase: the normalization of defensive tradecraft.
The implications are profound:
Creative labor now includes security labor: Artists must think like data custodians. Sharing work online is no longer just about promotion; it's a risk-assessment exercise. Sanctuary platforms face inherent pressure: By promising safety, they attract both targeted creators and ideological attackers, creating a sustainability crisis. Cara's server fees spiked due to the scraping traffic, and Zhang launched a GoFundMe that has raised over $100,000 for legal defense. The market will fracture: We will see a clear divergence between "AI-engaged" platforms (where scraping is assumed) and "verified-clean" platforms (where it is explicitly fought). Each will have its own economies, communities, and attendant risks.
This phase is defined by pragmatism over purity. The alliance between Zhang and Heft is a microcosm of that shift. It's an admission that in a war where the legal terrain is barren, you need every tactical advantage you can get, even if it comes from a former enemy.
The watchpoint is no longer whether art can be hidden, but whether tools like Lantern can create enough friction and consequence to make scraping a less attractive sport. The next test will be if this model of defensive co-creation scales, or if it remains a unique footnote in a conflict that grows more automated, and more impersonal, by the day.
The Stakes
- The case exposes the vulnerability of digital "safe spaces" and the limitations of consent when platforms can be forcibly scraped.
- It highlights the volatile ethics of AI data sourcing, where attackers can become collaborators, blurring lines of harm and solution.
- The outcome sets a precedent for how platforms under attack might turn adversaries into allies to build stronger technical and ethical defenses.
Primary Sources & Disclosures
Written by
XOOMAR Insights Team
Research and Editorial Desk
The XOOMAR Insights Team pairs automated research with human editorial judgment. We track hundreds of sources across technology, fintech, trading, SaaS, and cybersecurity, cross-check the facts, and explain what happened, why it matters, and what to watch next. We do not just rewrite headlines. Every article is fact-checked and scored for reliability before it goes live, and we link back to the original sources so you can verify anything yourself.










