Tools, standards, and infrastructure for protecting human-made music, art, language, and craft — from training-data extraction, deceptive substitution, and the quiet erosion of attribution. The manifesto's anti-extraction thesis rendered into operational practice.
Of the seven catalog domains, Preservation & Provenance launches first because the work in this domain is racing the deployment curve. Generative models are already trained on un-consented creative work; large-language models routinely answer in the voice of named writers without their permission; synthetic media is reaching the maturity to substitute for human performers without disclosure. The tools that protect, attribute, and compensate are operational today — not speculative. The catalog's job here is to make them findable.
Open standards and verification infrastructure that let images, video, audio, and documents carry cryptographic proof of origin and edit history — so platforms, publishers, and end-readers can distinguish authentic media from synthetic substitution.
The Coalition for Content Provenance and Authenticity. Open technical standard for embedding tamper-evident provenance into media files. Steering members include Adobe, Microsoft, BBC, Intel, and the New York Times. The reference spec other tooling builds on.
The end-user-facing implementation of C2PA. A small "CR" badge that opens a panel showing how a piece of content was made, what edits were applied, and which AI tools (if any) touched it. Increasingly visible in pro creative tooling and major publisher CMSs.
A consortium of news organizations (BBC, CBC, Microsoft, the New York Times) building newsroom-grade provenance pipelines on the C2PA stack. Targets the long-form, attribution-sensitive content where deepfake substitution does the most damage.
Capture-time provenance for high-stakes use cases — insurance claims, humanitarian documentation, supply-chain verification. SDKs and capture apps that bake C2PA assertions into media at the moment of recording, before any opportunity for tampering.
Open-source tools that let creators defend their work against unauthorized inclusion in training datasets — either by cloaking style so models can't reliably mimic it, or by poisoning the data so models that ingest it degrade rather than improve.
Adversarial perturbations applied to an artist's images that are imperceptible to humans but disrupt style-mimicry models. From Ben Zhao's lab at the University of Chicago. Free desktop app and webGlaze service. Now the default first step many illustrators take before posting work.
The aggressive companion to Glaze. Where Glaze defends a single artist's style, Nightshade introduces poison samples designed to corrupt models that scrape unconsented training data — making large-scale unauthorized scraping costlier than licensing.
Mist
An independent style-cloaking project preceding Glaze. Different perturbation methodology, similar goal. Worth noting because real-world creators often combine multiple defenses, and the field is evolving fast enough that one-tool reliance is fragile.
Infrastructure that lets creators see whether their work was included in major training datasets and assert opt-out preferences that downstream model trainers honor. The consent layer the open-web's training-data pipeline never built.
The organization behind Have I Been Trained, the Source.Plus dataset transparency tools, and the Do Not Train opt-out registry honored by Stability, Hugging Face, Stable Diffusion 3, and others. The closest the field has to a unified opt-out plumbing.
The public-facing dataset-search tool from Spawning. Drop in an image or URL and see which of the major training datasets (LAION-5B, etc.) included it. The first time many creators discovered their work had already been ingested.
Higher-resolution view of what's in major image training sets — metadata, source, captioning. Useful for researchers, journalists, and creators trying to understand what was actually trained on, beyond a simple yes/no answer.
The Distributed AI Research Institute, founded by Timnit Gebru. Independent research on the data, labor, and harm patterns behind dominant AI systems — the academic and policy foundation a lot of consent-and-provenance work builds on.
Community-owned or community-aligned platforms with native provenance, anti-scraping protections, and economics built around the creator rather than the algorithmic ad market. Where artists, writers, and filmmakers actually want to publish in 2026.
The portfolio-and-social platform that absorbed a wave of illustrator migration when other platforms quietly opted creators into AI training. Native Glaze integration, explicit no-train policy, community moderation. The current default for serious illustrators.
ActivityPub-based photo-sharing across the fediverse. No central owner can change the terms of service to opt creators into training overnight. Lower friction than self-hosting, higher control than commercial platforms.
Network-style research and reference catalog. Member-supported (no ads, no algorithmic feed). The default sketchbook for designers, writers, and curators who want a slow, shaped, attributable creative practice rather than the doomscroll.
A protocol-based social network where users own their handles and can move between providers. Not a creator platform per se, but the underlying AT Protocol is increasingly the substrate for portable, scrape-resistant creator identity across the open web.
Open speech datasets and audio archives for languages and oral traditions at risk — recorded and stewarded with the communities that speak them, on terms the communities set. The opposite of extractive scraping: consented, attributed, returnable.
A volunteer-built oral-history archive of every language in the world. Open recordings, collaborative documentation, partnerships with revitalization efforts. The closest thing to a free public substrate for endangered-language speech data.
Field linguistics work focused on indigenous and endangered languages. Talking dictionaries, mobile recording infrastructure, and capacity-building grants for community linguists. Long-term partner relationships, not extractive fieldwork.
A catalog and collaborative documentation hub backed by the Alliance for Linguistic Diversity. Aggregates resources from dozens of academic, indigenous, and community-led archives.
Crowdsourced multi-language speech dataset under CC0. Coverage is uneven across languages, but where it exists it's one of the few large-scale speech corpora available without restrictive licensing — and the pipeline for adding under-represented languages is open.
Capital aligned with cultural-heritage stewardship, creator rights, journalism integrity, and provenance infrastructure. The funder ecosystem that's already moving against the extraction-first AI playbook — and the most realistic source of multi-year support for projects in this domain.
Long-running funder of journalism infrastructure, civic technology, and the institutions of an informed democracy. Active in newsroom-side AI provenance work — including support for newsroom adoption of content credentials and synthetic-media detection tooling.
Active funder of journalism, criminal justice, and technology in the public interest. The On Nigeria, Nuclear Challenges, and Journalism & Media programs all touch the integrity-of-information work that provenance and creator-rights tooling supports.
The largest US funder of the humanities and cultural heritage work. Public Knowledge and Higher Learning programs fund archives, libraries, and the institutional stewardship that endangered-language and creator-rights work depends on.
Hewlett's Cyber Initiative and US Democracy program have moved meaningfully into responsible-AI work, including funder-collaborative work on AI's effects on civic information, creator economies, and democratic institutions.
The largest private funder of digital-rights work globally. Independent media, civil society technology, and the human-rights infrastructure that creator-consent and provenance tools ultimately rest on.
This page is a sample of the Preservation & Provenance domain. Inclusion is editorial — not a commercial endorsement, not a vetting of every claim a project makes, and not a recommendation that any single tool fits every creator's situation. Foundations listed here are not partners; they're documented as the funder ecosystem most aligned with this work, public information, no relationship implied. If you build in this space and we missed you, submit your project below.
Beyond the catalog-wide exclusions, Preservation & Provenance has its own line. Tools that wear the language of authenticity but operate as surveillance, persuasion, or rights-laundering infrastructure are not the work. Naming the line is part of the work.
If your work fits the Preservation & Provenance domain — and is honest about which line it sits on — we want it in the catalog. Submissions are reviewed editorially before inclusion. We may reach out for clarifying details.
Pulled live from the AI for Planet catalog — reviewed editorially, with maturity and evidence flagged. Inclusion is not endorsement; see the curatorial note below. Submissions for the Preservation & Provenance domain open Q3 2026.