diff --git a/public/.well-known/agent-skills/index.json b/public/.well-known/agent-skills/index.json index 1406b70b..f9a7ce19 100644 --- a/public/.well-known/agent-skills/index.json +++ b/public/.well-known/agent-skills/index.json @@ -6,7 +6,7 @@ "type": "skill-md", "description": "Query and apply The Website Specification — a platform-agnostic specification of what a good website does. Use when the user asks what their site should have, whether something is required, how to audit a URL, what's missing for agent readiness, or anything else where you'd otherwise be guessing at web best practice. Backs answers with primary sources and ships an MCP server with search, list, fetch, checklist, and audit tools.", "url": "/.well-known/agent-skills/specification-website/SKILL.md", - "digest": "sha256:9d40c04e08404c5e1d3fc754e9f007c0900230f1fffc73786e192c03ca74087b" + "digest": "sha256:ac28752d4ce916f4ca8dff31cae705ef4bc7d882c1656df685b5b92556c5ac0e" } ] } diff --git a/public/.well-known/agent-skills/specification-website/SKILL.md b/public/.well-known/agent-skills/specification-website/SKILL.md index 4612d0e0..359bb62a 100644 --- a/public/.well-known/agent-skills/specification-website/SKILL.md +++ b/public/.well-known/agent-skills/specification-website/SKILL.md @@ -5,7 +5,7 @@ description: Query and apply The Website Specification — a platform-agnostic s # specification.website -The Website Specification is a single source of truth for what a good website does. Ten categories, 164 pages, every item tagged with a status. It ships in three machine-readable forms: per-page Markdown, llms.txt / llms-full.txt, and an MCP server. +The Website Specification is a single source of truth for what a good website does. Ten categories, 165 pages, every item tagged with a status. It ships in three machine-readable forms: per-page Markdown, llms.txt / llms-full.txt, and an MCP server. ## When to use this skill diff --git a/public/_headers b/public/_headers index 7abdda2b..aba31e22 100644 --- a/public/_headers +++ b/public/_headers @@ -33,6 +33,13 @@ # /reports means a script shipped without integrity. Flip to enforcing # `Integrity-Policy` once the /reports stream confirms a clean run. Worked # example for /spec/security/subresource-integrity/. +# +# tdm-reservation: 0 is the W3C TDMRep declaration that we do NOT reserve the +# right to object to text and data mining — the machine-readable counterpart to +# the CC BY 4.0 content licence, and the inverse of the Article 4 CDSM opt-out. +# Keep it at 0: a site whose whole point is being read by machines has no +# business reserving against them. Worked example for +# /spec/agent-readiness/tdmrep/. /* Strict-Transport-Security: max-age=63072000; includeSubDomains @@ -47,6 +54,7 @@ Cross-Origin-Resource-Policy: same-site X-Frame-Options: DENY No-Vary-Search: params=("utm_source" "utm_medium" "utm_campaign" "utm_content" "utm_term" "gclid" "fbclid" "msclkid" "mc_cid" "mc_eid" "ref"), key-order + tdm-reservation: 0 Link: ; rel="describedby"; type="text/markdown"; title="Site index for LLMs", ; rel="alternate"; type="text/markdown"; title="Full content as Markdown", ; rel="api-catalog"; type="application/linkset+json", ; rel="mcp"; type="application/json"; title="MCP server card", ; rel="related"; title="MCP endpoint", ; rel="service-desc"; type="application/json"; title="A2A agent card", ; rel="related"; title="A2A endpoint", ; rel="agent-skills"; type="application/json"; title="Agent Skills discovery index", ; rel="ai-catalog"; type="application/ai-catalog+json"; title="Agentic Resource Discovery catalog", ; rel="sitemap"; type="application/xml", ; rel="alternate"; type="application/rss+xml"; title="Feed of spec changes", ; rel="alternate"; type="application/rss+xml"; title="Changelog feed", ; rel="security"; type="text/plain" # Self-hosted Plausible tracker. Not fingerprinted and refreshed in place by diff --git a/public/og-default.png b/public/og-default.png index 6ae92ff3..ac9b7623 100644 Binary files a/public/og-default.png and b/public/og-default.png differ diff --git a/public/og/checklist.png b/public/og/checklist.png index f5d34d1a..87956f64 100644 Binary files a/public/og/checklist.png and b/public/og/checklist.png differ diff --git a/public/og/spec.png b/public/og/spec.png index 339214f2..d113e0c7 100644 Binary files a/public/og/spec.png and b/public/og/spec.png differ diff --git a/public/og/spec/agent-readiness.png b/public/og/spec/agent-readiness.png index da787bd9..f50a5641 100644 Binary files a/public/og/spec/agent-readiness.png and b/public/og/spec/agent-readiness.png differ diff --git a/public/og/spec/agent-readiness/tdmrep.png b/public/og/spec/agent-readiness/tdmrep.png new file mode 100644 index 00000000..c5c08026 Binary files /dev/null and b/public/og/spec/agent-readiness/tdmrep.png differ diff --git a/src/content/changelog/2026-07-28-tdmrep.md b/src/content/changelog/2026-07-28-tdmrep.md new file mode 100644 index 00000000..419cb861 --- /dev/null +++ b/src/content/changelog/2026-07-28-tdmrep.md @@ -0,0 +1,8 @@ +--- +title: Reserving mining rights is a legal act, not a crawler block +date: "2026-07-28" +type: added +relatedSlugs: [tdmrep, content-signals, robots-for-ai-crawlers] +--- + +Added a page on [TDM reservation](/spec/agent-readiness/tdmrep/), the W3C convention for declaring in machine-readable form whether you reserve the right to object to text and data mining. It is easy to mistake for another crawler blocker, and it is not one: `robots.txt` decides whether a bot may fetch a page, while `tdm-reservation` decides whether what it fetched may be mined — and under Article 4 of the EU copyright directive, a reservation that is not machine-readable does not count at all. This site now sends `tdm-reservation: 0` on every response, declining to reserve, which is the honest counterpart to a CC BY 4.0 licence. diff --git a/src/content/spec/agent-readiness/agent-readiness-overview.md b/src/content/spec/agent-readiness/agent-readiness-overview.md index e9cc8751..e785f4a3 100644 --- a/src/content/spec/agent-readiness/agent-readiness-overview.md +++ b/src/content/spec/agent-readiness/agent-readiness-overview.md @@ -6,7 +6,7 @@ summary: "Agent readiness is the set of choices that make a site legible to AI a status: recommended order: 10 appliesTo: [all] -relatedSlugs: [llms-txt, markdown-source-endpoints, robots-for-ai-crawlers, stable-urls, structured-data-for-agents, machine-readable-formats, agent-skills-discovery, mcp-and-tool-discovery, a2a-agent-cards, web-bot-auth, webmcp, nlweb, server-side-rendering, agentic-resource-discovery] +relatedSlugs: [llms-txt, markdown-source-endpoints, robots-for-ai-crawlers, tdmrep, stable-urls, structured-data-for-agents, machine-readable-formats, agent-skills-discovery, mcp-and-tool-discovery, a2a-agent-cards, web-bot-auth, webmcp, nlweb, server-side-rendering, agentic-resource-discovery] updated: "2026-06-08T00:00:00.000Z" sources: - title: "Is It Agent Ready?" diff --git a/src/content/spec/agent-readiness/content-signals.md b/src/content/spec/agent-readiness/content-signals.md index 2f06208b..09f8f5f4 100644 --- a/src/content/spec/agent-readiness/content-signals.md +++ b/src/content/spec/agent-readiness/content-signals.md @@ -6,7 +6,7 @@ summary: "Add Content-Signal directives to robots.txt to declare whether AI craw status: optional order: 42 appliesTo: [all] -relatedSlugs: [robots-for-ai-crawlers, robots-txt, agent-readiness-overview, web-bot-auth] +relatedSlugs: [robots-for-ai-crawlers, robots-txt, tdmrep, agent-readiness-overview, web-bot-auth] updated: "2026-07-28T00:00:00.000Z" sources: - title: "IETF AI Preferences WG (aipref) — drafts" diff --git a/src/content/spec/agent-readiness/robots-for-ai-crawlers.md b/src/content/spec/agent-readiness/robots-for-ai-crawlers.md index fcb49cdb..ef0db443 100644 --- a/src/content/spec/agent-readiness/robots-for-ai-crawlers.md +++ b/src/content/spec/agent-readiness/robots-for-ai-crawlers.md @@ -6,7 +6,7 @@ summary: "Major AI vendors publish named user-agents for their crawlers. Setting status: recommended order: 40 appliesTo: [all] -relatedSlugs: [robots-txt, content-signals, agent-readiness-overview, llms-txt, web-bot-auth] +relatedSlugs: [robots-txt, content-signals, tdmrep, agent-readiness-overview, llms-txt, web-bot-auth] updated: "2026-05-29T12:14:17.000Z" sources: - title: "OpenAI — GPTBot" diff --git a/src/content/spec/agent-readiness/tdmrep.md b/src/content/spec/agent-readiness/tdmrep.md new file mode 100644 index 00000000..b3cf54d3 --- /dev/null +++ b/src/content/spec/agent-readiness/tdmrep.md @@ -0,0 +1,90 @@ +--- +title: "TDM reservation (TDMRep)" +slug: tdmrep +category: agent-readiness +summary: "Declare in machine-readable form whether you reserve the right to object to text and data mining, using the `tdm-reservation` HTTP header, a `` element, or `/.well-known/tdmrep.json`. Its force is legal, not technical." +status: optional +order: 43 +appliesTo: [all] +relatedSlugs: [content-signals, robots-for-ai-crawlers, robots-txt, agent-readiness-overview] +updated: "2026-07-28T00:00:00.000Z" +sources: + - title: "TDM Reservation Protocol (TDMRep) — Final Community Group Report" + url: "https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20240202/" + publisher: "W3C" + - title: "Directive (EU) 2019/790 on copyright in the Digital Single Market — Article 4" + url: "https://eur-lex.europa.eu/eli/dir/2019/790/oj" + publisher: "EUR-Lex" + - title: "IANA Well-Known URIs registry" + url: "https://www.iana.org/assignments/well-known-uris/well-known-uris.xhtml" + publisher: "IANA" + - title: "European Commission — Consultation on protocols for reserving rights from text and data mining" + url: "https://digital-strategy.ec.europa.eu/en/consultations/commission-launches-consultation-protocols-reserving-rights-text-and-data-mining-under-ai-act-and" + publisher: "European Commission" +--- + +## What it is + +TDMRep is a small, finished convention for stating one thing: whether you reserve the right to object to **text and data mining** of your content. It carries a single value, `1` (reserved) or `0` (not reserved), plus an optional URL pointing at a human- or machine-readable licensing policy. + +Three places can carry it, in ascending order of specificity: + +``` +tdm-reservation: 1 +tdm-policy: https://example.com/policies/tdm.json +``` + +```html + + +``` + +```json +[ + { "location": "/", "tdm-reservation": 1 }, + { "location": "/press/*.jpg", "tdm-reservation": 0 } +] +``` + +The last of these lives at `/.well-known/tdmrep.json` and uses `robots.txt` path-matching rules, including `*` and `$`. When more than one mechanism applies to the same resource, the more specific wins: an HTML `` beats an HTTP header, which beats the well-known file. The file is the site-wide default; the header and the meta element are the per-resource overrides. + +**This is not a crawler blocker, and reading it as one is the most common mistake.** `robots.txt` and AI-crawler directives answer *may you fetch this?* TDMRep answers a later and separate question: *having fetched it, may you mine it?* A crawler that respects `Disallow:` never sees your content. A crawler that respects `tdm-reservation: 1` reads the page quite happily and is then on notice that training on it was not licensed. One is a gate; the other is a notice. Shipping TDMRep blocks nobody, and a site that wants a gate still needs [robots directives for AI crawlers](/spec/agent-readiness/robots-for-ai-crawlers/). + +## Why it matters + +The consumer of this signal is mostly a legal system rather than a crawler, and that is the honest way to judge it. + +Article 4 of the EU Copyright in the Digital Single Market Directive permits text and data mining of lawfully accessible works **unless the rightsholder has expressly reserved that right in a machine-readable way**. The reservation has to be machine-readable to count. TDMRep exists to be that machine-readable form. If you do nothing, Article 4's default is that mining is permitted. + +That obligation now has a counterpart on the other side. Providers of general-purpose AI models must have a policy for identifying and respecting Article 4 reservations, and in December 2025 the European Commission opened a consultation to agree a list of protocols that qualify — with TDMRep among the named candidates. The output of that process, not crawler adoption statistics, is what will decide whether this convention matters. + +Technical uptake, meanwhile, is thin and worth stating plainly. Spawning AI reads TDMRep in the API it operates for model trainers, and Common Crawl measures its prevalence when surveying opt-out signals. No major crawler operator has published a commitment to honour it in the way they publish commitments about `robots.txt`. Do not ship this expecting enforcement. Ship it because an unstated reservation is, under Article 4, no reservation at all. + +## How to implement + +**Decide the value first, then pick the mechanism.** The value is a licensing decision, not a technical one. `1` means anyone wanting to mine your content must seek permission or follow your `tdm-policy`. `0` means they may mine it freely. There is no "ask me nicely" middle option, and inventing one breaks parsers. + +**Prefer the HTTP header for a whole-site answer.** It applies to every response, including images, PDFs, and JSON that have no `` to put a `` in. That coverage is the reason to reach for it first. + +**Use `/.well-known/tdmrep.json` when the answer varies by path**, and remember it is the weakest of the three — anything you set in a header or a `` on a specific resource overrides it. + +**Set `tdm-policy` whenever you set `tdm-reservation: 1`.** A bare reservation tells a would-be licensee that the answer is no, without telling them who to ask. That is a wasted opportunity and, if you are hoping to license, self-defeating. + +**Keep it consistent with everything else you publish.** A `tdm-reservation: 1` next to a permissive content licence, or next to `Content-Signal: ai-train=yes`, is a contradiction someone will eventually have to resolve — probably not in your favour. + +> **This site ships it.** specification.website sends `tdm-reservation: 0` on every response, set in [`public/_headers`](https://github.com/jdevalk/specification.website/blob/main/public/_headers). The content is CC BY 4.0 and we want it mined, indexed, and trained on; declaring `0` states that in the same machine-readable form that a refusal would take, rather than leaving it to be inferred from a licence file. + +## Common mistakes + +- Treating it as a block. It stops no fetch and triggers no `403`. If you want fetching stopped, that is a different mechanism. +- Sending `tdm-reservation: 1` while your `robots.txt` invites AI crawlers in, or vice versa. Pick a position and express it consistently. +- Using `true`, `yes`, or `reserved` as the value. Only the integers `1` and `0` are valid; anything else means agents treat the field as unset, so a malformed reservation is the same as no reservation. +- Putting the well-known file in place and assuming it wins. It is the lowest-precedence mechanism of the three. +- Expecting the declaration to be retroactive. It says nothing about copies already in a training corpus. + +## Verification + +- `curl -sI https://example.com/ | grep -i tdm-` shows the header on a real response. +- Fetch a non-HTML asset — an image, a PDF — and confirm the header is present there too. That is the coverage the `` form cannot give you. +- If you publish `/.well-known/tdmrep.json`, confirm it parses as a JSON array and that every entry has a `location` and an integer `tdm-reservation`. +- Check the value against your actual licence. The two should say the same thing.