# Enfilade — enfilade.guide # # This file is assembled. Editing it here is pointless: the next assembly # overwrites it. # # WHAT IS CLOSED HERE: THE FILTER STATES, THE MAP TILES, AND THE CRAWLERS # THAT GATHER TEXT TO TRAIN MODELS. NOTHING ELSE. # # A page of the catalogue narrowed by one argument — /{lang}/{city}?value=5 # and its like — is not a page. Nobody wrote a heading for it, it does not # keep, and it lists the same museums in another order. Every page stays # open, and so does the pagination that leads through the entries. Of the # arguments, only these are closed; of the paths, only the tiles of the map, # which are data for the map and not pages. # # Until 4 September 2026 nothing here was closed at all, and the reason # given for that was a good one: to read "do not index" a crawler has to # fetch the address, and a rule in this file forbids the fetch and with it # the reading. The reason was true of Google. Two things were measured on # 4 September, and after them it is no longer enough. # # THE FIRST IS THE SHARE. Over four days the filter states took 29 per cent # of everything Bing asked for, 35 of Amazon, 22 of GPTBot and the same of # Meta, 16 of Apple. Google spent 8 per cent, which is bearable. Bing asked # for sixty-two pages in those four days and eighteen of them were states: # a crawler with a small budget was spending a third of it on addresses # that must never appear in a result. # # THE SECOND IS WHO READS WHAT. `` is # an instruction to a search index. The crawlers of language models read # this file and do not read that tag, so for them the instruction is not # there at all. Of the 1229 pages GPTBot took in those four days, 267 were # filter states — the body of text that stands for Enfilade inside a model, # diluted with our own switches. # # THE TAG STAYS IN THE PAGE. Google prefers a rule here for filters and # says the tag spends crawl budget to no purpose, and that is true of the # addresses it has yet to find. For the ones it already knows the tag is # all there is until it asks again, and it is also what is read by anyone # who ignores this file. # # WHAT IS LEFT OPEN ON PURPOSE. The pagination, because the pages of a list # are how the entries are found. `simple` and `all` — the plain view of the # same list, and the scope of the top bar — because both carry a canonical # to the clean address, and a crawler that may not fetch the page never # reads it. The exports of the catalogue, because we offer them to # assistants in /llms.txt, and a thing offered in one file and forbidden in # another is believed in neither. # # NO REDIRECT LANDS IN A CLOSED SPACE. The section's former addresses at # m1kle.ru answer 301 into this one, and not one of them lands on an address # closed below. # # All of this is measured rather than remembered. One watchdog walks the # site, counts the states a crawler can reach, and fails if any link # combines two arguments or if any state omits its tag. A second reads every # address of the sitemap and every link of /llms.txt back through the rules # below, and fails if a single one of them comes out closed. It was written # on the day the rules returned, and it caught the first version of them: # "t=" also matches "export=". # # NO PRIVATE PAGE IS NAMED HERE, AND THAT IS DELIBERATE. All of them sit # behind a password and answer 401 to anyone, so there is nothing to forbid: # a robot meets a door and leaves of its own accord, at the cost of one # request. And robots.txt is read by everybody, so a file listing what is # closed is a ready-made signpost for anyone looking for exactly that. # # ASSISTANTS BUILT ON LANGUAGE MODELS ARE WELCOME, AND THE WELCOME NOW HAS # A SHAPE. Until 4 September 2026 this file set no prohibition against any # crawler, and said so. It sets one now, and the reason is not a change of # heart: the licence was rewritten that day and says two things plainly. # # Allowed, and we want it: indexing this site for search, finding it as a # source, and quoting it briefly with a link to the page quoted. Not # allowed without written permission: training or fine-tuning models, # building training datasets, model distillation, text and data mining, # and reprinting our texts in full or in substantial part. # # The terms in full are on /en/licences, in all four languages. The same # terms are stated for machines three times over, because three different # readers look in three different places: the groups below and the # Content-Signal line here, the RSL licence at /licence.xml, and the # reservation at /.well-known/tdmrep.json — the last of these being what # European law reads when it asks whether the text and data mining # exception was reserved against. # # The index for assistants is /llms.txt. # # The License line below is RSL 1.0 (rslstandard.org): a machine-readable # licence, global for the whole site, and the standard puts it here — at # the top, outside any User-agent group. Validators that know only the # RFC 9309 records call it an unknown directive; crawlers ignore records # they do not know (RFC 9309, section 2.2.4), so nothing about crawling # or indexing changes. License: https://enfilade.guide/licence.xml User-agent: * Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=no # Filter states. Two lines to an argument, because the only wildcards # robots.txt understands are "*" and "$": the boundary before the name # has to be spelt out, or "t=" would also match "export=". Disallow: /*?access= Disallow: /*&access= Disallow: /*?age= Disallow: /*&age= Disallow: /*?archive= Disallow: /*&archive= Disallow: /*?building= Disallow: /*&building= Disallow: /*?category= Disallow: /*&category= Disallow: /*?collection= Disallow: /*&collection= Disallow: /*?collection_extended= Disallow: /*&collection_extended= Disallow: /*?company= Disallow: /*&company= Disallow: /*?day= Disallow: /*&day= Disallow: /*?distinction= Disallow: /*&distinction= Disallow: /*?exact= Disallow: /*&exact= Disallow: /*?free= Disallow: /*&free= Disallow: /*?gorod= Disallow: /*&gorod= Disallow: /*?interiors= Disallow: /*&interiors= Disallow: /*?kids= Disallow: /*&kids= Disallow: /*?kind= Disallow: /*&kind= Disallow: /*?mexact= Disallow: /*&mexact= Disallow: /*?mvalue= Disallow: /*&mvalue= Disallow: /*?now= Disallow: /*&now= Disallow: /*?q= Disallow: /*&q= Disallow: /*?r= Disallow: /*&r= Disallow: /*?scope= Disallow: /*&scope= Disallow: /*?since= Disallow: /*&since= Disallow: /*?t= Disallow: /*&t= Disallow: /*?time= Disallow: /*&time= Disallow: /*?ulica= Disallow: /*&ulica= Disallow: /*?until= Disallow: /*&until= Disallow: /*?upto= Disallow: /*&upto= Disallow: /*?value= Disallow: /*&value= Disallow: /*?when= Disallow: /*&when= Disallow: /*?where= Disallow: /*&where= # Map tiles. /karta/{city}/{z}/{x}/{y}.mvt are vector tiles fetched by # the map script: not pages, no words in them, listed in no index. On # 14 September 2026 they took 43 of the 151 requests Googlebot made in # a day — more than a quarter of a small crawl budget. The map page # itself stays open; what is indexed is its text, not its canvas. Disallow: /karta/ # Arguments that change nothing in the page: the scope of the top bar # (`all`), the plain view of the same list (`simple`), and the tracking # marks that other sites append to links to us. All carry a canonical to # the clean address, which is enough for Google. Yandex treats a canonical # as a hint and does not merge duplicates by it, so it is told here. # A Disallow would be worse than useless: it would forbid the fetch and # with it the reading of the canonical. Clean-param: simple&all&utm_source&utm_medium&utm_campaign&utm_term&utm_content&yclid&gclid&fbclid&_openstat&from # Crawlers that gather text to train models. The terms are on # https://enfilade.guide/en/licences and they are short: # search, being found as a source and short quotations with a link are # allowed; training, fine-tuning, dataset building, distillation and text # and data mining are not, without written permission. # # Where an operator separates its crawlers by purpose, only the training # one is named here: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, # Claude-User, Applebot, Googlebot and the rest keep the access above. # Where no such separation is offered, the crawler is named in full. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Amazonbot Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: cohere-ai Disallow: / User-agent: AI2Bot Disallow: / User-agent: PanguBot Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: Omgilibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Timpibot Disallow: / User-agent: Diffbot Disallow: / Sitemap: https://enfilade.guide/sitemap.xml