# ZNetLab — https://znetlab.com/ # # This file belongs at the ROOT of the domain. It used to sit inside the # teaching half's folder, where crawlers never look for it — a robots.txt at # /learn/robots.txt is not read by anything. # # Everything public is crawlable. The exclusions below are not hiding # anything: they are pages that would be noise in an index, because there is # nothing to read on them until somebody signs in or JavaScript runs. User-agent: * Allow: / # The lab platform is an application behind a sign-in. The home page is what # should rank for it, and it links there. # # NOT disallowed, deliberately: /lab/ carries `noindex` in its own head, and # the two do not combine the way people expect. A disallowed URL is never # fetched, so the noindex inside it is never read — and a page that is linked # from elsewhere can still end up listed, as a bare URL with no description, # precisely because the crawler was forbidden from reading the instruction # telling it not to. Letting it in to read `noindex` is what actually keeps # it out. # # /learn/lab.html is the same case and is also left crawlable. # Build artefacts and test harnesses, not pages anyone searches for. Disallow: /learn/test-engine.html # The PDFs. A PDF cannot carry a robots meta tag — the only way to keep one # out of an index without a server header is here. Their HTML counterparts # in the same folder are left crawlable so their own `noindex` can be read, # for the reason given above. Disallow: /docs/*.pdf$ # Query-string variants of the catalogue are the same page with a filter # applied — indexing them splits the ranking of one page across many. Disallow: /*?q= Disallow: /*&sort= Disallow: /*?tag= # Nothing here changes fast enough to need aggressive crawling, and the # simulator is a large page. Crawl-delay: 1 # ---------------------------------------------------------------- answer engines # # These are already allowed by `User-agent: *` above, and they are named here # anyway, for two reasons. The first is that silence is ambiguous: a reader # checking whether this site minds being summarised finds nothing either way, # and some operators treat an unnamed crawler as a grey area. The second is # that Google-Extended is NOT covered by the wildcard in the way people # assume — it is a separate opt-out token with no crawling of its own, and # leaving it unmentioned is the one case where saying nothing is read as a # refusal. # # The position this takes: everything public here may be read, quoted and # summarised. A product nobody is allowed to describe does not get # recommended, and being recommended accurately is worth more than the # traffic a summary diverts. User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Google-Extended Allow: / User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / User-agent: CCBot Allow: / User-agent: meta-externalagent Allow: / User-agent: Bingbot Allow: / # The plain-text summary written for exactly these readers: what ZNetLab is, # what it does NOT do, how it compares to GNS3, EVE-NG and Packet Tracer, and # which figures are safe to quote. It is not a standard, and it costs one file # to offer; a summariser that ignores it is no worse off than before. # https://znetlab.com/llms.txt Sitemap: https://znetlab.com/sitemap.xml