Machine-readable access to Need to Know IT
This page is for AI agents, assistants and automated crawlers, and for anyone checking what one can see here. It says plainly what's exposed today and what isn't yet, rather than describing a roadmap as if it already shipped.
The site index: llms.txt
/llms.txt lists the main sections of the site and points to the full XML sitemap for every live page. It's generated from the live page tree on every request, using the same filtering the sitemap itself uses, so it can't drift out of date the way a hand-maintained file would.
What's here today
- Full-text HTML articles and buying guides, each with structured data (JSON-LD: article, product, review, FAQ, breadcrumbs where relevant) already embedded in the page.
- Clean Markdown twins of 121 articles, covering
buying guides, reviews, VS comparisons, plus
1 individually opted-in article
of other kinds. Append
.mdto the article's URL, e.g./blog/best-nas-australia.md: full text, no nav/footer/ad markup, a source line pointing back to the canonical page. Not yet built for brand hubs, how-to guides, 96 of 97 informational articles, those are HTML only for now. - An XML sitemap (/sitemap.xml) covering every indexable live page.
- An Atom feed of the latest 30 articles at
/feeds/blog.atom, advertised via
<link rel="alternate">on every page. - "Ask AI about this page" links on every article that has a Markdown twin, pre-filled to open the twin in ChatGPT, Claude or Perplexity.
- Real Australian pricing, retailer availability and consumer-law context on relevant pages, sourced and refreshed on a stated schedule rather than written once and left to go stale.
- Conditional GET on every machine endpoint here.
llms.txt, this page, the sitemap, the feed and every Markdown twin sendETagandLast-Modifiedand answerIf-None-Match/If-Modified-Sincewith304, so re-checking for changes is cheap. - Three separate provenance dates in each Markdown twin's header: Published, Last updated and Facts last checked against sources.
What "facts last checked" means, and what it does not
These are deliberately three different dates, because editing a page is not the same as re-verifying what it claims. Last updated moves when the page is substantively changed. Facts last checked against sources moves only when this site's fact index actually re-checks that page's claims against their sources, which happens nightly. It is never set by saving or republishing a page.
The date shown is the oldest check across the page's claims, not the newest, because a page is only as current as its least-recently-checked figure. If a page shows no fact-checked date, that means no checkable claim on it has been indexed - it does not mean the page was checked and found clean. We would rather say nothing than imply a check we did not run.
Crawler policy
Every AI crawler is allowed, and that is a decision rather than an oversight.
robots.txt names the retrieval crawlers
(OAI-SearchBot, Claude-SearchBot,
PerplexityBot), the user-initiated agents
(ChatGPT-User, Claude-User,
Perplexity-User) and the training crawlers
(GPTBot, ClaudeBot, Google-Extended)
explicitly, all allowed, so that the position is legible rather than implied by
silence. Only /cms/ and /admin/ are disallowed.
What isn't here yet
- Markdown twins for brand hubs, how-to guides, 96 of 97 informational articles.
- No public API.
- A machine-readable record of when a page's external citations were last confirmed reachable. The fact index covers pricing and specification claims; it does not yet re-check cited URLs, and we would rather name the gap than publish a date that means less than it looks like it means.
These are tracked as later build phases, not promises with a date attached. This page will be updated to describe them once they're actually live, not before.
Who this is
Need to Know IT is an independent Australian information resource covering NAS, home server, backup and networking topics. No pricing, margin or stock information from any employer or distributor relationship is ever published here.