People reach your site and still do not act. Sometimes the visitor between them and your page is now software: an assistant searching for an answer or an agent trying to complete a form. This guide shows what such systems can read, what they may miss and which checks belong before AI-specific extras.
An agent may arrive through search or the page itself
Some agents do not visit a website directly. A Copilot Studio agent that uses a public website as knowledge relies on Grounding with Bing Search, so its answers come from content indexed by Bing. Pages behind sign-in or absent from that index cannot serve as that source (Microsoft on public website knowledge, checked: 2026-10-05).
Google gives similar advice for its AI search features. A page needs no special AI markup for AI Overviews or AI Mode. It must be indexed and eligible to show a snippet, while ordinary controls such as noindex, nosnippet, data-nosnippet and max-snippet limit what can appear (Google AI features guidance, checked: 2026-10-05).
Browser agents can take another route. Google describes them as systems acting for people, for example to compare products or make a booking. They may interpret a screenshot, the document object model and the accessibility tree, which is the semantic view also used by assistive technology (Google AI optimisation guide, checked: 2026-10-05).
This means visibility is not one switch. Search-grounded assistants need indexable pages. Browser agents need understandable controls. A person still needs an honest offer and a form that works.
Ordinary web foundations carry most of the load
A JavaScript site is not automatically empty to AI. Google says it generally processes JavaScript-generated content when that content is not blocked, although JavaScript SEO is more complex and semantic HTML remains recommended (Google AI optimisation guide, checked: 2026-10-05).
The narrower risk is content that appears only after an interaction, a failed script or a sign-in. A price written as zero in HTML and replaced by animation later may expose the wrong state to a simple reader. A product description locked inside a download form may never reach a search index.
The same foundations guide our work on a website that people and machines can read: meaningful HTML, visible text, stable addresses, labelled controls and a clear route from an offer to an enquiry. None guarantees a mention in an AI answer. They make the actual page understandable.
Check the route from question to action
Read the page without interaction
Open the home page, service page and contact page with JavaScript disabled. The test shows content with no usable initial state; it does not predict how Googlebot renders the page.
Can you still read the service name, scope, price if one is published, contact details and the purpose of every field? Then reload normally and compare. Record any fact that appears only after clicking, scrolling or accepting optional tracking.
Check whether an important answer lives only in a PDF behind a form or inside a client portal. Microsoft says signed-in and non-indexed URLs do not work as public website knowledge for a Copilot Studio agent (Microsoft on public website knowledge, checked: 2026-10-05).
Inspect controls and labels
Try the form with a keyboard. Every interactive element should receive focus in a sensible order. A visible label should be associated with each input, and links and buttons should use their native HTML elements.
Google's web.dev guidance says browser agents benefit from semantic <button> and <a> elements, correctly associated <label> elements and stable layouts. Continuous shifts can confuse systems working from screenshots, while transparent overlays can block the element an agent tries to use (web.dev agent UX guidance, checked: 2026-10-05).
These checks also help people who use keyboards or screen readers. The shared benefit is practical: the control announces what it is and behaves as expected.
Review robots, sitemap and feeds
Open /robots.txt and the sitemap it declares. The Robots Exclusion Protocol is standardised in RFC 9309, but the specification explicitly says its rules are not access authorisation. Private content needs real access control (RFC 9309, checked: 2026-10-05).
A sitemap gives crawlers a list of canonical locations. Its required fields are urlset, url and loc; lastmod, changefreq and priority are optional. The protocol also allows the sitemap address in robots.txt (Sitemaps protocol, checked: 2026-10-05).
If the site publishes articles or updates, check its RSS feed too. RSS 2.0 is an XML syndication format whose channel requires a title, link and description (RSS 2.0 specification, checked: 2026-10-05). Framewright publishes blog feeds at /blog/feed.xml and /pl/blog/feed.xml, not at the site root.
Ask the questions a customer would ask
Ask two assistants what the company does, who the service is for and how someone starts. Compare each answer with the site. This is a spot check, not a ranking test.
When the answer is wrong, trace the source. Is the page missing, blocked, stale or redirected? Microsoft documents that a public knowledge URL redirecting to another top-level domain can fail to return content, and the configured path can contain no more than two path levels (Microsoft on public website knowledge, checked: 2026-10-05).
Keep the review with the person who already monitors platform changes. The same ownership model appears in the process for reading change notices: a feed or dashboard matters only when someone compares it with the business.
Treat AI-specific files as optional layers
llms.txt is a proposal published by Jeremy Howard, not a web standard. Its format starts with an H1 and can add a summary plus sections of links. Google states that Search does not use llms.txt and that the file neither helps nor harms Google Search visibility (llms.txt proposal, checked: 2026-10-05; Google AI optimisation guide, checked: 2026-10-05).
NLWeb goes further. It is an open project announced by Microsoft in May 2025, designed around Schema.org and RSS so a site can expose a natural-language interface. Its reference implementation provides /ask and /mcp APIs, and every instance acts as a Model Context Protocol server (Microsoft introduction to NLWeb, checked: 2026-10-05; NLWeb repository, checked: 2026-10-05).
The current repository calls the code proof-of-concept demonstrations and lacks a bundled production CI/CD pipeline. Treat it as an active experimental project, not a supported shortcut to AI visibility (NLWeb repository, checked: 2026-10-05).
What not to do
Do not pay for a promise of guaranteed placement in AI answers or plan around one. Google says ordinary indexing and snippet eligibility apply, and no published guidance promises inclusion.
Do not rebuild a page merely because it uses JavaScript. Google usually renders JavaScript. Test the initial state and interactive dependencies instead.
Do not block robots "just in case." Google-Extended controls use of fetched content for future Gemini training and grounding in Gemini apps. It does not control Google Search or act as a ranking signal (Google crawler documentation, checked: 2026-10-05).
Do not add llms.txt, schema markup or NLWeb before fixing missing service text and broken controls. Google says structured data is not required for generative search, although it can support eligible rich results (Google AI optimisation guide, checked: 2026-10-05).
When you can handle this yourself
You can run the first pass internally. Compare the site with JavaScript off, inspect labels, open robots.txt and the sitemap, and find the RSS feed. Ask two assistants the three customer questions. Write down each mismatch and fix the page that owns it.
Bring in help when content is rendered only after complex client-side actions, a domain migration has left cross-domain redirects, or the form lacks semantic controls. NLWeb and MCP endpoints are software projects with deployment and security decisions, not content settings.
The useful standard is uncomplicated. Publish accurate information at stable addresses, make the controls understandable and keep discovery files current. An agent can then work from the same evidence a person and a search engine already need.