Site Check in Cursor
Site SEO and landing checks. The MCP server is https://sitecheck.openkrill.app/mcp. No key. The steps below are only for Cursor. A call to check_page_tags was checked against that server on 2026-10-01.
What this server answers
Look at what your website or online store tells crawlers, link previews and AI agents. Ask "can ChatGPT and Claude crawl example.com?", "is this product page set up for search?", "can AI shopping agents read this store?", "are any pages in this store's sitemap broken?" or "what would make this landing page convert better?".
Site Check fetches public pages and files and reports what it finds. For AI crawler access it evaluates robots.txt for fourteen published AI crawler names (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-Extended and others) against the page path, shows the rule that decided each one, and adds the X-Robots-Tag header and the robots meta tag of the page. For page tags it reads the title, description, canonical link, language and icons, plus the Open Graph and Twitter share tags, and lists the common ones that are missing. For redirects it follows the chain hop by hop and shows each status code, the final address and the main response headers.
For online stores, one product page can be checked for its title and description length, h1 heading, canonical link and Product structured data (name, brand, SKU, GTIN, MPN, price, currency, availability, variants). One domain can be checked for llms.txt, agents.md, the Universal Commerce Protocol file that Shopify stores publish, the sitemap and AI crawler rules in robots.txt. A sample of up to 40 addresses from the sitemap can be checked for broken pages and redirects. These work on any public store and add Shopify-specific fields where the store is on Shopify. They are not a speed test.
For landing pages, one page can be checked for the items a conversion review looks at, each with the evidence found and a fix: one clear h1 and a supporting sentence under it, a primary call to action that is near the top and repeated, whether the call-to-action links work, the number of form fields, trust signals (testimonials, customer logos, review markup, guarantees), a visible price or next step, the mobile viewport tag, HTML size and render-blocking hints, and a way to contact the business. The checklist topics follow the conversion-rate-optimization skill in coreyhaines31/marketingskills (MIT licence); the checks and wording are our own. Contact details, review text and page copy are not returned, only counts and the headline.
Only public websites on the standard web ports are checked. Private, local and internal addresses, other schemes such as file or ftp, and redirects into any of those are refused, and pages are fetched with a User-Agent that names Agent Tools. The store checks read robots.txt first and do not request a page that it disallows for Agent Tools. A check makes a handful of requests, and the sitemap sample or the landing page check at most 46.
It shows what the site publishes to a request from Agent Tools servers right now. It does not test whether a firewall or bot filter blocks real crawlers, does not run JavaScript, does not return page text, does not sign in to anything and does not scan for security vulnerabilities. robots.txt is a request that crawlers follow by choice, not access control.
It stores nothing about you or the site, and no fetched page is kept: answers are cached for five minutes, and the only lasting record is a daily count of calls per tool.
What it can do
- Shows whether robots.txt allows or blocks 14 AI crawlers such as GPTBot, ClaudeBot and PerplexityBot
- Shows the rule that decided each crawler, the sitemaps and any noindex or noai directives
- Reads a page's title, description, canonical link, language and icons
- Reads Open Graph and Twitter share tags and lists the common ones that are missing
- Follows the redirect chain with each status code and shows the final address and headers
- Checks one product page's title, description, h1, canonical link and Product structured data (price, GTIN, availability)
- Checks a domain for llms.txt, agents.md, the Shopify UCP file, a sitemap and AI crawler rules in robots.txt
- Checks up to 40 sitemap addresses for broken pages and redirect chains, skipping pages robots.txt disallows
- Checks a landing page's h1, call to action, form fields, trust signals, listed prices, viewport and contact path
- Tests up to five call-to-action links of a page and gives each finding with evidence and a fix
- Refuses private, local and internal addresses and redirects into them
Tools
check_ai_crawler_access, Check AI crawler access. Checks whether 14 AI crawlers may fetch a page. Use it when the user asks whether AI crawlers can read a website or page, for example "can ChatGPT crawl example.com?", "is GPTBot blocked in my robots.txt?" or "which AI bots are allowed on https://example.com/blog?". Pass the page address. Returns, for 14 AI crawlers, whether robots.txt allows the page and the rule that decided it, plus the sitemaps, the redirect chain, and the page's X-Robots-Tag header and robots meta tag. It reads robots.txt only and does not test whether a firewall blocks crawlers. Do not use it for private or local addresses, for security scans, or to collect a site's content.check_page_tags, Check page and share tags. Reads title, description and share tags of a page. Use it when the user asks about the meta tags or link preview of a web page, for example "what will the preview look like when I share this page?", "does my homepage have a meta description and canonical tag?" or "which Open Graph tags is my page missing?". Pass the page address. Returns the title, meta description, canonical link, language, viewport, icons, Open Graph and Twitter tags, and a list of the common tags that are missing. Tags added by JavaScript are not seen. Do not use it to read a page's text or contact details, for private or local addresses, or for rankings and traffic.trace_redirects, Trace redirects and headers. Traces the redirects and headers of a web address. Use it when the user asks where a web address redirects to or why it lands somewhere else, for example "why does http://example.com end up on another address?", "show the redirect chain for this link" or "does this site redirect http to https?". Pass the address. Returns each hop with its status code and target, the final address and status, whether http is upgraded to https, and the main response headers such as cache-control and strict-transport-security. At most 6 redirects are followed. Do not use it for private or local addresses, to read page content, or to scan a site for vulnerabilities.product_page_seo, Check a product page for search. Checks the SEO and Product data of a product page. Use it when the user asks whether one product page of an online store is set up for search, for example "is https://shop.example.com/products/blue-mug set up for search?", "does my product page have Product structured data?" or "why does Google show no price for this product?". Pass the product page address. Returns title and meta description length, h1 headings, canonical link, Open Graph tags, and the Product structured data (JSON-LD): name, brand, SKU, GTIN, MPN, price, currency, availability, variant count and rating summary, plus a list of issues found. On Shopify it adds the product handle and whether a variant parameter is in the address. Works on any store. It reads one page, skips it if robots.txt disallows it for Agent Tools, and does not run JavaScript. Do not use it for page speed, rankings or traffic, to collect prices or catalogs, or for private or local addresses.ai_readiness_check, Check AI readiness of a store. Checks if AI agents can read a store or website. Use it when the user asks whether AI agents and assistants can read an online store or website, for example "can AI shopping agents read my store?", "does shop.example.com have an llms.txt?" or "does my store publish a UCP file?". Pass the domain. Returns whether the domain has /llms.txt, /agents.md, the Universal Commerce Protocol file at /.well-known/ucp (published by Shopify stores; version and listed services are shown) and a /sitemap.xml, plus which of 14 AI crawlers robots.txt blocks for the whole site, and lists present and missing files. It requests at most five files of one domain, skips any that robots.txt disallows for Agent Tools, and shows what is published, not whether an agent can complete a purchase. Do not use it for private or local addresses, to read page content, or to scan for vulnerabilities.link_sample_check, Check a sample of sitemap links. Finds broken or redirected pages in a sitemap sample. Use it when the user asks whether pages of an online store or website are broken or redirected, for example "are any pages of shop.example.com broken?", "check my sitemap for 404s" or "do my product pages redirect?". Pass the domain. Reads the sitemap (for a sitemap index, the product, page, collection and blog sitemaps first) and checks up to 40 addresses spread across it, on that host only, returning counts by status, and for each broken, unreachable or redirected address its status and final address. It is a sample, not every page; it uses HEAD requests at low concurrency, skips addresses that robots.txt disallows for Agent Tools, and makes at most 46 requests. Do not use it for private or local addresses, to read page content, or to monitor a site over time.landing_page_check, Check a landing page for conversion. Checks a landing page for conversion problems. Use it when the user asks what could be improved on a landing page to get more sign-ups, sales or enquiries, for example "review https://example.com/ for conversion", "does my landing page have a clear call to action?" or "why might visitors not sign up on this page?". Pass the page address. Returns twelve checks with a status (pass, warn, fail or info), the evidence found and a fix: the h1 and the supporting sentence under it, the primary call to action and how often it repeats, whether up to five call-to-action links work, form field count, trust signals, a visible price or next step, the mobile viewport tag, HTML size, render-blocking hints and a contact path. It reads one page and tests a few of its links, skips what robots.txt disallows for Agent Tools, does not run JavaScript, and does not return page copy, review text or contact details. It is a checklist, not a prediction of conversion rate, and not a speed or design test. Do not use it for private or local addresses, rankings or traffic.
Add Site Check in Cursor
Cursor reads MCP servers from mcp.json. A project file at .cursor/mcp.json applies to that project. A user file at ~/.cursor/mcp.json applies everywhere. These steps were checked against the Cursor MCP docs on 2026-10-01. Streamable HTTP is the transport: a url, not a shell command.
Create or edit the file so it contains this server. Merge it with servers you already have. Do not replace the whole file if other servers are listed.
{
"mcpServers": {
"site-check": {
"url": "https://sitecheck.openkrill.app/mcp"
}
}
}
No headers and no auth block. Those are for servers that require a token or a pre-registered OAuth client. Site Check answers without either. Saving the file is the install. Open Cursor Settings, then MCP, and confirm Site Check is enabled. If it stays disconnected, reload the window. You should see check_ai_crawler_access, check_page_tags, trace_redirects, product_page_seo, ai_readiness_check, link_sample_check, landing_page_check.
In Agent chat, ask: Can GPTBot and ClaudeBot crawl https://example.com/blog? Cursor calls the tool over streamable HTTP and shows the arguments before it runs them, depending on your run mode. Approve the call the first time and read the result. The server is public and read-only. It does not see your repository unless the model puts repository text into the arguments, so do not paste secrets into the question.
A one-click Marketplace install does not exist for this server. The JSON above is the whole setup. Team admins can distribute the same url through a team marketplace, but that is an admin step, not something this page can do for you. OAuth redirect URLs in the Cursor docs matter only for servers that require sign-in. Skip them here.
A call checked on 2026-10-01
example.com is a public page. The tool reads the head only. The request below was posted to https://sitecheck.openkrill.app/mcp as tools/call. HTTP 200. Source: the Site Check server, read 2026-10-01.
{
"method": "tools/call",
"params": {
"name": "check_page_tags",
"arguments": {
"url": "https://example.com"
}
}
}
- status. 200
- source. The website itself, requested from Agent Tools servers
- notice. This shows what the site publishes to a request from Agent Tools servers at this moment. Firewalls and CDNs can treat real crawlers or visitors differently, and robots.txt is a request that crawlers follow by choice, not
- url. https://example.com/
- finalUrl. https://example.com/
Repeat the call yourself if you need a newer reading. Cached answers expire. A rate limit is not a result: wait and try again. Nothing in the call is a ranking, a filing, or advice.