Get a free website with any plan

See how
INTERNET ESSENTIALS

robots.txt best practices

Last updated

IN SHORT

A robots.txt file is a plain text file at your domain root that tells search crawlers which paths they may request. For any Flashcloud site, keep rules minimal: allow everything by default, disallow only admin or internal paths, and publish your sitemap URL. Incorrect syntax can accidentally deindex your entire site.

robots.txt is a plain text file at the root of your domain (yourdomain.com/robots.txt) that tells search engine and other crawlers which paths they may request. The fix for most problems: keep it short, allow everything by default, disallow only admin and internal paths, and publish a Sitemap: line. Get the syntax wrong and you can accidentally deindex your entire site, so treat this file with the same care as a DNS record.

The file crawlers actually respect

robots.txt is a request, not a lock. Well-behaved crawlers (Googlebot, Bingbot) honor it. It does not stop a page from being accessed directly, and it does not remove a page that's already indexed. If you need to hide something from prying eyes, put it behind a login or use Directory Privacy in cPanel, not robots.txt.

A minimal, safe file for most sites looks like this:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yourdomain.com/sitemap.xml

User-agent: * applies the rules to every crawler. Disallow lines block a path prefix. Allow lines carve out exceptions inside a blocked path, useful for the WordPress AJAX endpoint that themes and plugins rely on even though it lives under /wp-admin/.

Common mistakes that tank SEO

The single most damaging mistake is Disallow: / left in place after a site goes live. That one line blocks crawling of the entire domain. It's a default in a lot of staging setups and page builders, so check it the moment you point a domain at production.

Other mistakes worth checking for:

  • Blocking CSS and JS. Google renders pages to judge layout and mobile-friendliness. If you disallow /wp-content/themes/ or similar, you can block the assets it needs to render the page correctly.
  • Confusing noindex with Disallow. Disallowing a page in robots.txt stops crawling, but if other sites link to it, Google can still index the URL with no description. If the goal is "don't show this in search results," use a noindex meta tag on the page instead, and don't disallow it, since a crawler has to fetch the page to see the noindex tag.
  • Case sensitivity. Paths in Disallow are case-sensitive. /Private/ and /private/ are different rules.
  • One rule per line. You can't combine paths with commas or wildcards beyond the basic * and $ that major crawlers support.

Wildcards and pattern matching

* matches any sequence of characters, and $ anchors the end of a path. This lets you target file types without listing every file:

User-agent: *
Disallow: /*.pdf$
Disallow: /search?*

The first line blocks any URL ending in .pdf. The second blocks any URL containing a query string starting with search?, useful for keeping internal search-result pages out of the index. Test these patterns before shipping them; a stray wildcard can match far more than intended.

AI crawlers need their own decision

Search crawlers, AI assistant crawlers, and AI training crawlers are three different categories now, and you don't have to treat them the same. Because Flashcloud domains are proxied through Cloudflare by default, the CDN page in the portal has a dedicated AI Crawlers panel with separate controls for search crawlers, AI assistants, and model training, each with Allow, Block, or Block on ad pages, plus a "robots.txt" option for training crawlers that asks nicely rather than firewalling. That's a control point at the CDN level, separate from hand-editing robots.txt.

If you're managing this by hand instead, general web knowledge (not Flashcloud-specific) identifies user-agents like GPTBot and CCBot as commonly associated with AI training crawlers:

User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

Leave User-agent: * untouched if you still want normal search indexing.

Where to put it and how to verify it

The file must live at the domain root, not in a subfolder, and it must be named exactly robots.txt in lowercase. Upload it via File Manager or an FTP client to your site's document root (the same folder index.php or index.html lives in). After uploading, load yourdomain.com/robots.txt <p>If you're publishing content through Content Studio and want the underlying mechanics of getting an article live and indexed, <a href="https://portal.flashcloud.com/knowledgebase/article/writing-your-first-article">writing your first article</a> covers the publish flow end to end.</p> <h2>When to get help</h2> <p>If you changed <code>robots.txt and traffic or indexing drops afterward, or you're not sure whether a rule is blocking something it shouldn't, open a ticket from Support in the portal. It goes to a real person, and they can help diagnose what's being blocked.

Common questions

Can I use robots.txt to hide private pages?

No, robots.txt is only a crawler request and not a security barrier. It cannot prevent direct access to URLs or remove already indexed pages. Use cPanel Directory Privacy or a login system to protect private data.

Why did my entire website drop out of search results?

You likely have a Disallow: / rule in your robots.txt file. This single line blocks crawlers from your entire domain and often gets left behind by staging environments. Check the file at your domain root and remove that rule.

How do I stop AI crawlers from scraping my content?

Manage them through the AI Crawlers panel on your Flashcloud CDN page, or add custom Disallow rules for user-agents like GPTBot and CCBot. The CDN settings let you allow, block, or set robots.txt rules without editing files manually.

Where do I upload my robots.txt file?

Upload it to your primary document root folder alongside index.php or index.html using File Manager or an FTP client. It must be named robots.txt in lowercase letters to function.

CAN'T FIND IT?

Real humans answer fast.

Hosting with us? Open a ticket and a real person replies - no scripts, no upsells. Still choosing a host? The same team is included with every plan, from day one.