GatPilotBot
The GatPilot program that reads a website's pages when a company adds them to its assistant's knowledge.
Last updated: 8 October 2026 · Version 1
In short
GatPilotBot, in five lines
The summary helps you find what matters; the full text below is the one that applies.
It reads a website's public pages only when a company that is a GatPilot customer asks for it from the dashboard.
It identifies itself as GatPilotBot and reads robots.txt at the start of every crawl.
It sends at most 3 requests at a time, then pauses for at least 300 ms, or for as long as Crawl-delay asks.
You stop it with two lines in robots.txt: User-agent: GatPilotBot and Disallow: /.
The text it reads helps a single assistant answer. We do not train AI models on it.
01 · What it is
What GatPilotBot is
The GatPilot crawler: it reads a website's pages for a company's AI assistant.
GatPilot is a service operated by TOPFED WEB S.R.L., a limited liability company registered in the Republic of Moldova under IDNO 1026023125543. The company's registered office is in Chișinău, Republic of Moldova.
GatPilotBot is the GatPilot crawler, that is, the program that automatically reads a website's public pages.
It reads a website to turn its text into knowledge for the AI assistant of a single company that is a GatPilot customer.
It identifies itself to the server with this name, called the user agent:
GatPilotBot/1.0 (+https://gatpilot.com/bot)From a page it keeps:
- the title and the main text, without menus, header, footer, forms and scripts;
- the page's address and the date it was read.
What then happens to the text:
- it becomes a source in the knowledge of that company's assistant, which answers from it;
- it is split into chunks, that is, short pieces of text, usually a paragraph or two;
- the chunks go through OpenAI, which turns them into series of numbers, so that the assistant can search them by meaning;
- the company sees, in the dashboard, the text that was read and can delete any page;
- other companies do not see it, and we do not use it to train AI models.
02 · When it comes
When it reads your website
Only when a company sends it, from its dashboard.
GatPilotBot comes only when someone from a customer company, the owner or an administrator, enters a website's address in the dashboard and presses “Read the site”.
The bot does not check whether the company owns the website. The Terms of service require it to add only content it has the right to use.
It has no re-crawl schedule. It comes back only when the company starts the crawl again.
If a crawl fails with an error, it is retried automatically at most twice. If our server restarts during a crawl, the crawl starts again from the beginning.
For one assistant it reads at most 20 pages in total on the Free plan, 200 on Pro and 2,000 on Custom. Pages read before count too.
03 · Behaviour
How it reads
What it reads first, what it skips and how fast it goes.
At the start of every crawl, the bot downloads /robots.txt from the root of the website that was added. It is the file in which the owner says what robots may read.
- It uses the User-agent: GatPilotBot group. If that group is missing, it uses the User-agent: * group.
- It reads Disallow and Allow rules as the start of a path. The longest rule wins, and in a tie Allow wins.
- The characters * and $ have no special meaning for it, so write rules as the start of a path.
- Rules apply only to the path, not to the parameters after the question mark.
- It takes Crawl-delay into account for its pace, as described below.
- If robots.txt is missing, does not respond within 10 seconds or returns an error, the bot reads as if there were no rules.
It first takes the addresses from the sitemap: the one named in robots.txt or, if there is none, /sitemap.xml.
- It downloads the sitemaps at the start, one after another, without a pause: as a rule, at most about 50 files. It opens no further sitemaps after 5,000 addresses, but a single large sitemap can yield more.
- If the sitemap yields no address, it follows the links in the pages, at most 3 levels from the first page.
- It takes only addresses on the same domain, with or without www, and checks each one against robots.txt.
- If a page sends it elsewhere (a redirect), the bot follows at most 4 redirects, even to another domain.
- It skips the cart and checkout, account and sign-in, search and filters, feeds, tags, categories and administration.
- By default it also skips product pages. The company can include them with the “Also read product pages” option.
- It skips files: images, PDFs, documents, archives, style sheets and scripts.
- It keeps only pages with at least 200 characters of text.
It sends at most 3 requests at a time.
- After each batch it pauses for at least 300 ms, or for as long as Crawl-delay asks, if that is longer.
- It waits at most 15 seconds for each request and does not keep pages larger than 5 MB.
- It makes only read requests (GET), without cookies and without authentication. It does not run JavaScript and does not submit forms.
- It does not open addresses on private networks, nor ports other than 80 and 443.
Good to know
With Crawl-delay: 10, the bot sends at most 3 requests every 10 seconds, not one.
04 · Stopping it
How to stop it
Two lines in robots.txt are enough.
To stop GatPilotBot on the whole website, write in robots.txt:
User-agent: GatPilotBot
Disallow: /To stop it on only part of the website, write that part's path:
User-agent: GatPilotBot
Disallow: /internal/To slow it down, ask for a longer pause, in seconds:
User-agent: GatPilotBot
Crawl-delay: 10Write the name without the version: a group for “GatPilotBot/1.0” does not apply. A rule takes effect from the next crawl; a crawl already under way uses the robots.txt downloaded at its start.
Even with Disallow: /, the bot still downloads robots.txt and the sitemap at the start. It does not, however, open any page.
Pages read before stay in the assistant's knowledge until the company deletes them. A new crawl updates them only as far as the plan's limit still allows: at the limit, it reads nothing. It does not delete pages that have disappeared from the website.
If you want your pages that have already been read to be removed, write to us at info@gatpilot.com, with the domain and the addresses.
05 · Contact
Contact for website owners
Write to us if GatPilotBot is putting load on your server or reading what it should not.
Write to us at info@gatpilot.com. Send the domain, the date and time of the requests, with the time zone, and what happened.
A few lines from the server log help us find the crawl.
We reply within one business day, in Romanian or Russian.
For a wrong or harmful answer from a GatPilot assistant, see the Legal notice.
Version history
- : first publication.
Documents
All GatPilot legal documents
The same company and the same rules on gatpilot.md and gatpilot.com.
Legal notice
Who operates GatPilot, how to contact us and where to turn.
Terms of service
The contract for GatPilot: account, plans and payment, your AI key, your content, liability and the applicable law.
Privacy Policy
What data we process, why, for how long and what rights you have.
Data Processing Agreement
How we process your customers' and employees' data on your behalf.
Sub-processors
The providers that help us run GatPilot and the data each one receives.
Cookies and local storage
What GatPilot keeps in your browser, why and for how long.
GatPilotBot
What our crawler reads and how to stop it.
Have a question about this document? Write to us at info@gatpilot.com.