polite.aiDocs
BUILD / Knowledge

Website knowledge

Point us at a site and its pages become answers your agents can give — crawled, summarised, and kept fresh automatically.

Your site is usually already there

The website address you gave at signup is crawled automatically the moment your account is approved — no action needed. It becomes your organisation's first knowledge base, and it also gives the Builder a brief on your business — what you do, how you describe it, your real contact details — so new agents sound like you from the first draft.

That happens because you ticked the crawl-consent box on the waitlist form. If you didn't, nothing was crawled: the base shows Consent withheld and stays empty until you add the site yourself.

Add another site

Open the form

Knowledge → Add knowledge base, then the Website tab. A bare domain is fine — example.co.uk works, and we add https:// for you.

Confirm you're authorised

Tick I'm authorised to have this site crawled — you own the site, or the owner has said yes. The button stays disabled until you do.

Add website

The crawl starts straight away, and you can watch it progress on the list — see Knowledge bases.

One base per site: adding a host you already have is politely refused. How many bases your plan includes is in Plans & limits.

What the crawler reads — and skips

The crawler reads the text of your pages — cleaned, summarised and indexed so agents can search it — and picks out contact details like emails and phone numbers as it goes. On public sites it honours robots.txt, so pages you've told crawlers to leave alone stay unread. Crawls are capped at a sensible size, and each base holds up to 100 MB of extracted text — plenty for almost any business site.

Password-protected sites

A knowledge base can sit behind a login — a staff area or an intranet. Tick This site needs a login when adding the site, choose the login type — Username & password form or HTTP Basic auth — and enter the credentials. For form logins you can also give the Login page address, if the form isn't on the site's front page.

Good to know

The password is stored encrypted and used only by the crawler — it's never shown anywhere in the interface again. The base simply carries a Private badge. And because it's your own consented content, authenticated crawls don't apply robots.txt.

If the login fails, the crawl fails with a check-your-credentials error on the row — fix the details and try again.

Keeping it fresh

Website bases are re-crawled automatically on a schedule, so day-to-day drift takes care of itself. Just updated your site and want agents answering from the new content now? Press Re-crawl on the row — a fresh crawl queues immediately, with live progress as it runs.

The crawl consent you give is revocable at any time. To withdraw it, delete the base — Delete, then Confirm delete, on the row. Everything we stored from the crawl — pages, summaries, the search index — is purged permanently, and we won't crawl the site again unless you re-add it.

Last updated 2026-07-30