Train your chatbot on your website content. Boei discovers pages on your site, lets you select which ones to train, and extracts the useful content automatically.
How to get there: Go to Setup → Chatbot in the top menu → click your chatbot → Training → Website tab.
Click Add New Pages to see the available discovery methods:
Boei's crawler visits your website and follows links to discover pages, similar to how a search engine works.
The crawler discovers pages by following links from your starting URL. Pages behind login walls or not linked from anywhere won't be found.
If your crawl fails, it's usually because a firewall is blocking the crawler. See Crawl errors for troubleshooting steps.
If your website has a sitemap (most do), you can import pages directly from it.
When you add a URL, Boei automatically looks for your sitemap, so you rarely need to paste the sitemap URL by hand.
Extract and import all links found on a specific page. Useful for sites that have an HTML sitemap or a links/resources page.
Paste a list of URLs directly.
| Site Crawl | Sitemap Import | Page Links | Bulk Upload | |
|---|---|---|---|---|
| Best for | Quick start, small sites | Larger sites, ongoing accuracy | HTML sitemaps, resource pages | Known URL lists |
| Discovery | Follows links automatically | Reads sitemap.xml | Extracts links from one page | Manual URL list |
| Auto-sync | Not supported | Supported | Not supported | Not supported |
| New pages | Only on manual re-crawl | Auto-detected if auto-sync is on | Manual re-import | Manual |
| Removed pages | Manual cleanup | Auto-cleaned on sync | Manual | Manual |
Recommendation: Use Sitemap Import if your site has a sitemap, especially for sites that change regularly. Pair it with auto-sync (Growth plan and above) to keep your chatbot current without any manual work. Use Site Crawl for a quick start if you're not sure about your sitemap.
Boei automatically strips common boilerplate elements when crawling pages: navigation menus, footers, sidebars, cookie banners, comment sections, ads, and similar non-content elements. This happens by default so your training data stays clean without any setup.
For advanced control, you can set your own CSS include/exclude selectors per website source to further refine what content gets extracted. See CSS filters for details.
Each page shows a status indicator:
Every 5,000 characters of cleaned page content counts as 1 page credit (minimum 1 per page). Each page shows its credit cost in the list. See Training credits for plan limits and details.
Long crawls resume from where they left off if interrupted, so a large site doesn't have to restart from zero after a timeout.