v2.0.33
2.0.33 (2026-08-02)
Bug Fixes
-
indexing: stop a partial crawl replacing a complete index (ecbd90f)
A crawl that lost a couple of pages threw away every page it had
successfully crawled, so a 124-page site indexed nothing. Tolerating
those failures on its own was unsafe, because addDocument replaces
rather than merges - a partial crawl would have quietly overwritten a
complete index and reported success.Crawler:
- Tolerate a small number of failed pages instead of failing the crawl
- Never tolerate a loss of half the crawl or more, however small it is
- Always fail when the root page cannot be loaded, with its own message
Indexing:
- Refuse a reindex that would replace a document with one less than half
its size, counting pages rather than chunks so the check does not move
when chunk sizing changes - Skip that check only when the path prefix genuinely narrowed, since
widening or clearing it should return more pages, not fewer - Keep indexing when the stored page count cannot be read, so a storage
hiccup cannot discard a finished crawl
Storage:
- Add countDocumentPages, counting distinct published pages for a URL
Normalise a trailing slash off pathPrefix so '/docs' and '/docs/' are one
spelling, matching what the crawl already does.