Releases: HarperFast/prerender-plugin
Release list
prerender-v0.98.1 — @harperfast/prerender
@harperfast/prerender 0.98.1
PageCheck stores a digest of the observation, not the observation (#234).
- A row is written for nearly every URL the nightly pass probes while
changeProbe.serveCheckis enabled (dry run included), and for every serve-time check, and it replicates to every node. The observation, ~1 KB on a product endpoint, was most of each row. - The row now carries
observedDigest: 96 bits of SHA-256 over the observation's UTF-16 code units, 16 characters. A row drops from ~1 KB to ~228 bytes. - The sweep compares
observedDigestwith the digest of its own baseline. It skips exactly the rows it skipped before.
Schema: PageCheck.signature is replaced by PageCheck.observedDigest. This is safe whether or not 0.98.0 rows exist: an old row's digest reads as null, so it spares no probe, and it is replaced on its next write. Deploy 0.98.1 rather than 0.98.0.
prerender-v0.98.0 — @harperfast/prerender
@harperfast/prerender 0.98.0
Serve-time checks: a cached page is checked against the origin when a bot asks for it (#232). Everything here is off by default; changeProbe.serveCheck ships in dry run.
changeProbe.serveCheck: a page served from cache that has not been checked or rendered since the last anchor (or, for thesampleof URLs, withinmaxAge) is served exactly as before and checked in the background. The check compares the rule endpoint's mapped fields with the facts of the page that was just served, or, on adocumentCheckroute, the origin's own document head. A disagreement expires and re-files the page like a detected change. The same disagreement again, on a page rendered after it, with the origin unchanged, isheldrather than re-rendered. Dry run countsserve_checkwould-checkand asks nothing.PageCheck(new table,verificationdatabase, replicated): every decided check, so a URL is asked once per window cluster-wide. The nightly pass skips a URL only when a check since the pass began observed exactly its baseline.- Origin budget: checks share the node's
ratePerSecondwith the pass, which leaves them what they use (at mostserveCheck.share). Origin pushback pauses them node-wide. util/documentFacts: reads the renderer's page facts off HTML bytes (head only, exact or null), held to Chrome parity by tests.- Raw cache:
ingress.routes[].rawFactsstores page facts with a raw document (new nullableRawPage.facts), and a raw hit is checked from them.render.raw.ignoreNoStorestores despite the origin'sno-store. pageCheck.fields[].pathPattern: a mapped field compares only on URLs whose path matches; elsewhere it makes no claim.- Metric:
prerender_opsserve_check(outcome by source).
Schema: the new PageCheck table and a nullable RawPage.facts column, both additive. Deploy console 0.23.0 for the Serve-time checks card.
prerender-v0.97.3 — @harperfast/prerender
@harperfast/prerender 0.97.3
A URL too long to be a cache key is proxied to the origin, not answered 500 (#230).
- Harper refuses a primary key over 1978 encoded bytes, and the first table read on the bot serve path threw on it, so the crawler got a 500 for a URL the origin serves.
CacheKey.fitsKeyLimit(url)is checked where a URL enters: a bot request over the limit is proxied and nothing else (bot_servebypass,origin_fetchreasonkey-too-long, one warn); a sitemap<loc>over it is invalid; a render redirecting to one discards the render and keeps the target (a 301 retires the source and adopts nothing,unkeyable-destination); the admin explain and revalidate answer without reading.
No schema or config change: a worker restart loads it.
prerender-console-v0.24.0 — @harperfast/prerender-console
@harperfast/prerender-console 0.24.0
- Read cache on the console (#236). The console now answers repeat reads itself instead of passing every one to the prerender nodes, which also serve crawlers. Answers are stored in a node-local table (
prerender_console.ProxyRead) that every worker on the console host shares.- Retention:
analyticsandcrawl-breadth60s,metrics10 min,pages15s, everything else 5s.sessionandpage-contentare never cached. - Who sees a cached answer: only an operator whose own session that node has confirmed as a super_user in the last 60s.
- After a write: every POST that can change state clears the cache before it answers, so the reload after an action reads the node.
- New field: each entry in
sources.nodes[]now carriesageMs. - New option:
cache: falseturns the cache off.
- Retention:
- Deploy note: the component now has a schema, the
prerender_consoledatabase. The table is created on first load; there is nothing to migrate.
prerender-console-v0.23.0 — @harperfast/prerender-console
@harperfast/prerender-console 0.23.0
- Serve-time checks card on the Probe view (plugin 0.98.0,
prerender_opsserve_check): would-check or queued, agreed, mismatched, raw documents deleted, held, inconclusive, failed, and the checks that asked nothing, with mismatches by source (#232). - The Traffic view describes
bypassas including a URL too long to be a cache key (plugin 0.97.3, #230).
prerender-v0.97.2 — @harperfast/prerender
@harperfast/prerender 0.97.2
An uncompressed origin body is gzipped before it is stored, in the raw cache and the negative cache (#229).
- A body the origin sent with no
content-encoding(oridentity) is gzipped at level 6 on the detached store path and stored undercontent-encoding: gzip. A body the origin already encoded is stored exactly as before. - The serve path is unchanged: it re-encodes from the stored header, so gzip clients get the stored bytes and others a decode.
maxBytesstill counts bytes as received. It bounds the capture's memory, and for these bodies that memory is uncompressed. Sizerender.negative.maxBytesfor an uncompressed error page: its default, 1 MB, fits the pages that prompted this.
Why: one production origin gzips its pages but sends every 404 uncompressed, whatever Accept-Encoding asks for (16 of 16 sampled, 79-699 KB). Stored as sent, those bodies took 7-13x the storage, and answering from storage would re-encode each one for the crawler. Measured: gzip takes 1.7-2.4 ms of CPU per 348-699 KB body and produces 47-54 KB, once per stored body.
No schema change, no config change. Rows stored before this release keep serving as they are.
1,823/1,823 plugin tests. Console unchanged (0.22.0).
prerender-v0.97.1 — @harperfast/prerender
@harperfast/prerender 0.97.1
A snapshot's own validators are opt-in: page.snapshotValidators, default false. (#227)
- Default (
false): a rendered page (a cache serve, a peer rescue, a render-now result) carries noETagand noLast-Modified, so a conditional request never gets a 304 from a snapshot. true: 0.97.0's behaviour. The snapshot carriesETag: W/"<lastCachedMs>-<deviceType>"andLast-Modified= its render time, and a crawler holding the current render gets a 304 from the cache.- Unchanged from 0.97.0: a snapshot never carries the origin document's
ETag/Last-Modified. Raw-cache documents and proxied responses keep the origin's validators.
Why: 0.97.0 gave every snapshot validators of its own. On a site whose origin sends none, that is the first time crawlers get validators from this cache, so they start revalidating and it starts answering 304s. That is a change to what crawlers receive, to make and measure per site (response_304 beside bot_serve), not a side effect of an upgrade.
Upgrading from 0.97.0: set page.snapshotValidators: true to keep 0.97.0's behaviour. It is a live option; no restart needed.
1,816/1,816 plugin tests. Console unchanged (0.22.0).
prerender-v0.97.0 — @harperfast/prerender
Fixes from a full review of the plugin, reviewed again as one branch in several rounds (#226). The full list, with every option and metric, is in the PR.
- Sitemap walk safety.
- A failed child no longer turns shifted URLs into departures. It holds back only the URLs on the routes it last held.
- A truncated document counts as a failed child.
sitemap.shrinkGuardrefuses a much shorter document, or a dropped child that still holds URLs listed nowhere else. A refusal is accepted only after a new origin version (newVersionAfter) oracceptAge.- One lossless departure budget per walk,
sitemap.departure.maxPerWalk(default 5000). It is spent only on departures that get an action and shared round-robin across children, and the rest are deferred to later walks. - The bypass token goes only to the root's host or
domains, only over https, and redirects are followed manually. - Every hold-back is listed on the run row (
holdBack).
- Change probe.
- A pass skips only rows probed since it began.
reprobeAfteris retired. - The anchored pass is chained, caught up at boot, and resumed after a throw.
- The reseed after a canary trip acts on changes, except covered or healed pages.
- The render check (
renderCheck, on) confirms a suspect render by one re-probe drawn from the sweep's own origin budget. - Per-batch counters, plus
probe_errors/caught_up/ignored/covered/anchor/detection_lag/render_mismatch.
- A pass skips only rows probed since it began.
- Render queue.
- The slow-lane retry waits the effective cadence.
- "Render this now" asks get a head start (
urgentAt, 0.5) below detected changes. fileDueNowoff the owner forwards to the owner.- A result from a lease older than a change mark doesn't store the old page. Grant times are exact and releases are generation-checked.
render change_lag_msandrender_size.
- Serve path.
- Proxied 3xx keep
Location. - Pages found but not served are proxied without validators, and cached pages carry a weak ETag.
- Peer rescue checks the owner copy's serve status.
- HEAD goes upstream as HEAD.
- Invalidation state is held per worker from a subscription.
- Proxied bodies go through
originBodyStream, becauseReadable.toWeb's cancel can throw.
- Proxied 3xx keep
- Admin and config.
POST /prerender_admin/revalidateworks again (broken since 0.94.0).- Override rows report
masking/redundant, and a set that would newly drop routes is refused. - Timer and read limits have maxima, and the analytics scan streams.
Deploying
- Worker restart required. New attributes on
RenderSchedule,ProbeState,ConfigOverrideandSitemap. The lease buffer is renamed, so expect one duplicate-render burst. - Don't write config overrides during the rolling restart.
- Check stored child sitemaps are https on a host in
domains. - Departures are metered at 5000 per walk even with
maxActions: -1. SetmaxPerWalk: -1for the old behaviour. - Don't raise
page.swrTtlyet: the probe's hard expiry is still relative to it.
prerender-console-v0.22.0 — @harperfast/prerender-console
Panels for every series plugin 0.97.0 adds (#226):
- Change probe → Change safety:
- action errors, caught up, ignored and covered changes;
- anchored-pass outcomes, with a skipped anchor flagged;
- detection lag (median and ≈p95, per rule);
- render-check outcomes.
- Queue:
- Change to cache: change mark to re-render, per route.
- Page size: share over 1 MB, by route and device.
- Filing: off-owner forwarding outcomes and unschedulable rows.
- Sitemaps → Departures:
- the departure family, grouped by what it means for a page: acted, shear, held back, deferred, skipped, lost;
- shrink refusals and acceptances;
- the live hold-back list.
- Also:
- the probe view describes per-batch counters;
- sitemap detail pages over the stored entries.
Against an older plugin, the new cards show as empty.
prerender-v0.96.0 — @harperfast/prerender
A negative cache for origin 404/410s, and gone targets reopened when the origin answers 200 again (#225).
- The negative cache (
render.negative, route opt-innegativeCache: true). On a true miss, the origin's own 404/410 is stored, and the next crawler asking for the same dead URL is answered from it.- Three windows, measured from the last time the origin confirmed the status (
checkedAt):- inside
freshMs: answered from storage, and the origin is not asked. This is the only part that saves origin requests; - inside
lifeMs: answered at once while a background re-check asks the origin. A 404/410 confirms the entry, a 200 or a redirect drops it, and a 5xx or timeout keeps it serving (stale-if-error); - past
lifeMs: proxied, and a 404 stores again.
- inside
- The stored body ages out too. Once the bytes are older than
lifeMs, the next check that goes to the origin anyway replaces them (the re-check is a GET instead of a HEAD). No body is served more than about twicelifeMsold, at no extra origin requests. - It never answers a URL whose Target a sitemap lists, checked on every read and before every store (
skipTargets: anyalso skips unlisted Targets). It never answers a bot inexcludeBots. A bulk-invalidation epoch refuses entries confirmed before it, for every bot. - Store refusals are counted by name: has-cookie, private, no-store (
ignoreNoStorefor an origin that sends it on every document), staging, oversize, empty.assumeSharedworks as it does for the raw cache. - Re-checks are single-flight per key and capped per worker (
maxConcurrentChecks); captures are capped like the raw cache's. - Counted honestly: a fresh answer is
bot_servesourcenegative. An answer whose background re-check went to the origin is sourceorigin, cacheStatusnegative-revalidate, one perorigin_fetchreasonrevalidate. - Off by default.
dryRun(default on) stores and re-checks but never answers from storage, and countswould-serve(the saving) andwould-serve-live(the risk: an armed cache would have answered a 404 while the origin said 200).negative_gapis the crawlers' re-ask curve to choosefreshMsfrom.
- Three windows, measured from the last time the origin confirmed the status (
- Reopening gone targets (
render.suppression.gone.reopen, on and dry-run by default). An origin 200 for a gone-suppressed target files its recheck due now, and the render's own verdict flips it.- The 200 can come from a proxied bot request (any bot, even one the discovery gate refuses), a negative-cache re-check, or the target rejoining its sitemap (the arrival check used to skip every suppressed row).
- Gone verdicts only: a noindex or canonical-mismatch page answers 200 by definition. Bounded by
dedupeMsper URL (6h) andmaxPerMinuteper worker (60); a failed filing is retried by the next 200.
- Both outcomes of a suppressed target's render are counted, with the reason and the age since the last verdict:
suppression_liftedandsuppression_held. A gone target's own recheck lands in14d+, sohttp-gonein the younger buckets is an early recheck, and lifted / (lifted + held) there is how often the evidence that filed it was right. - Metrics:
prerender_opsseriesnegative_cache,negative_gap,gone_reopen,suppression_lifted,suppression_held;raw_cacheis now declared (it was emitted but missing from the catalog).
Deploying
- Schema: a new database
negative_cache(NegativePage, node-local, plus a replicated anchor table). It needs a worker restart per node to load. - Nothing changes behaviour until a route sets
negativeCache: truewithrender.negative.enabled, orrender.suppression.gone.reopen.dryRunis turned off. - With reopen on (the default), a bot-gated true 200 miss on a route that adds targets pays one detached local
Target.get. - Suggested rollout: enable the negative cache in dry run and read
would-serve-liveagainstwould-servefor a few days, then arm it; arm the reopen trigger oncegone_reopen would-filelooks sane.