Skip to content

Releases: HarperFast/prerender-plugin

prerender-v0.98.1 — @harperfast/prerender

Choose a tag to compare

@harper-joseph harper-joseph released this 01 Oct 21:34
6b8c624

@harperfast/prerender 0.98.1

PageCheck stores a digest of the observation, not the observation (#234).

  • A row is written for nearly every URL the nightly pass probes while changeProbe.serveCheck is enabled (dry run included), and for every serve-time check, and it replicates to every node. The observation, ~1 KB on a product endpoint, was most of each row.
  • The row now carries observedDigest: 96 bits of SHA-256 over the observation's UTF-16 code units, 16 characters. A row drops from ~1 KB to ~228 bytes.
  • The sweep compares observedDigest with the digest of its own baseline. It skips exactly the rows it skipped before.

Schema: PageCheck.signature is replaced by PageCheck.observedDigest. This is safe whether or not 0.98.0 rows exist: an old row's digest reads as null, so it spares no probe, and it is replaced on its next write. Deploy 0.98.1 rather than 0.98.0.

prerender-v0.98.0 — @harperfast/prerender

Choose a tag to compare

@harper-joseph harper-joseph released this 01 Oct 20:54
289083e

@harperfast/prerender 0.98.0

Serve-time checks: a cached page is checked against the origin when a bot asks for it (#232). Everything here is off by default; changeProbe.serveCheck ships in dry run.

  • changeProbe.serveCheck: a page served from cache that has not been checked or rendered since the last anchor (or, for the sample of URLs, within maxAge) is served exactly as before and checked in the background. The check compares the rule endpoint's mapped fields with the facts of the page that was just served, or, on a documentCheck route, the origin's own document head. A disagreement expires and re-files the page like a detected change. The same disagreement again, on a page rendered after it, with the origin unchanged, is held rather than re-rendered. Dry run counts serve_check would-check and asks nothing.
  • PageCheck (new table, verification database, replicated): every decided check, so a URL is asked once per window cluster-wide. The nightly pass skips a URL only when a check since the pass began observed exactly its baseline.
  • Origin budget: checks share the node's ratePerSecond with the pass, which leaves them what they use (at most serveCheck.share). Origin pushback pauses them node-wide.
  • util/documentFacts: reads the renderer's page facts off HTML bytes (head only, exact or null), held to Chrome parity by tests.
  • Raw cache: ingress.routes[].rawFacts stores page facts with a raw document (new nullable RawPage.facts), and a raw hit is checked from them. render.raw.ignoreNoStore stores despite the origin's no-store.
  • pageCheck.fields[].pathPattern: a mapped field compares only on URLs whose path matches; elsewhere it makes no claim.
  • Metric: prerender_ops serve_check (outcome by source).

Schema: the new PageCheck table and a nullable RawPage.facts column, both additive. Deploy console 0.23.0 for the Serve-time checks card.

prerender-v0.97.3 — @harperfast/prerender

Choose a tag to compare

@harper-joseph harper-joseph released this 01 Oct 20:52
5935a46

@harperfast/prerender 0.97.3

A URL too long to be a cache key is proxied to the origin, not answered 500 (#230).

  • Harper refuses a primary key over 1978 encoded bytes, and the first table read on the bot serve path threw on it, so the crawler got a 500 for a URL the origin serves.
  • CacheKey.fitsKeyLimit(url) is checked where a URL enters: a bot request over the limit is proxied and nothing else (bot_serve bypass, origin_fetch reason key-too-long, one warn); a sitemap <loc> over it is invalid; a render redirecting to one discards the render and keeps the target (a 301 retires the source and adopts nothing, unkeyable-destination); the admin explain and revalidate answer without reading.

No schema or config change: a worker restart loads it.

prerender-console-v0.24.0 — @harperfast/prerender-console

Choose a tag to compare

@harper-joseph harper-joseph released this 01 Oct 23:04
3c53d1e

@harperfast/prerender-console 0.24.0

  • Read cache on the console (#236). The console now answers repeat reads itself instead of passing every one to the prerender nodes, which also serve crawlers. Answers are stored in a node-local table (prerender_console.ProxyRead) that every worker on the console host shares.
    • Retention: analytics and crawl-breadth 60s, metrics 10 min, pages 15s, everything else 5s. session and page-content are never cached.
    • Who sees a cached answer: only an operator whose own session that node has confirmed as a super_user in the last 60s.
    • After a write: every POST that can change state clears the cache before it answers, so the reload after an action reads the node.
    • New field: each entry in sources.nodes[] now carries ageMs.
    • New option: cache: false turns the cache off.
  • Deploy note: the component now has a schema, the prerender_console database. The table is created on first load; there is nothing to migrate.

prerender-console-v0.23.0 — @harperfast/prerender-console

Choose a tag to compare

@harper-joseph harper-joseph released this 01 Oct 20:54
289083e

@harperfast/prerender-console 0.23.0

  • Serve-time checks card on the Probe view (plugin 0.98.0, prerender_ops serve_check): would-check or queued, agreed, mismatched, raw documents deleted, held, inconclusive, failed, and the checks that asked nothing, with mismatches by source (#232).
  • The Traffic view describes bypass as including a URL too long to be a cache key (plugin 0.97.3, #230).

prerender-v0.97.2 — @harperfast/prerender

Choose a tag to compare

@harper-joseph harper-joseph released this 30 Sep 18:41
30f9520

@harperfast/prerender 0.97.2

An uncompressed origin body is gzipped before it is stored, in the raw cache and the negative cache (#229).

  • A body the origin sent with no content-encoding (or identity) is gzipped at level 6 on the detached store path and stored under content-encoding: gzip. A body the origin already encoded is stored exactly as before.
  • The serve path is unchanged: it re-encodes from the stored header, so gzip clients get the stored bytes and others a decode.
  • maxBytes still counts bytes as received. It bounds the capture's memory, and for these bodies that memory is uncompressed. Size render.negative.maxBytes for an uncompressed error page: its default, 1 MB, fits the pages that prompted this.

Why: one production origin gzips its pages but sends every 404 uncompressed, whatever Accept-Encoding asks for (16 of 16 sampled, 79-699 KB). Stored as sent, those bodies took 7-13x the storage, and answering from storage would re-encode each one for the crawler. Measured: gzip takes 1.7-2.4 ms of CPU per 348-699 KB body and produces 47-54 KB, once per stored body.

No schema change, no config change. Rows stored before this release keep serving as they are.

1,823/1,823 plugin tests. Console unchanged (0.22.0).

prerender-v0.97.1 — @harperfast/prerender

Choose a tag to compare

@harper-joseph harper-joseph released this 30 Sep 16:26
886fbef

@harperfast/prerender 0.97.1

A snapshot's own validators are opt-in: page.snapshotValidators, default false. (#227)

  • Default (false): a rendered page (a cache serve, a peer rescue, a render-now result) carries no ETag and no Last-Modified, so a conditional request never gets a 304 from a snapshot.
  • true: 0.97.0's behaviour. The snapshot carries ETag: W/"<lastCachedMs>-<deviceType>" and Last-Modified = its render time, and a crawler holding the current render gets a 304 from the cache.
  • Unchanged from 0.97.0: a snapshot never carries the origin document's ETag / Last-Modified. Raw-cache documents and proxied responses keep the origin's validators.

Why: 0.97.0 gave every snapshot validators of its own. On a site whose origin sends none, that is the first time crawlers get validators from this cache, so they start revalidating and it starts answering 304s. That is a change to what crawlers receive, to make and measure per site (response_304 beside bot_serve), not a side effect of an upgrade.

Upgrading from 0.97.0: set page.snapshotValidators: true to keep 0.97.0's behaviour. It is a live option; no restart needed.

1,816/1,816 plugin tests. Console unchanged (0.22.0).

prerender-v0.97.0 — @harperfast/prerender

Choose a tag to compare

@harper-joseph harper-joseph released this 30 Sep 15:33
1f75f2e

Fixes from a full review of the plugin, reviewed again as one branch in several rounds (#226). The full list, with every option and metric, is in the PR.

  • Sitemap walk safety.
    • A failed child no longer turns shifted URLs into departures. It holds back only the URLs on the routes it last held.
    • A truncated document counts as a failed child.
    • sitemap.shrinkGuard refuses a much shorter document, or a dropped child that still holds URLs listed nowhere else. A refusal is accepted only after a new origin version (newVersionAfter) or acceptAge.
    • One lossless departure budget per walk, sitemap.departure.maxPerWalk (default 5000). It is spent only on departures that get an action and shared round-robin across children, and the rest are deferred to later walks.
    • The bypass token goes only to the root's host or domains, only over https, and redirects are followed manually.
    • Every hold-back is listed on the run row (holdBack).
  • Change probe.
    • A pass skips only rows probed since it began. reprobeAfter is retired.
    • The anchored pass is chained, caught up at boot, and resumed after a throw.
    • The reseed after a canary trip acts on changes, except covered or healed pages.
    • The render check (renderCheck, on) confirms a suspect render by one re-probe drawn from the sweep's own origin budget.
    • Per-batch counters, plus probe_errors / caught_up / ignored / covered / anchor / detection_lag / render_mismatch.
  • Render queue.
    • The slow-lane retry waits the effective cadence.
    • "Render this now" asks get a head start (urgentAt, 0.5) below detected changes.
    • fileDueNow off the owner forwards to the owner.
    • A result from a lease older than a change mark doesn't store the old page. Grant times are exact and releases are generation-checked.
    • render change_lag_ms and render_size.
  • Serve path.
    • Proxied 3xx keep Location.
    • Pages found but not served are proxied without validators, and cached pages carry a weak ETag.
    • Peer rescue checks the owner copy's serve status.
    • HEAD goes upstream as HEAD.
    • Invalidation state is held per worker from a subscription.
    • Proxied bodies go through originBodyStream, because Readable.toWeb's cancel can throw.
  • Admin and config.
    • POST /prerender_admin/revalidate works again (broken since 0.94.0).
    • Override rows report masking / redundant, and a set that would newly drop routes is refused.
    • Timer and read limits have maxima, and the analytics scan streams.

Deploying

  • Worker restart required. New attributes on RenderSchedule, ProbeState, ConfigOverride and Sitemap. The lease buffer is renamed, so expect one duplicate-render burst.
  • Don't write config overrides during the rolling restart.
  • Check stored child sitemaps are https on a host in domains.
  • Departures are metered at 5000 per walk even with maxActions: -1. Set maxPerWalk: -1 for the old behaviour.
  • Don't raise page.swrTtl yet: the probe's hard expiry is still relative to it.

prerender-console-v0.22.0 — @harperfast/prerender-console

Choose a tag to compare

@harper-joseph harper-joseph released this 30 Sep 15:33
1f75f2e

Panels for every series plugin 0.97.0 adds (#226):

  • Change probe → Change safety:
    • action errors, caught up, ignored and covered changes;
    • anchored-pass outcomes, with a skipped anchor flagged;
    • detection lag (median and ≈p95, per rule);
    • render-check outcomes.
  • Queue:
    • Change to cache: change mark to re-render, per route.
    • Page size: share over 1 MB, by route and device.
    • Filing: off-owner forwarding outcomes and unschedulable rows.
  • Sitemaps → Departures:
    • the departure family, grouped by what it means for a page: acted, shear, held back, deferred, skipped, lost;
    • shrink refusals and acceptances;
    • the live hold-back list.
  • Also:
    • the probe view describes per-batch counters;
    • sitemap detail pages over the stored entries.

Against an older plugin, the new cards show as empty.

prerender-v0.96.0 — @harperfast/prerender

Choose a tag to compare

@harper-joseph harper-joseph released this 29 Sep 19:48
b312c1d

A negative cache for origin 404/410s, and gone targets reopened when the origin answers 200 again (#225).

  • The negative cache (render.negative, route opt-in negativeCache: true). On a true miss, the origin's own 404/410 is stored, and the next crawler asking for the same dead URL is answered from it.
    • Three windows, measured from the last time the origin confirmed the status (checkedAt):
      • inside freshMs: answered from storage, and the origin is not asked. This is the only part that saves origin requests;
      • inside lifeMs: answered at once while a background re-check asks the origin. A 404/410 confirms the entry, a 200 or a redirect drops it, and a 5xx or timeout keeps it serving (stale-if-error);
      • past lifeMs: proxied, and a 404 stores again.
    • The stored body ages out too. Once the bytes are older than lifeMs, the next check that goes to the origin anyway replaces them (the re-check is a GET instead of a HEAD). No body is served more than about twice lifeMs old, at no extra origin requests.
    • It never answers a URL whose Target a sitemap lists, checked on every read and before every store (skipTargets: any also skips unlisted Targets). It never answers a bot in excludeBots. A bulk-invalidation epoch refuses entries confirmed before it, for every bot.
    • Store refusals are counted by name: has-cookie, private, no-store (ignoreNoStore for an origin that sends it on every document), staging, oversize, empty. assumeShared works as it does for the raw cache.
    • Re-checks are single-flight per key and capped per worker (maxConcurrentChecks); captures are capped like the raw cache's.
    • Counted honestly: a fresh answer is bot_serve source negative. An answer whose background re-check went to the origin is source origin, cacheStatus negative-revalidate, one per origin_fetch reason revalidate.
    • Off by default. dryRun (default on) stores and re-checks but never answers from storage, and counts would-serve (the saving) and would-serve-live (the risk: an armed cache would have answered a 404 while the origin said 200). negative_gap is the crawlers' re-ask curve to choose freshMs from.
  • Reopening gone targets (render.suppression.gone.reopen, on and dry-run by default). An origin 200 for a gone-suppressed target files its recheck due now, and the render's own verdict flips it.
    • The 200 can come from a proxied bot request (any bot, even one the discovery gate refuses), a negative-cache re-check, or the target rejoining its sitemap (the arrival check used to skip every suppressed row).
    • Gone verdicts only: a noindex or canonical-mismatch page answers 200 by definition. Bounded by dedupeMs per URL (6h) and maxPerMinute per worker (60); a failed filing is retried by the next 200.
  • Both outcomes of a suppressed target's render are counted, with the reason and the age since the last verdict: suppression_lifted and suppression_held. A gone target's own recheck lands in 14d+, so http-gone in the younger buckets is an early recheck, and lifted / (lifted + held) there is how often the evidence that filed it was right.
  • Metrics: prerender_ops series negative_cache, negative_gap, gone_reopen, suppression_lifted, suppression_held; raw_cache is now declared (it was emitted but missing from the catalog).

Deploying

  • Schema: a new database negative_cache (NegativePage, node-local, plus a replicated anchor table). It needs a worker restart per node to load.
  • Nothing changes behaviour until a route sets negativeCache: true with render.negative.enabled, or render.suppression.gone.reopen.dryRun is turned off.
  • With reopen on (the default), a bot-gated true 200 miss on a route that adds targets pays one detached local Target.get.
  • Suggested rollout: enable the negative cache in dry run and read would-serve-live against would-serve for a few days, then arm it; arm the reopen trigger once gone_reopen would-file looks sane.