Skip to content

1.0

Choose a tag to compare

@adhikjoshi adhikjoshi released this 15 Aug 07:33
· 4 commits to main since this release

Liveness release. Two independent fixes, both merged off the 2026-08-15 outage.

fix: bound the Redis read (#3)

A dropped TCP connection could park a worker forever on a blocking read.

blPop(['ml_tasks'], 1) looked safe, but that 1 is the server's timeout — how long Redis holds the pop before answering. It says nothing about how long the client waits for that answer. phpredis defaults its read timeout to 0 (wait forever) and connect() was called with no timeouts, so when a connection was silently dropped the reply could never arrive and the read never returned.

  • connect() now passes CONNECT_TIMEOUT (5s) and READ_TIMEOUT (10s)
  • OPT_TCP_KEEPALIVE (60s) so the kernel probes an idle peer, guarded by defined() — needs phpredis 5+
  • the blocking pop uses BLPOP_TIMEOUT (1s), deliberately held below READ_TIMEOUT; invert that ordering and every idle poll aborts the read before Redis replies

Callers injecting their own Redis client are unaffected.

perf: stop reading every task_result payload to prune (#2)

pruneOldTaskResults() scanned the whole task_result:* keyspace and GET the multi-MB value of every key just to read a timestamp — to delete keys Redis already expires on its own. Measured on one production shard over 26 days: 196.7M SCAN + 228.8M GET to issue 2 DELs.

Upgrading

No API changes. The new constants are public const on ModelQ and can be overridden by subclassing if your network needs different bounds.