Repository navigation
Release 1.1.0
Adds a time-to-live to the per-worker inference cache to bound how long
stale entries can sit in memory between requests, in addition to the
existing size cap.
- Replaced
functools.lru_cachein
api/services/classification_service.pywith
cachetools.TTLCache(maxsize=10_000, ttl=3600)+@cached. Entries
now expire one hour after insertion, in addition to being evicted
when the size cap is reached. - A
threading.Lockis wired in via thelock=argument of@cached.
The lock is held only around cache reads and writes, not around the
wrapped inference call, so concurrent requests still run in parallel. - Expiry is lazy: stale entries are dropped when their key is next
accessed or when an insertion scans the cache. There is no background
sweep, so cache memory is reclaimed on use rather than on a timer. - Added
cachetools (>=5.3.0,<6.0.0)as a runtime dependency.
Note: if a worker is still being OOM-killed after this change, the cache
is unlikely to be the cause — at 10,000 entries × ~1–2 KB it is bounded
at roughly 10–20 MB. The more common culprits under sustained
transformers load are PyTorch allocator fragmentation and thread-stack
growth; mitigations there include gunicorn's --max-requests to
recycle workers periodically, restricting OMP_NUM_THREADS, and
tuning MALLOC_TRIM_THRESHOLD_.