Skip to content

Release 1.0.0

Choose a tag to compare

@cedricbonhomme cedricbonhomme released this 17 Apr 09:38
· 45 commits to main since this release
v1.0.0
e9bb8c4

First stable release. The focus of this release is production deployment:
serving many concurrent clients on a multi-core host without paying the
full model-loading cost once per worker, and without re-running inference
for inputs the server has already seen.

Production deployment with gunicorn + --preload

  • Models are now eagerly loaded at module import time via a new
    preload_models() function in api/models/severity_model.py, called
    from api/main.py before the FastAPI app is constructed.
  • Combined with gunicorn's --preload flag, every model is loaded once
    in the master process before workers are forked. Forked workers
    inherit the already-loaded tensors via copy-on-write, so the models
    occupy memory roughly once instead of once per worker. On a 16-core
    host running 4 workers this is a 4× reduction in model memory.
  • Without --preload, the new eager loading means each worker still
    loads every model at startup rather than lazily on first request.
    This trades a slower boot for predictable first-request latency and
    removes a class of "first request after restart is slow" issues.
  • The README now recommends a production command tuned for 16 cores:
    4 uvicorn workers with OMP_NUM_THREADS=4 / MKL_NUM_THREADS=4
    (4 workers × 4 PyTorch intra-op threads = 16 threads, keeping every
    core busy during inference while avoiding the memory overhead of
    8+ worker processes), plus --preload, --reuse-port and
    --proxy-protocol.

Per-worker LRU cache for inference results

  • Added an in-process functools.lru_cache (bounded at 10,000 entries)
    in api/services/classification_service.py, keyed by
    (model_name, description). Repeat requests for the same
    description against the same model now return the cached result
    without re-running the model.
  • Tradeoff: the cache is per worker, not shared across processes. With
    gunicorn's --reuse-port the kernel scatters incoming connections
    across workers, so duplicate requests only benefit from the cache
    when they happen to land on the same worker. In practice this is
    still a win because hot descriptions (e.g. re-enrichment passes over
    the same CVE) repeat often enough to land on the same worker
    multiple times, but users who need cross-worker deduplication should
    front ML-Gateway with an external cache such as Redis.
  • Memory footprint is bounded: 10,000 entries × ~1–2 KB per
    description ≈ 10–20 MB of cache per worker. The cache can be
    inspected or cleared at runtime with
    _cached_predict.cache_info() and _cached_predict.cache_clear().
  • Exceptions (e.g. unknown model name) are not cached, so error
    handling behaviour is unchanged.

Model registry cleanup

  • Removed CIRCL/vulnerability-severity-classification-distilbert-base-uncased
    from the LABELS registry. The corresponding Hugging Face repository
    is no longer published, so any attempt to load the model (including
    the new eager preload) failed with a 404. The entry was already
    commented out of the CLI's refresh-all command and was effectively
    unreachable.