Skip to content

Continuous batching #1333

Description

@andreapiso

Recently, a lot of benchmarks point to the fact that if you want to serve your models behind an API, continuous batching grants higher throughput and lower latency compared to static batching. Some examples of systems that implement continous batching:

In order to enable continuous batching, it is necessary to be able to:

  1. add requests to an existing running batch, if there are enough resources to take it (compared to static batching where requests need to be submitted all together)
  2. remove a request early from the batch when it reaches the stop token (as opposed to returning all requests at the same time).

Is this concept compatible with CTranslate2 architecture? I am keen to build an inference engine on top of CTranslate2, would love to hear some thoughts around this before I deep dive into it.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions