Skip to content

Deployment Concurrency Limits #14934

Description

@jeanluciano

Describe the current behavior

Today, Prefect users can set concurrency for work pools and/or tags. Work pool concurrency is sufficient for throttling deployment execution in infrastructure with homogenous work. For this reason, users hoping to create deployment concurrency often create one-workpool-per-deployment, which is cumbersome, unergonomic, and a resiliency burden.

Describe the proposed behavior

Key components:

  1. Concurrency Limit: Allow users to set the maximum number of concurrent executions for a workflow

  2. Collision Handling: Provide options for managing executions when the concurrency limit is reached:

    • Enqueue: Add new executions to a queue and run them when a slot becomes available

    • Cancel: Cancel new execution attempts that exceed the concurrency limit

    • Cancel Running: Cancel the currently running execution(s) and start the new one

When a worker sees a scheduled run, it should check the deployment's concurrency limit and use the specified behavior:

If using concurrency_limit:

  1. Use a global rate limit inside of workers' run method, such that workers only spin up infrastructure and run a flow if they succeed in acquiring a slot for the limit

  2. Attempt to acquire a global concurrency limit, passing concurrency(…, timeout=0) to time out the concurrency check immediately if a slot isn't available

  3. If we acquire the lock, run the flow. If we can't acquire a lock, propose a new state, Scheduled("AwaitingConcurrencySlot"), handle with the corresponding collision strategy i.e re-queue the flow, then continue to the next flow run.

Otherwise, run the flow.

Additional configuration options:

Queue Priority: Set priority levels for queued executions when using the Enqueue option

Timeout: Set a maximum wait time for queued executions

Dynamic Concurrency: Allow concurrency limits to be adjusted based on system load or time of day

Example Use

Allows users to have fine-grained control over flow concurrency limits:

from prefect import deploy, flow
@flow()
def expensive_compute():
    ...

@flow()
def cheap_compute():
...

if __name__ == "__main__":
    deploy(
       expensive_compute.to_deployment(name="buy-deploy",concurrency_limit=3),
        cheap_compute.to_deployment(name="sell-deploy"),
        work_pool_name="my-dev-work-pool"
        image="my-registry/my-image:dev",
        push=False,
    )

Additionally, users would also be able to pass in one of ["enqueue", "cancel", "cancel-running"] to concurrency_limit:

if __name__ == "__main__":
    hello_world.deploy(
        name="LatestFlow",
        work_pool_name="my_pool",
        parameters=dict(name="Prefect"),
        image="my_registry/my_image:my_image_tag",
        concurrency={limit=1, collision_strategy="cancel-running"},
    )

Metadata

Metadata

Assignees

Labels

enhancementAn improvement of an existing feature

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions