Skip to content

v1.6.0

Choose a tag to compare

@PawelPeczek-Roboflow PawelPeczek-Roboflow released this 11 Sep 18:43
· 95 commits to main since this release
0b20131

🚀 Added

🧠 VLM blocks decode their own predictions

Until now a VLM block emitted a raw output string and a separate vlm_as_detector / vlm_as_classifier step turned it into predictions. Ten VLM blocks now parse the model answer themselves, through a shared decoding package keyed by the box-coordinate contract the vendor speaks (xyxy_absolute, xyxy_0_1000, yxyx_0_1000, xyxy_percent, named_0_1000, named_normalized), so two vendors asking for the same contract share one parser. Each new block version adds three outputs next to the existing output, classes and token counts: predictions (object-detection or classification kind depending on the task, None for the other tasks), error_status and inference_id. An answer in the wrong contract now sets error_status=True instead of silently returning zero detections (@Erol444, #2909).

The new versions. roboflow_core/anthropic_claude@v5, roboflow_core/open_ai@v7, roboflow_core/google_gemini@v6, roboflow_core/openrouter@v3, roboflow_core/kimi_openrouter@v3, roboflow_core/google_gemma@v4, roboflow_core/qwen_vlm@v4, roboflow_core/zai_vlm@v2, roboflow_core/meta_vlm@v3 and roboflow_core/spacexai@v3. Swap the step to the new version, read $steps.<name>.predictions directly and delete the formatter step. openrouter@v3 gains a detection_format field (default named_normalized, any of the six contracts or a selector) and an unknown value fails validation before a paid request is sent. The superseded versions and the vlm_as_* blocks are deprecated; see the notices at the bottom.

🐘 Native PostgreSQL sink

roboflow_core/postgresql_sink@v1 writes workflow results straight into a PostgreSQL table, replacing the hand-written insert code people kept putting in Custom Python blocks. It inserts one row or a list of rows with identical column names in a single transaction. Required inputs are host, database, username, table_name and data (a dict or a list of dicts); optional ones are password (a secret-kind selector, None means libpq-configured auth), port (5432), schema_name (public), fire_and_forget (True), sslmode (require), connect_timeout (10 s) and statement_timeout (10000 ms). Outputs are error_status and message. The block's description carries a full guide with a CREATE TABLE snippet and an example workflow (@maxschridde1494, #2967).

Where it runs. It is an enterprise block: it loads only with LOAD_ENTERPRISE_BLOCKS=True and refuses to run on the Roboflow hosted platform, so it is for self-hosted servers and dedicated deployments. The psycopg[binary] driver now ships in every server image and in the inference, inference-cpu and inference-gpu packages, so it is importable from Custom Python blocks as well (#2966). Outbound connections are subject to the address controls described in the security section below.

At this moment, PostgreSQL sink is not available on the serverless platform

🔁 Remote dispatch for inner workflows

roboflow_core/inner_workflow@v1 gains a second execution mode. The default embedded mode is unchanged: the child workflow is inlined into the parent graph at compile time. The new remote_dispatch mode keeps the step as an outputless sink that serialises the bound child inputs and posts the child workflow to another inference server in a background task. It is fire-and-forget: no outputs, failures are logged only, no durable queue and no delivery acknowledgement. Set execution_mode: "remote_dispatch" on the step and, optionally, a per-step remote_target base URL (no credentials, query or fragment allowed). The server-wide default target is WORKFLOWS_INNER_WORKFLOW_REMOTE_TARGET (https://serverless.roboflow.com) with WORKFLOWS_INNER_WORKFLOW_REMOTE_DISPATCH_REQUEST_TIMEOUT (300 s). The parent API key is forwarded only when the resolved target equals that configured default; a per-step target pointing elsewhere is called without it, and redirects are never followed. Dispatch chains are capped by the existing WORKFLOWS_MAX_INNER_WORKFLOW_DEPTH (4) through a new inner_workflow_dispatch_depth field on the workflow request, so every server in a chain must run a version that accepts it (@hansent, #2915).

📈 Per-source completed-frame telemetry for stream pipelines

GET /inference_pipelines/{pipeline_id}/status now reports cumulative per-camera completed-frame counters, so a client can compute true per-source FPS and spot a stalled camera. The existing inference_throughput counts submitted work, which can run ahead of asynchronous predictions; the new counters advance after a prediction resolves and before sink delivery, and are returned as one atomic snapshot under report.completion_statistics with sampled_at_monotonic and a sources list of source_id, completed_frames, last_completed_at_monotonic and last_frame_id. Monotonic values are process-local, so compare them only between snapshots of the same pipeline. In the same change, POST /inference_pipelines/{pipeline_id}/consume keeps None placeholders in outputs and frames_metadata for cameras that produced no frame in a multi-camera batch, instead of dropping them and misaligning the lists (@grzegorz-roboflow, #2976).

🎬 Zero-shot action recognition is nvidia/cosmos-3-edge-action-recognition

The hosted zero-shot action-recognition model id changes from cosmos-3-edge/action_recognition to nvidia/cosmos-3-edge-action-recognition. The old id was only ever registered on US staging, so it failed the serverless access check in production and in both EU environments; the new id is a platform alias of the base nvidia/cosmos-3-edge, so the access check passes and the weights provider serves the base package. The old id is not accepted any more on any surface (POST /infer/action_recognition, the SDK, or roboflow_core/roboflow_action_recognition_model@v1), so update saved workflows and client code (@leeclemnet, #2965).

🤖 GPT-6 Astra in the OpenAI block

gpt-6-astra is a first-class option of the OpenAI block, with reasoning effort low, medium, high, xhigh or max and the structured-absolute detection prompt style, instead of being reachable only through a selector-supplied id with no validation. Pick it in roboflow_core/open_ai@v7, which is the version carrying the in-block decoding above (@SkalskiP, #2935).

🔒 Security hardening

This release closes a series of findings from our security review of the self-hosted server. Several of them change defaults, so a 1.5.x deployment can behave differently after the upgrade. Each change is described below, and the administrator checklist at the end of this section lists every action you may need to take. Read the checklist before you upgrade.

🏠 inference server start publishes the server on localhost by default

The CLI used to publish the container port on every interface of the host, so a server started with inference server start was reachable from the whole network the moment it came up. It now publishes on the host's loopback address, which is the equivalent of docker run -p 127.0.0.1:9001:9001. This is a host-side port mapping, not a change inside the image: the server process still binds 0.0.0.0 inside the container, HOST keeps its 0.0.0.0 default in every Dockerfile, and nothing changes for deployments you run yourself with docker run -p 9001:9001, for Kubernetes and the Helm chart, or for the Jetson images, which keep publishing on all interfaces because a Jetson is usually the camera host itself. inference server start --tunnel also keeps 0.0.0.0, since the tunnel container reaches the server through the host gateway. The macOS and Windows app bundles now honour HOST and default it to 127.0.0.1; before, they set it and then passed a hardcoded 0.0.0.0 to uvicorn anyway (@grzegorz-roboflow, #2943).

How to expose a CLI-started server on the network again. Pass the new --bind-address (-b) flag; the CLI then prints a warning that names the two settings you should review when a server is reachable by others, WORKSPACES_WHITELISTED_FOR_LOCAL_DEPLOYMENT and ALLOW_CUSTOM_PYTHON_EXECUTION_IN_WORKFLOWS.

inference server start --bind-address 0.0.0.0

In the same change the server gained a startup security notice: when custom Python execution is enabled in local mode and the server is neither whitelisted to a workspace nor bound to a dedicated deployment, a SECURITY: line is logged at boot. It is a log warning, so INFERENCE_WARNINGS_DISABLED does not hide it. The earlier deprecation warning that announced a flip of the ALLOW_CUSTOM_PYTHON_EXECUTION_IN_WORKFLOWS default has been removed; the default stays True, and the decision about the server's exposure is left to you through the binding and the settings above.

🔐 Runtime and model-loading hardening

One pull request tightens a set of runtime defaults that all touch how the server trusts its configuration and its inputs (@alexnorell, #2952). These are the changes that alter behaviour:

  • Stream API video references are validated. POST /inference_pipelines/initialise now rejects raw GStreamer launch strings, bare identifiers, bare relative file names such as video.mp4, and camera indices passed as strings such as "0". Integer indices, /dev/video0, ./video.mp4, rtsp://, csi://0 and file:// references are accepted. The in-process InferencePipeline.init() Python API is not gated. Set ALLOW_UNSAFE_GSTREAMER_PIPELINES=True to restore the old behaviour if you rely on raw pipelines.
  • SECURE_GATEWAY (and the legacy LICENSE_SERVER) defaults to HTTPS. A bare host[:port] value used to be proxied over plain HTTP; it is now treated as https:// with a warning. An explicit http:// URL is refused at startup unless the host is loopback, and so are values with embedded credentials, a query string, a fragment or whitespace. This applies to the server and to the inference_models library alike.
  • SSL_CA_CERTS now requires client certificates. When ENABLE_HTTPS=True and SSL_CA_CERTS is set, uvicorn is started with --ssl-cert-reqs 2, so every client must present a certificate signed by that CA. Previously the CA bundle was loaded but clients were not verified.
  • WORKFLOWS_CUSTOM_PYTHON_EXECUTION_MODE is validated. The value is normalised and must be local or modal; anything else fails at startup. OFFLINE_MODE=True combined with modal is now a startup error instead of a warning and a silent downgrade to local.
  • Jetson images no longer load untrusted inference_models packages by default. The jetson-5.1.1, jetson-6.2.0 and jetson-7.2.0 images flip ALLOW_INFERENCE_MODELS_UNTRUSTED_PACKAGES from True to False, which is what the x86 images already did. Set it back to True in the container environment if you load client-registered or non-Roboflow packages on Jetson. This is particularly important for TRT packages compiled outside the platform and registered back - those are considered untrusted, yet Roboflow platform providing packages makes sure that untrusted packages exposed to specific client are only the ones created by client or us (w.r.t. pre-trained models).
  • inference benchmark inference-models-speed defaults to --no-allow-untrusted-packages; pass --allow-untrusted-packages for the old behaviour.
  • The enterprise parallel image refuses SSL_KEYFILE_PASSWORD together with ENABLE_HTTPS, because gunicorn cannot take an encrypted key safely; use an unencrypted key file or the uvicorn launcher.

The rest of the change is transparent: the OPC UA connection pool keys sessions on an HMAC over URL, user and password so two users sharing a username never share a session; the usage collector no longer deadlocks on a saturated queue; the SSL arguments of the entrypoint are passed positionally so certificate paths with spaces work; and the release workflows moved every ${{ }} expression into env: blocks, validate tags against a strict pattern and pin the review action to a SHA.

🌐 Outbound connections made by Workflows are gated

Four blocks reach out of the server to a destination taken from the workflow definition. Each of them now has a boundary the operator controls.

  • OpenAI-compatible block: an operator allowlist. roboflow_core/openai_compatible@v1 used to pass a workflow-supplied base_url straight to the OpenAI client, so a workflow author could send the prompt, the images and the API key anywhere. The server now checks the base URL against OPENAI_COMPATIBLE_ALLOWED_BASE_URLS on every call, rejects URLs with credentials, a query string, a fragment or a non-HTTP scheme regardless of the list, and never follows redirects. The default is *, which allows any well-formed URL, so nothing changes until you restrict it; a rejected endpoint surfaces as error_status on the step with a message naming the variable, not as a failed request. Matching is exact after trailing-slash removal: include the /v1 suffix, match the scheme, host case and port, and separate values with commas. An empty value blocks every endpoint. The openai client floor rises to 1.17.0 (@grzegorz-roboflow, #2957).
  • WebRTC MJPEG input: SSRF protection and open deadlines. A caller-supplied mjpeg_url no longer reaches FFmpeg's networking stack. The stream is opened through the same SSRF-protected HTTP adapter the URL image input uses, redirects are followed manually with re-validation on every hop, the resolved address is pinned so DNS rebinding cannot swap it, and FFmpeg only decodes what it is handed. With the new WEBRTC_MJPEG_ALLOW_NON_GLOBAL_ADDRESSES=False default, loopback, private, link-local (including the cloud metadata address), carrier-grade NAT, unique-local IPv6 and multicast destinations are refused. Each hop has a 5 second connect-and-read budget, the whole open has a 5 second wall-clock deadline so a trickling peer cannot hang the worker, and at most 5 redirects are followed (@grzegorz-roboflow, #2958).
  • PostgreSQL sink: address controls. The new sink opens a TCP connection to a host from the workflow definition, so it ships with ALLOW_POSTGRESQL_WORKFLOWS_SINK_TO_NON_GLOBAL_ADDRESSES (default True), POSTGRESQL_WORKFLOWS_SINK_BLACKLISTED_ADDRESSES and POSTGRESQL_WORKFLOWS_SINK_WHITELISTED_ADDRESSES (both unset). With the defaults the host is used as given. Setting the first to False or either list activates screening: the host is resolved, the denylist wins over the allowlist, matching is exact and case-sensitive on the raw host and on every resolved address, the connection is pinned to the resolved IP while keeping the hostname for TLS, and Unix-socket paths and multi-host strings are refused. Write the lists without spaces. Under the block's default fire_and_forget=True a rejected destination is reported in the server log only; set fire_and_forget=False to see it in the step output (@maxschridde1494, #2967).
  • S3 sink: explicit credentials only. roboflow_core/s3_sink@v1 no longer falls back to the server's ambient AWS identity. Because /workflows/run accepts caller-supplied definitions, any workflow author could previously write chosen content to a chosen bucket as the instance role, task role or configured profile of the server. The block now requires aws_access_key_id and aws_secret_access_key from the workflow, ideally through an Environment Secrets Store step, and fails the step when either is missing. There is no opt-out (@Jonathan-Roboflow, #2964).

🧱 Resource bounds against denial of service

  • Stream manager pipeline cap. The stream manager spawned one OS process per initialise that could not reuse an idle pipeline, with no ceiling. STREAM_MANAGER_MAX_ACTIVE_PIPELINES (default 8, floored at STREAM_API_PRELOADED_PROCESSES) now caps live pipelines per container; idle pipelines count, terminating one frees a slot, and an initialise at the cap returns HTTP 500 with the descriptive message in the container log. The same change fixes a TypeError in the RAM guard that turned an initialise issued within a second of a spawn into an internal error on servers without STREAM_MANAGER_MAX_RAM_MB (@grzegorz-roboflow, #2944).
  • Non-realtime RTSP decoding is bounded by backpressure. For WebRTC sessions with webrtc_realtime_processing=False, RTSP frames were pushed into an unbounded queue, so a camera faster than inference grew memory without limit. A dedicated decode thread now feeds a 60-frame queue and blocks when it is full, preserving order and dropping nothing; decode failures are reported as a fixed message so an rtsp://user:secret@host URL cannot leak through an FFmpeg error. Realtime RTSP and file sources are unchanged (@digaobarbosa, #2959).
  • Keypoint padding is bounded. Keypoints are ragged and get right-padded to the batch maximum, so an 840 KB workflow input could demand about 700 MB of dense arrays. A shared validator now caps the padded product at 1,000,000 cells before any allocation, in all six numpy and torch padding paths; an oversized workflow input is rejected with HTTP 400. A 17-keypoint model would need about 58,000 detections in one image to reach it (@grzegorz-roboflow, #2955).
  • Nearest-neighbour matching is capped. roboflow_core/detections_nearest_neighbor@v1 builds a full query-by-target distance matrix. It now refuses more than 1,000 detections per set, more than 10,000 matched pairs, and more than 100 matched pairs when either side carries instance-segmentation masks, raising an error with a remediation message rather than truncating. The mask cap is the one to watch: a dense segmentation scene with more than 100 matched objects now fails, and there is no configuration knob, so limit detections upstream (@bczifra, #2963).
  • Track Class Lock is bounded in CPU and memory. roboflow_core/track_class_lock@v1 rescanned the whole per-video track store for every new tracker id, an O(n²) path reachable through /workflows/run with caller-supplied detections, and kept unbounded state. Eligible sources are now computed once per frame, state is capped at 4,096 tracks per video with least-recently-seen eviction, state_ttl is capped at 1,000,000 frames (validation error above it), and re-attachment considers the 256 most recently lost tracks. Worst-case re-attachment on a 4,000-id frame drops from about 3.5 s to 0.02 s. Both the numpy and the tensor-native block are covered (@alexeialexandrovich, #2937).
  • Local model paths cannot be authorised per API key. With MODELS_CACHE_AUTH_ENABLED=True the server authorises every model id against the platform, but a model id that names a local package directory was accepted without that check when ALLOW_INFERENCE_MODELS_DIRECTLY_ACCESS_LOCAL_PACKAGES=True. The server now refuses to start with both enabled on an online server; both default to False, so stock deployments are unaffected, and the OFFLINE_MODE opt-in through ALLOW_OFFLINE_MODEL_CACHE_AUTH_BYPASS keeps working (@grzegorz-roboflow, #2956).

🛠️ Administrator checklist for upgrading from 1.5.x

Work through this list before you roll v1.6.0 out. Items are ordered by how likely they are to affect a typical deployment.

  1. Servers started with inference server start are reachable from localhost only. Clients on other machines get connection refused. Add --bind-address 0.0.0.0 to expose the server on the network. Deployments started with your own docker run -p 9001:9001, Kubernetes, Helm, --tunnel and the Jetson images are not affected (#2943).
  2. Every workflow that uses the S3 Sink must pass aws_access_key_id and aws_secret_access_key. Instance roles, task roles, web-identity roles, ~/.aws/credentials and ambient environment credentials are no longer used, and there is no opt-out. Audit your workflows and move the keys into an Environment Secrets Store step (#2964).
  3. SECURE_GATEWAY / LICENSE_SERVER values without a scheme now mean HTTPS, and plain http:// refuses to start unless the host is loopback. Air-gapped deployments must terminate TLS on the gateway and configure an explicit https:// URL (#2952).
  4. SSL_CA_CERTS now enforces client certificates. If you set it together with ENABLE_HTTPS=True and your clients do not present certificates, every request fails the TLS handshake. Unset it or issue client certificates (#2952).
  5. More than 8 concurrent stream pipelines on one server now fails. Set STREAM_MANAGER_MAX_ACTIVE_PIPELINES to your intended concurrency before upgrading. The explanatory message is in the container log; the HTTP response is a generic internal error (#2944).
  6. MJPEG cameras on LAN, loopback or link-local addresses stop working in WebRTC sessions. Set WEBRTC_MJPEG_ALLOW_NON_GLOBAL_ADDRESSES=True on the server or worker that opens the stream, and only where untrusted workflow requests are not accepted, because the setting lifts the block for the whole process (#2958).
  7. Jetson images no longer load untrusted model packages. Set ALLOW_INFERENCE_MODELS_UNTRUSTED_PACKAGES=True in the container environment if your Jetson fleet loads client-registered or non-Roboflow packages (#2952).
  8. Stream API video references are validated. Raw GStreamer pipeline strings, bare relative file names, bare identifiers and string camera indices are rejected over HTTP. Use an integer index, an explicit path or a supported URL, or set ALLOW_UNSAFE_GSTREAMER_PIPELINES=True (#2952).
  9. OFFLINE_MODE=True with WORKFLOWS_CUSTOM_PYTHON_EXECUTION_MODE=modal no longer starts, and neither does any value other than local or modal. Set the mode to local explicitly (#2952).
  10. MODELS_CACHE_AUTH_ENABLED=True together with ALLOW_INFERENCE_MODELS_DIRECTLY_ACCESS_LOCAL_PACKAGES=True no longer starts on an online server. Drop one of the two. Stock deployments leave both at False and are unaffected (#2956).
  11. detections_nearest_neighbor@v1 errors above 100 matched pairs when masks are present, and above 1,000 detections per set or 10,000 pairs otherwise. There is no knob; filter detections before the block (#2963).
  12. The OpenAI-compatible block keeps allowing any endpoint by default, so no action is required to keep it working. Base URLs with credentials, query strings or fragments, endpoints reached through a redirect, and environments pinning openai below 1.17.0 now fail. To restrict, set OPENAI_COMPATIBLE_ALLOWED_BASE_URLS to the exact base URLs including the /v1 suffix and restart (#2957).
  13. The PostgreSQL sink ships permissive. If you expose Workflows to untrusted authors, set ALLOW_POSTGRESQL_WORKFLOWS_SINK_TO_NON_GLOBAL_ADDRESSES=False and pin your databases with POSTGRESQL_WORKFLOWS_SINK_WHITELISTED_ADDRESSES=db1.example.com,db2.example.com, written without spaces. Rejections are log-only under the default fire_and_forget=True (#2967).
  14. Two smaller defaults moved: inference benchmark inference-models-speed needs --allow-untrusted-packages for the old behaviour, and the enterprise parallel image refuses SSL_KEYFILE_PASSWORD with ENABLE_HTTPS (#2952).
  15. Dependency floors: environments pinning GitPython below 3.1.59 or openai below 1.17.0 no longer resolve (#2972, #2957).

All new variables introduced in this section: ALLOW_UNSAFE_GSTREAMER_PIPELINES (False), STREAM_MANAGER_MAX_ACTIVE_PIPELINES (8), OPENAI_COMPATIBLE_ALLOWED_BASE_URLS (*), WEBRTC_MJPEG_ALLOW_NON_GLOBAL_ADDRESSES (False), ALLOW_POSTGRESQL_WORKFLOWS_SINK_TO_NON_GLOBAL_ADDRESSES (True), POSTGRESQL_WORKFLOWS_SINK_BLACKLISTED_ADDRESSES (unset) and POSTGRESQL_WORKFLOWS_SINK_WHITELISTED_ADDRESSES (unset). Unchanged, because people will ask: HOST stays 0.0.0.0 in every image, and ALLOW_CUSTOM_PYTHON_EXECUTION_IN_WORKFLOWS stays True.

🔧 Fixed

  • WebRTC SDK sessions no longer leak after a failed start — when the first connection attempt of an inference_sdk WebRTCSession failed, the session stayed STARTED, its private event loop and daemon thread kept running, and a peer connection created late in startup was never closed because teardown could race the startup task. A failed start now moves the session to CLOSED under the state lock, the startup task is retained and awaited on its own loop before the peer and the source are closed, and the loop thread drains pending tasks before shutting down. A later call on that session raises "Cannot use closed WebRTCSession" instead of operating on a half-dead one; the original HTTP diagnostics are preserved (@voropaevv, #2808).
  • Empty VLM detections keep their image dimensions, in own and in parent coordinates — sv.Detections.data is per row, so a VLM that returned zero detections produced {"image": {"width": null, "height": null}, "predictions": []} because the serialiser's per-row loop never ran. The seven VLM-as-detector parsers (Anthropic, Gemini, Muse, OpenAI, Qwen, SpaceXAI and the generic LLM path) now build their empty result through a helper that stores the dimensions in object-level metadata, and the serialiser falls back to it (@davidnichols-ops, #2892). The follow-up extends the same idea to lineage: an empty result from a VLM run on a crop carried no parent information, so outputs declared in parent coordinates reported the crop's dimensions instead of the root image's. Parent id, parent coordinates and parent dimensions are now written to metadata for zero-row detections and rewritten to root values during coordinate conversion (@grzegorz-roboflow, #2979).

⚙️ Execution Engine v1.15.2

The Workflows Execution Engine version moves from v1.15.1 to v1.15.2 alongside the runtime hardening in #2952. The bump records the engine build that ships with the stricter runtime validation described above; block contracts and the workflow definition schema are unchanged, so existing definitions run as before. The full version-by-version record lives in the Execution Engine changelog.

🚧 Maintenance

  • Security dependency refresh (2026-09-11) — inference-models moves to 0.37.1, raising its GitPython floor to 3.1.59 (3.1.62 in the lock); every server image pins inference-models~=0.37.1. Alongside it, the openai client floor rises to 1.17.0 (#2957), psycopg[binary] joins the shared requirements for the new PostgreSQL sink (#2966), and the landing page is rebuilt on Next.js 16.3.4 and sharp 0.35.4 (@PawelPeczek-Roboflow, #2972).

  • CI cancels superseded runs — one canonical concurrency block now sits on 49 workflows: one live run per workflow per ref, a newer push to the same PR or to main cancels the older run, manual dispatches are keyed by run id so they never cancel or get cancelled, and a guard script fails CI when a pull_request- or push-triggered workflow lacks the block (@PawelPeczek-Roboflow, #2970).

  • Test rename reverted pending CLA — a rename of five rfdetr-seg static-crop tests whose names shadowed the letterbox tests (@Anai-Guo, #2948) was reverted the same minute because the CLA was not signed (#2971). Net effect on this release: none; the fix is welcome back once the CLA is in.

  • Usage rows carry model latency — the registry's optional modelLatencyMs is read from platform model metadata and carried through the metadata cache into usage telemetry as resource_details.model_latency_ms, so Serverless can price NAS images without a billing-time registry lookup. Nothing changes in any HTTP response, SDK entity or workflow output (@maxschridde1494, #2969).

⚠️ Deprecations and platform notices

  • vlm_as_detector and vlm_as_classifier are deprecated, along with the VLM block versions they served. roboflow_core/vlm_as_detector@v1/@v2, roboflow_core/vlm_as_classifier@v1/@v2, their tensor variants, and anthropic_claude@v4, open_ai@v6, google_gemini@v5, openrouter@v2, kimi_openrouter@v2, google_gemma@v3, qwen_vlm@v3, zai_vlm@v1, meta_vlm@v2 and spacexai@v2 are marked deprecated in favour of the in-block decoding above. Deprecated here means a flag in the block metadata that the Workflows editor uses to steer new workflows to the successors: every one of these blocks still loads and runs unchanged, and no removal date is set. roboflow_core/llama_vision@v1/@v2 are deprecated without a successor because OpenRouter delisted the model; use roboflow_core/meta_vlm@v3 (#2909).
  • depth-anything-v2/base and depth-anything-v2/large are sunset. Usage was zero, so the weights are being withdrawn from the platform registry; depth-anything-v2/small and the bare depth-anything-v2 alias stay. Nothing in the server rejects or remaps the two ids, so a request for them fails at load time once the weights are gone. Move to depth-anything-v2/small or to Depth Anything V3 (#2980).
  • Jetson images no longer load untrusted inference_models packages by default. All three Jetson Dockerfiles (jetson-5.1.1, jetson-6.2.0, jetson-7.2.0) flip ALLOW_INFERENCE_MODELS_UNTRUSTED_PACKAGES from True to False, bringing them in line with the x86 images. See the administrator checklist above if your fleet relies on it (#2952).
  • JetPack 5 and JetPack 6.2 end of life stays at the end of 2026, as announced in v1.5.2. JetPack 7.2 is the platform going forward; the JetPack 5 image remains on transformers 5.7.0 and keeps the security caveat described in the v1.5.2 notes.

🏅 New Contributors

Full Changelog: v1.5.2...v1.6.0