v1.7.2
🚀 Added
🔬 Ask a Workflow what it would do, without doing it
POST /workflows/describe_workload reads the recipe card without turning on the oven. It compiles the definition and hands back the graph, how deep each step goes, which models it needs, what kind of work it does, and where it refuses to run. No block is initialised, no model is loaded, no custom Python runs. Point a cost estimator, a scheduler, or plain curiosity at it. Experimental: the contract may still move. Field guide in docs/workflows/workload_introspection.md (@PawelPeczek-Roboflow, #3025).
curl -X POST localhost:9001/workflows/describe_workload -H "Authorization: Bearer $KEY" -d @workflow.json📡 Workflows can now listen to the factory floor
The MQTT Writer could only shout. roboflow_enterprise/mqtt_reader@v1 listens: one message per run, newest or next unread, retained messages on the first run, and a full backlog after a restart when you give it a client_id. TLS landed for both Reader and Writer. Operators get a broker allowlist. Enterprise only, not on hosted, happiest inside an InferencePipeline (@tewanchia, #3034).
🏎️ RF-DETR: one execution plan, every backend, plus a SIMD lane
Torch and ONNX now run the same five-stage plan TensorRT had. A new pillow-simd-v1 preprocessor ships inside the CPU, GPU and CUDA 13 images; auto picks Triton on CUDA, SIMD otherwise, base last. The author's T4: 4K ONNX infer() 101 ms → 31 ms. Every model now carries a receipt, optimization_runtime_metadata, saying which implementation actually ran each stage and whether it fell back. Triton staging buffers are capped at 8K, so a request can no longer size the GPU's memory for it. Model-eval requests stop being bounced to the slow path for turning off switches that were already off (@dkosowski87, #3017, #2903, #2985, #3063).
🏷️ Describe "a small household feline", get cat back
roboflow_core/open_ai@v7 gains output_classes. Prompt with a description, receive a clean label downstream. On GPT models with structured outputs the label is a JSON-schema enum, so no creative spellings (@lucas-fochesatto, #3036).
🧠 Four new brains
Native Qwen 3.8 VL 27B in qwen_vlm@v2–v4, with thinking; it was already in the building, only the v1 block had it on the directory board (@lrosemberg, #3040). GPT-6 Sol and Luna in open_ai@v7, Claude Opus 5.5 in anthropic_claude@v5 (#3051). Grok 4.7 with reasoning effort in spacexai@v3, plus PNG level 9 so a 4K frame fits under xAI's 25 MB limit (#3046). Hosted has had all four since 2026-09-22; pip gets them today.
🐍 Python 3.13
Colab flipped to 3.13 on 2026-09-21. On 2026-09-22, pip install inference on 3.13 stopped resolving to nothing. All packages declare <3.14, CI runs it. Fine print: pybase64 builds from source, Docker images stay on 3.11/3.12 (@PawelPeczek-Roboflow, #3041, #3067; @grzegorz-roboflow, #3065).
🦭 Podman
inference server start no longer assumes a Docker daemon. Podman is auto-detected (or forced with INFERENCE_CONTAINER_RUNTIME=podman), the GPU goes through CDI, SELinux stays enforcing. Not yet: --tunnel and Jetson. Community contribution (@jim60105, #2994).
🔧 Fixed
- Blur follows the silhouette, not the rectangle — the block's description always promised it; the code drew a courtroom-sketch box. New
paddingfor when the mask runs tight around hair and fingers (@alessandro-j-ds, #3061). - Keypoint skeletons stop buttoning the shirt one hole off — a missing nose used to shift every bone after it; frames where nobody had all 17 COCO points drew no bones at all (@alessandro-j-ds, #3062).
- A Workflow that calls itself gets a 400, not a stack overflow — cycles are caught, and depth (4) and count (32) limits fire before anything is fetched (@dkosowski87, #3055).
- WebRTC SDK picks a TURN server that answers — behind a TCP-443-only firewall,
maindelivered 0 frames in 90 s; now the first frame arrives (@japrescott, #3058). - MQTT Writer hangs up once on a wrong password — instead of redialling every second until the broker rate-limits you (@tewanchia, #3034).
- Executor-less runs keep one crew for the whole run — no more hiring a thread pool per wave of steps. Author's numbers: −35% on a 10-step chain, −41% on 50 (@davidnichols-ops, #3052).
- Template Matching remembers how big the picture was when it finds nothing (@davidnichols-ops, #3038).
- Editor restriction lists tell the truth — BoT-SORT, PostgreSQL sink, SAM2/SAM3 video, action recognition, LMM and Continue If now declare what they always did (#3025).
⚙️ Execution Engine v1.16.0
1.15.2 → 1.16.0: workload introspection, the inner-workflow cycle guard, the run-scoped thread pool, and one stricter rule: a requested engine version is a minimum, and asking for a newer one than installed fails before compilation. Definitions run as before. The canonical record now lives in workflows/CHANGELOG.md; history up to 1.15.2 stays in the Execution Engine changelog.
🚧 Maintenance
- Versions —
inference1.7.2,inference-models0.39.0,roboflow-workflows0.2.3(@grzegorz-roboflow, #3069). - Where did 1.7.0 and 1.7.1 go? PyPI only, both on 2026-09-22, no tags, no images.
1.6.2-post1and1.6.2-post2went the other way: images and hosted, never PyPI. These notes cover the whole gap. inference-modelsis pinned with==now — a pre-release built on a feature branch reachedmainthrough~=in three hours. Once was enough (@PawelPeczek-Roboflow, #3068).- JetPack 5 image builds again — its GCC 9.4 cannot compile
zxing-cpp2.3.0, so the pin is per Python version (@PawelPeczek-Roboflow, #3067). - Process — a label-gated fast-track CI lane for post-release fixes (#3053, #3054); automated review hands off to a maintainer in Slack, documented in
CONTRIBUTING.md(#3035, #3039);AGENTS.md, a plan template and a single canonical Workflows changelog (#3043, #3044, #2902); the Codeflash benchmark CI is gone (#3037).
📜 Licensing and pricing wording
One consistent statement across the READMEs and model pages: Roboflow Cloud products include a commercial license for the models Roboflow can relicense, for every user; a commercial license for self-hosted deployment is an Enterprise add-on; without it, the model's own license applies, and AGPL is AGPL. Inference on your own hardware does not use credits; the Serverless Cloud API bills per image at the model's rate (@SolomonLake, #3047, #3057).
⚠️ Notices
- RF-DETR
threaded-exact-v1is gone. Selecting it fails model load. INFERENCE_MODELS_RFDETR_POSTPROCESSOR=triton-fused-v1now breaks Torch and ONNX loads. It was silently ignored there before. Leave it onauto.- SIMD is the default preprocessor in the CPU, GPU and CUDA 13 images. Resize output can drift by 1 intensity level.
INFERENCE_MODELS_RFDETR_PREPROCESSOR=basefor pixel-exact. rfdetr_execution_plan→execution_plan. The old name warns until 2026-10-24, then disappears.- Blur and keypoint pixels changed. On purpose. Re-baseline any golden images.
- JetPack 5 and 6.2 reach end of life at the end of 2026. JetPack 7.2 is the road ahead.
🏅 New Contributors
- @lucas-fochesatto made their first contribution in #3036
- @jim60105 made their first contribution in #2994
- @alessandro-j-ds made their first contribution in #3062
Full Changelog: v1.6.2...v1.7.2