Multimodal industrial robot copilot prototype using RGB labels/images, thermal labels/images, radar point clouds, and metadata to produce world state, risk assessment, robot planning, scene graph context, and an explainable Gradio dashboard.
./install.sh
./start.shOpen the URL printed by start.sh, usually http://localhost:7860.
The top stream plays all frames as a browser-side canvas stream. The lower controls
inspect one selected frame through the full copilot pipeline, including boxed
RGB/thermal views, robot command, world state, risk decision, scene, plan, and
explanation. The selected RGB frame also overlays the estimated desired robot
path from Qwen when available, or a local aisle-based fallback path otherwise.
Limit the stream/selector with DEMO_FRAME_COUNT=10 if needed.
Use Play Model-Synced in the lower controls when you want each inspected
frame to render only after Qwen/control output is complete.
To show Qwen output in sync with the stream, precompute a lightweight cache:
VLM_BACKEND=qwen \
QWEN_VLM_BASE_URL=http://localhost:8000/v1 \
LOCATOR_BACKEND=labels \
python scripts/cache_qwen_outputs.py --frame-count 50Then start the app. The stream will show cached Qwen scene/risk/action/human
status for frames present in cache/qwen_stream.json.
Run the default five-frame check:
./scripts/run_smoke_tests.shRun a larger check:
FRAME_COUNT=20 ./scripts/run_smoke_tests.shThe main demo commands for benchmark, caching, Qwen, locate-anything.cpp, and
Gradio are collected in docs/COMMITTEE_RUNBOOK.md. A PDF copy can be generated
from the same file with:
pandoc docs/COMMITTEE_RUNBOOK.md -o docs/COMMITTEE_RUNBOOK.pdf --pdf-engine=xelatexThe AMD-safe final-demo default is:
VLM_BACKEND=heuristic
LOCATOR_BACKEND=labels
./start.shThis uses the deterministic local pipeline and dataset RGB/thermal labels. It is the safest mode when only AMD resources are available and LocateAnything has not been validated on ROCm.
If GroundingDINO or NVIDIA LocateAnything are not reliable on the AMD VM, use the live YOLO backend:
python -m pip install -r requirements-yolo.txt
VLM_BACKEND=heuristic \
LOCATOR_BACKEND=yolo_live \
YOLO_LIVE_MODEL=yolo11n.pt \
./start.shyolo_live runs Ultralytics YOLO on the RGB image and returns the same
located_objects shape as the other locators. Stock YOLO maps person to
human, so it is useful for live human detection and steering boxes. It is not
open-vocabulary; industrial classes such as control_panel, workbench, or
industrial_machine require a custom-trained YOLO model or a separate
open-vocabulary detector.
When LOCATOR_BACKEND=yolo_live, fusion can also use a live YOLO human box to
set world["human"]["present"] if no dataset RGB human label is available.
Distance remains dataset-derived when available; live YOLO detections currently
report distance: null.
Serve Qwen on the AMD resource using the helper:
./scripts/start_qwen_vllm.shIf the AMD image has the known vLLM dependency mismatch, the helper prints the environment alignment command. You can run the bundled alignment script explicitly:
./scripts/fix_amd_vllm_env.shThis pins the versions that avoid the observed huggingface-hub>=0.34,<1.0
and prometheus_fastapi_instrumentator / FastAPI route crashes.
Then run the app in another terminal:
./start.shstart.sh checks http://localhost:8000/v1/models; if Qwen is available it uses VLM_BACKEND=qwen, otherwise it falls back to the heuristic scene backend. Localization remains label-based unless changed.
Manual Qwen app mode:
VLM_BACKEND=qwen QWEN_VLM_BASE_URL=http://localhost:8000/v1 LOCATOR_BACKEND=labels ./start.shThe code supports nvidia/LocateAnything-3B, but the final AMD-only demo should keep:
LOCATOR_BACKEND=labelsuntil nvidia/LocateAnything-3B is validated on the AMD ROCm stack. The model is available on Hugging Face and is officially documented with NVIDIA-oriented hardware/runtime guidance.
If a compatible endpoint is available:
LOCATOR_BACKEND=nvidia_vllm \
NVIDIA_LOCATE_ANYTHING_BASE_URL=http://localhost:8001/v1 \
NVIDIA_LOCATE_ANYTHING_MODEL=nvidia/LocateAnything-3B \
./start.shIf local Transformers inference is validated:
LOCATOR_BACKEND=nvidia_transformers \
NVIDIA_LOCATE_ANYTHING_ATTN_IMPLEMENTATION=sdpa \
./start.shThe project also supports the mudler/locate-anything.cpp CLI backend:
Install/build example:
cd /workspace
git clone --recursive https://github.com/mudler/locate-anything.cpp
cd locate-anything.cpp
cmake -B build -DLA_BUILD_CLI=ON
cmake --build build -j
hf download mudler/locate-anything.cpp-gguf locate-anything-q8_0.gguf --local-dir modelsDirect CLI smoke test:
/workspace/locate-anything.cpp/build/examples/cli/locate-anything-cli detect \
--model /workspace/locate-anything.cpp/models/locate-anything-q8_0.gguf \
--input /workspace/AMD_Hackathon/data/industrial_subset/05_rgb/calibrated/013342.jpg \
--prompt "Locate all the instances that matches the following description: human</c>industrial machine</c>workbench</c>storage box." \
--mode hybrid \
--output /tmp/la_boxes.jsonApp mode:
LOCATOR_BACKEND=locate_anything_cpp \
LOCATE_ANYTHING_CPP_BIN=/path/to/locate-anything-cli \
LOCATE_ANYTHING_CPP_MODEL=/path/to/locate-anything-q8_0.gguf \
LOCATE_ANYTHING_CPP_MODE=hybrid \
./start.shThe CLI prompt is built from config/tracked_objects.json; categories are
separated with </c> as required by locate-anything.cpp. Use
LOCATE_ANYTHING_CPP_STRICT=1 to fail instead of falling back to dataset labels.
For repeated local runs, start the lightweight wrapper service. It caches detections by model, mode, image path, image mtime, and prompt:
python scripts/locate_anything_cpp_server.py \
--bin /workspace/locate-anything.cpp/build/examples/cli/locate-anything-cli \
--model /workspace/locate-anything.cpp/models/locate-anything-q8_0.gguf \
--mode fast \
--host 127.0.0.1 \
--port 8188Then use:
LOCATOR_BACKEND=locate_anything_cpp_http \
LOCATE_ANYTHING_CPP_URL=http://127.0.0.1:8188/detect \
./start.shThis wrapper avoids recomputing frames it has already seen. For maximum throughput on never-before-seen frames, use a native C-API server that keeps the GGUF model loaded in memory; the current wrapper is an integration-safe bridge around the CLI.
To cache Qwen plus locate-anything.cpp overlays for the top stream:
VLM_BACKEND=qwen \
QWEN_VLM_BASE_URL=http://localhost:8000/v1 \
LOCATOR_BACKEND=locate_anything_cpp_http \
LOCATE_ANYTHING_CPP_URL=http://127.0.0.1:8188/detect \
python scripts/cache_qwen_outputs.py --frame-count 621The cache script is resumable: it writes cache/qwen_stream.json after every
frame and skips frames already present in the cache. Use --force to recompute
everything.
Then start with the same locator settings:
DEMO_FRAME_COUNT=621 \
VLM_BACKEND=qwen \
QWEN_VLM_BASE_URL=http://localhost:8000/v1 \
LOCATOR_BACKEND=locate_anything_cpp_http \
LOCATE_ANYTHING_CPP_URL=http://127.0.0.1:8188/detect \
GRADIO_SHARE=1 \
./start.shBenchmark latency and GPU snapshots:
VLM_BACKEND=qwen \
QWEN_VLM_BASE_URL=http://localhost:8000/v1 \
LOCATOR_BACKEND=locate_anything_cpp_http \
LOCATE_ANYTHING_CPP_URL=http://127.0.0.1:8188/detect \
python scripts/benchmark_pipeline.py --frame-count 10Outputs are written to outputs/benchmark_pipeline.json and
outputs/benchmark_pipeline.csv. The benchmark reports cold locator latency,
cached locator latency, warm pipeline latency, and estimated cold end-to-end
latency. By default it asks the HTTP wrapper to bypass its cache for the first
locator timing per frame; add --use-existing-locator-cache if you want to
measure a fully warm run.
GroundingDINO can be used as a LocateAnything-style open-vocabulary locator:
VLM_BACKEND=qwen \
QWEN_VLM_BASE_URL=http://localhost:8000/v1 \
LOCATOR_BACKEND=grounding_dino \
GROUNDING_DINO_MODEL=IDEA-Research/grounding-dino-tiny \
./start.shBounding boxes from GroundingDINO appear on inspected RGB frames with DINO
labels. Qwen boxes appear with Qwen labels when Qwen returns approximate
located_objects; dataset RGB/thermal label boxes are always available in
LOCATOR_BACKEND=labels.
GroundingDINO is prompted for the tracked industrial objects: human, industrial machine, workbench, pipe, cabinet, storage box, control panel, forklift, and robot. The robot command panel labels its source separately: scene output comes from Qwen/heuristics, while command/steering comes from the rule-based control algorithm.
Tracked detection/VLM labels are configured in config/tracked_objects.json.
Add or remove labels in track_objects to change the object prompts used by
Qwen and open-vocabulary locators. Dataset-label fallback still only returns
objects that exist in the dataset annotations.
The scene and plan JSON include:
desired_path_source:qwen_vlmorheuristic_floor_projection.floor_region_source:qwen_vlmorheuristic_floor_projection.
The RGB overlay tints the estimated floor region and clips the speed-colored path ribbon to that floor polygon.
app.py: Gradio dashboard.src/pipeline.py: end-to-end orchestration.src/qwen_vlm.py: Qwen OpenAI-compatible scene backend.src/locate_anything.py: LocateAnything adapter plus label fallback.tests/: pipeline, module, and backend checks.analysis/: dataset analysis and CSV generation utilities.outputs/: generated analysis CSVs.docs/PROJECT_SPEC.md: project specification and architecture notes.scripts/check_env.py: environment and dataset checks.scripts/run_smoke_tests.sh: compile, environment, and multi-frame pipeline smoke test..env.example: runtime configuration template.