Skip to content

Releases: lateos-ai/reflex

v0.2.2-runpod-hub: hosted Hub icon

Choose a tag to compare

@leochong leochong released this 30 Sep 03:15

Hosted icon for the Runpod Hub listing; no engine or packaging changes.

  • .runpod/icon.jpg (1024x1024) added, and .runpod/hub.json's iconUrl now
    points at it (was an invalid TODO: placeholder that would fail validation).
  • .runpod/README.md records the verified Hub state (2026-09-30, via
    runpodctl): the repo is registered under lateos-ai and hub get
    resolves it, but it is not yet in the public catalog (hub list/search
    empty; the runpod.io/hub URL 404s) -- awaiting Runpod review.

The .runpod/Dockerfile build is unchanged from v0.2.1-runpod-hub
(portable PTX), so this release is just the icon + status fix for a new Hub
build to pick up.

v0.2.1-runpod-hub: fix Hub build GPU-arch mismatch

Choose a tag to compare

@leochong leochong released this 27 Sep 03:08

Fixes a real build/test failure found on the first Hub submission attempt (v0.2.0-runpod-hub).

What happened

Runpod's Hub build succeeded, but the automated test job failed: the worker landed on an RTX 2000 Ada Generation (sm_89) card, despite .runpod/hub.json's gpuIds exclusion list and .runpod/tests.json's gpuTypeId both trying to pin/exclude away from it. The sm_86-compiled reflex binary correctly detected the mismatch and panicked rather than silently misbehaving.

Fix

.runpod/Dockerfile now builds portable PTX (no REFLEX_CUDA_ARCH pin) instead of a fixed sm_86 cubin, so it JIT-compiles correctly on whatever GPU Runpod's Hub scheduler assigns. This trades away the zero-JIT-tax cold-start advantage for actual correctness, since a Hub listing (unlike a manually-deployed endpoint) has no reliable way to pin the GPU SKU.

Also fixes .runpod/hub.json's category field (language-models → language, matching the Hub UI's real accepted values).

See .runpod/README.md's "Real Hub build/test findings" section for full detail.

🤖 Generated with Claude Code

v0.2.0-runpod-hub: Runpod Hub listing

Choose a tag to compare

@leochong leochong released this 27 Sep 02:27

Adds an additive, queue-based Runpod Hub worker listing (.runpod/) so Reflex can appear in Runpod's public Hub catalog, alongside (not replacing) the existing, real-hardware-verified load-balancing deployment under serverless/runpod/.

What's in this release

  • Runpod Hub listing (.runpod/): a thin Python translation shim (handler.py) that spawns the same, unmodified reflex-openai-adapter binary and forwards jobs to its existing POST /v1/chat/completions — no engine logic reimplemented in Python. Real-hardware-verified via a temporary queue-based endpoint (correct end-to-end chat completion, GPU pin landed on RTX A4500 as configured). Not yet submitted to the Hub itself — see .runpod/README.md.
  • Runpod Serverless load-balancing deployment (serverless/runpod/): real-hardware-verified, ~3.4x faster cold start than Runpod's official worker-vllm on identical hardware (docs/serverless-cost-comparison.md).
  • Sidecar hardening: a real three-state /healthz (dead/loading/ready) backed by actual process liveness, a watchdog that fails the sidecar fast instead of degrading silently, and a GET /v1/models endpoint.
  • Always-warm AWS deployment guide, a sibling to the existing scale-to-zero pattern, for a public demo endpoint.
  • MoE per-expert LoRA support.

Known limitations

  • The Runpod Hub listing is non-streaming only (v1 scope).
  • .runpod/hub.json's iconUrl is a placeholder pending a real hosted icon asset.
  • Actual Hub submission (linking this repo in Runpod's console) is a separate, manual step not included in this release.

🤖 Generated with Claude Code