Repository navigation
Releases: Happenmass/duck-on-desk
Release list
Duck on Desk v1.2.0
Duck on Desk v1.2.0
第一课「行走」改为原样嵌入官方 mjlab 训练环境,应用内直接复现官方训练流程。
- 训练环境不再是自研的 CPU MuJoCo + 自定义 PPO,而是官方仓库
pollen-robotics/microduck_rl(commit53b8971,mjlab 1.3.0 / rsl_rl 5.0.1)的Mjlab-Velocity-Flat-MicroDuck任务:16 项奖励、域随机化、特权 critic、在线观测标准化、自适应 KL 的 PPO 与课程表全部沿用官方,策略从随机权重开始。 - 一套应用两种机器:Apple Silicon 上物理在 CPU(MuJoCo Warp 没有 Metal 后端)、网络在 MPS;NVIDIA 上走官方 CUDA。界面操作不变,仍由「用哪个设备训练」决定。
- 环境安装改为克隆钉死 commit 的官方仓库并按其
uv.lock执行uv sync --frozen,需要本机有 Git(macOS 装 Xcode 命令行工具,Windows 装 Git for Windows)。Windows CUDA 会额外把 torch 2.9.1 换成 cu128 wheel(PyPI 的 Windows torch 是纯 CPU)。 - 第 2 步的奖励编辑框对行走课改为
reward_weights权重字典,列出官方 16 项默认值;三个滑块对应track_linear_velocity、upright、action_rate_l2。action_rate_l2与head_pose_bias官方由课程表调整,一旦写出权重即固定,manifest 记录curricula_disabled。 - 「新训练的起点」与「PPO 更新方式」对行走课隐藏并固定为官方配方;倒立课程不受影响,继续在同一个环境里跑 CPU MuJoCo 路径。
- 行走默认 256 环境 × 24 步、学习率 0.001、seed 42、200 轮;CUDA 下「使用本机推荐参数」填回官方 4096 × 24。旧
workbench.json的行走设置一律忽略。 - 训练指标新增每回合平均时长、完成回合数、学习率、动作标准差、各损失项与官方逐项奖励分解;检查点为 rsl_rl 格式并附带环境步数与课程表进度,续训后轮数与步数连续。
- 官方导出的 ONNX(标准化已烘焙进图)由桌宠运行时零改动加载;训练前后的对比测试仍在应用自带的 CPU MuJoCo 环境里进行。
- MCP:
physics_backend新增"mjlab",行走草稿默认该值;lab_experiment_validate接受只声明reward_weights的 reward.py。
本机(Apple Silicon)实测:Warp CPU 物理约 640 env-steps/s,256 环境每轮约 11 s,200 轮约 37 分钟,只适合验证流程与配置;官方规模(4096 × 24 × 50000 轮)需在 NVIDIA 机器复现。CUDA 分支按官方脚本接入 cuda:0,尚未在实机验证。Windows 11 开启「智能应用控制」的机器会拦截 warp.dll(WinError 4551),需关闭该功能(不可逆)或改用 WSL2;两者都未接入应用。
实测数据与未验证项见 docs/diagnostics/mjlab-embedded-training-2026-09-11.md。
Duck on Desk v1.1.4
Duck on Desk v1.1.4
第二课从「静态保持 2 秒倒立」改为「翻倒后用脚蹬回来」的循环练习。
- 奖励整体重写。旧奖励的验收要求双脚离地,脚一沾地保持计时立即归零——而脚正是用来蹬回去的,这个动作在原设计里无法得分。新奖励取消保持计时器与角速度惩罚,脚触地不再扣分。
- 倒置程度的计分在 −0.8 处饱和。实测「折叠在自己头上」这种失败姿态比可持续的目标姿态更倒(−0.92 对 −0.80),继续为更大倒置加分会把失败排在目标之前。躯干高度是两者唯一的判别量(0.04 对 0.078),已作为计分因子。
- 新增恢复计分:每次真实跌倒后回到头部贴地给一次奖励,环境内以滞回判定,避免临界抖动被计入。只靠摆一个姿势不动的策略拿不到这一项。
- 环境新增参考状态初始化:从翻倒过程的随机时刻起步。此前只有「顶端平衡」和「完全趴地」两种状态会被采样,中间的小幅修正从未练到,约 96% 的样本来自躺地状态。
- 训练新增探索噪声保持步数。逐步独立噪声在多步蹬地动作上相互抵消,连贯的蹬地动作采不到。第二课默认保持 4 步。
- 第二课默认定时重置由 1500 步(30 秒)改为 120 步(2.4 秒),并默认开启随机起步。
- 验收指标改为循环口径:头部贴地时间占比、躺地占比、恢复次数、最长一次时长。旧的 2 秒静态保持无法登记这一行为。
第一课行走、第三课倒立进入与街舞课的奖励和验收标准均未改动。续训历史记录保留其原有奖励、重置规则与轮次长度。
本机 MPS 训练已复现目标行为,但 3 个种子中只有 1 个成功(约 140 万采样步、25 分钟)。成功的策略保持头部贴地、躯干高度 0.074、躯干不触地;把控制器切断 0.9 秒让它在重力下摔到 upright +0.03 后恢复控制,能爬回头部贴地并在结尾持续稳住。另外 2 个种子停在「折叠在自己头上」——躯干高度只有 0.050~0.054,摔倒后再也起不来。所以这是一个能训出来但不稳定的配方。
验收因此改为真实的击倒测试:评测中途切断控制器 0.9 秒,必须先真的摔下去(upright > -0.2),再回到头部贴地并稳住结尾 0.5 秒才算通过。只按「头朝下的时间占比」判定会把折叠姿态判为成功。
奖励另经 CEM-MPC 规划器从「近倒立」与「已经趴下」两个起点独立验证,最优解确实是目标行为(头部贴地时间 83.5% / 92.3%)。验收阈值仍是照规划器可达值设定的暂定值。CUDA 仍待 NVIDIA 实机验证。
分析、实测数据与方法教训见 docs/diagnostics/headstand-cycle-reward-2026-09-10.md。
Duck on Desk v1.1.3
Duck on Desk v1.1.3
机器人实验室的训练参数由用户决定,默认值不再作为强制范围。
- 移除新建训练、续训追加及累计训练轮数的人为上限,不再两小时自动停止。按填写的轮数训练,也可手动停止保存后继续。
- 移除同时练习的机器人数量、每轮采样步数、定时重置间隔的人为上限;界面、MCP 和训练程序使用相同规则。
- 放开学习率、目标速度和随机种子的原有范围限制。保留有效数字、正整数等必要校验;默认参数不变。
- 奖励权重改为数字输入,支持自定义数值;实验室预览不再把前进与转弯指令截断到旧范围。
- 支持单个样本的短轮次训练,修复小学习率或零学习率导致参数未变化时无法导出的问题。参数变化量仍如实记录,非有限数值继续报错。
- 保留 PPO 裁剪公式、官方网络结构和机器人模型定义的关节物理范围。
已验证真实 CPU 短训、检查点与 ONNX 导出,包括 64 个环境、300 步采样及单样本配置;这些验证不代表动作已学会,也不保证任意规模都能在当前硬件上运行。CUDA 训练仍待 NVIDIA 实机验证。
Duck on Desk v1.1.2
Duck on Desk v1.1.2
修复新机器首次训练提示 Microduck model / base policy unavailable; install model assets first 的问题。
- 训练开始前自动补齐固定版本的 Microduck 基础策略,保存到用户 Hugging Face 缓存。训练服务不再记住不存在的旧包内路径。
- “Python 环境”留空即可自动安装独立的 Python 3.12、PyTorch、MuJoCo 与 ONNX 依赖;也可点击“检查 / 安装训练环境”提前准备。已有自定义环境只检查,不修改。
- Windows x64 默认 CUDA,Apple Silicon 默认 MPS。CUDA 安装使用 PyTorch 2.9.1 / CUDA 12.8 wheel;启动前检查实际 GPU 计算。不可用时显示原因,不自动退回 CPU。
- 首次下载和环境检查显示进度,失败后可重试;训练的准备阶段也可以停止。MCP 新增环境安装与状态查询接口。
- 修复 Windows 默认编码无法读取中文训练配置和奖励脚本的问题,统一使用 UTF-8。
- 保留 1.1.1 的默认 30 秒重置、可配置重置时间与历史续训设置。
首次安装需要网络及可用磁盘空间,CUDA 依赖下载较大。CUDA 需要受支持的 NVIDIA 显卡和驱动;本版自动安装不代替显卡驱动安装。自动环境支持 Windows x64、Apple Silicon;Intel Mac / Windows ARM 用户可填写自行配置的 Python。CUDA 实际训练仍待 NVIDIA 机器验证,Windows CPU 验证不能代替它。
Duck on Desk v1.1.1
Duck on Desk 1.1.1
机器人实验室加入可配置的训练重置时间,并正式发布此前在本地验证的训练、续训和关节控制功能。
- 第二课「倒立保持」默认每 30 秒模拟时间回到近倒立初态。在「开始练习」步骤可设置 0~1000 秒,0 表示关闭定时重置。身体倾斜、脚触地时仍可以尝试蹬地恢复。
- 训练记录显示实际使用的重置规则。继续训练恢复权重、优化器、仿真和随机状态,沿用原记录的参数;修改界面时间只影响新训练。
- 四步强化学习教程、完整自定义奖励、奖励分项、策略/价值 loss 曲线,以及默认关闭的真实训练 3D 预览。
- 14 路关节输出中文映射、半透明模型和即时手动摆姿;镜头跟随机器人移动。实验室使用独立普通窗口,支持 Dock/任务栏切换。
- 两课使用官方同构策略网络;行走默认加载官方权重,倒立默认随机初始化。PPO 支持分批更新、自适应学习率和 MPS;CUDA 接口保留,尚待 NVIDIA 实机验证。
- Duck 右键动作菜单支持预置啄地及行走时随机动作。倒立基础练习仍为实验课程,未训练成功,不会自动替换桌宠行走或加入动作库。
- 安装包不复制模型权重;首次使用所需策略时下载固定版本到用户级 Hugging Face 缓存,之后复用。首次使用需要网络。
从「设置 → 关于 → 打开开发者模式」进入实验室。训练需要已有的 PyTorch、MuJoCo、ONNX 和 ONNX Runtime Python 环境。现有模型、训练记录及自定义奖励保留。
提供 macOS Apple Silicon/Intel 和 Windows arm64/x64 安装包。macOS 使用 ad-hoc 签名,未公证;Windows 窗口切换与 CUDA 的实机验证仍待完成。仿真训练和软件验证不代表实体机器人安全验收。
Duck on Desk v1.0.0
Duck on Desk 1.0.0
这次更新把机器人实验室带进正式安装版,并重新整理了「关于」页面的开发者入口和项目署名。
开发者模式与机器人实验室
- 在「设置 → 关于」开启「打开开发者模式」,进入独立的 3D 机器人实验室。开关默认关闭,保存后跨重启保留;窗口关掉后可从同一页面再次进入。
- 完整机器人、地面和轨道相机,支持运行、暂停、单步、重置以及 14 个关节的检查和手动控制。
- 可编辑奖励脚本与策略网络,设置 PPO 训练参数,查看实际采样步数和奖励曲线,停止任务并保留历史。
- 训练后可预览候选策略,完成原生与浏览器物理评测,再应用到普通脚桌宠;支持导出 ONNX 和恢复上一策略。
- Agent 管理页可以为 Claude Code、Codex CLI、OpenCode 安装 Robot Lab MCP。Coding Agent 可查询接口文档、版本化写入奖励/策略脚本、启动和管理真实训练任务。
训练需要自行选择装有 torch、mujoco、onnx、onnxruntime 的 Python 环境。物理使用 CPU MuJoCo,训练和 Python 评测推理可分别选择 MPS、CUDA、CPU;桌宠及实验室预览仍用 Web WASM。MPS 已实测,CUDA 待 NVIDIA 实机验证。第一版支持平地普通脚的 PPO 动作修正,尚未提供自动续训或任意 RL 算法插件。
关闭开发者模式会销毁实验室窗口;已启动的训练继续,由所属 App 或 MCP 进程管理。退出所属进程会停止其任务。默认不开启实验室,也不会自动训练或提交付费云作业。
关于页面
- 本项目作者、维护者和贡献者更新为 Happenmass。
- 上游 Clawd on Desk 作者、维护者及贡献者移入「上游项目与致谢」,保留 Pollen Robotics 来源及现有授权说明。
- 发布下载按平台分别获取并核对大小、SHA-512 和更新元数据,防止截断安装包进入更新通道。
验证与更新
本机已通过 MPS 20 轮、2,560 个真实物理步的训练、ONNX 导出、浏览器评测,以及隔离 Electron 虚拟桌宠中的实际加载、导出和回退。短训证明链路有效,不代表步态已经明显改善。Windows 和 CUDA 尚未进行本机硬件验收。
v0.2.3 用户可以在「设置 → 关于」检查更新。v0.2.2 及更早版本的更新地址有误,需要手动下载本次安装包;替换应用不会主动清空已有设置。macOS 安装包沿用 ad-hoc 签名,未作 Developer ID 公证。
本更新不包含实体机器人硬件安全验收。程序代码为 AGPL-3.0-only;机器人资产的单独授权和非商业使用边界见仓库 NOTICE.md。
Duck on Desk v0.2.3
v0.2.3
Checking for updates works. It never had, in any release of this fork.
- The update check asked GitHub for releases of
rullerzhou-afk/duck-on-desk.
That repository does not exist — the owner was inherited from the upstream
project this fork is derived from and left hardcoded, while the builds were
published here. Every check got a 404, which the app reported as "Duck 更新时
遇到未知错误 / No releases found". Checks now go toHappenmass/duck-on-desk,
where the releases actually are - Settings -> About linked the same nonexistent repository. It now points at
this fork's source, which is what a user needs for the build they are running.
The upstream author credit, copyright and AGPL-3.0 license are unchanged - A test now ties the update URL to the
build.publishtarget in
package.json, so the two cannot drift apart again
Updating from 0.2.2 or earlier has to be done once by hand. The fix lives in
the new build, so an installed 0.2.2 still asks the dead repository and still
reports the same error — it cannot see this release. Download the installer
below once; automatic checks work from 0.2.3 onward.
No behaviour, rendering or physics changes in this release.
macOS (arm64/x64 dmg+zip) and Windows (x64/arm64 NSIS).
Not real-machine validated on Windows or Linux for this release; the update
check was verified against the live GitHub API and the smoke pass was done on
macOS arm64.
Full Changelog: v0.2.2...v0.2.3
Duck on Desk v0.2.2
v0.2.2
The 3D pet costs about half the CPU it used to while it is just standing there.
- The renderer ran at whatever the display offered — 120 fps on a ProMotion
panel — and drew all 70 rig meshes a second time every frame for a real
shadow map. Both robot runtimes now render at 30 fps, and the shadow map is
replaced by a painted blob under the feet at the same size and weight - Measured on an M-series Mac with a 120 Hz display, the duck idling: total app
CPU 53.8% of a core -> 26.5% (2 samples before, 10 after; before was tight at
53.3-54.2%, after spreads 23.2-30.5% because the remaining cost now tracks
what the duck is doing rather than the frame rate). The GPU process alone
falls 25.5% -> 8.4%. The win is smaller on a 60 Hz display, where the frame
cap has half as much to remove. Memory is unchanged - The physics and the control policy were never the expensive part: MuJoCo plus
the ONNX policy measure ~1.1% of a core together, so nothing about how the
duck balances, walks, falls or gets picked up has changed - The camera's ease-down after the duck is dropped is now driven by elapsed
time instead of by frame count. It had been a flat factor per rendered frame,
which silently tied the settle to the display's refresh rate and would have
stretched it four-fold under the new frame cap
macOS (arm64/x64 dmg+zip) and Windows (x64/arm64 NSIS).
Not real-machine validated on Windows or Linux for this release; the
measurements and the smoke pass were done on macOS arm64.
Full Changelog: v0.2.1...v0.2.2
Duck on Desk v0.2.1
v0.2.1
A fix for the duck freezing after a restart, and new step sounds.
- The 3D duck no longer stands still after the app restarts or the renderer reloads while a session is already working: main now re-sends the current state once the renderer's MuJoCo/ONNX runtime is listening, instead of the on-load message being dropped and the duck waiting on a 9.75 s retry that could be dropped too
- Steps and landings are now single moves of a real hobby servo ("Small servo (arduino)" by gpag1, CC0 on freesound) instead of the Kenney footstep and thump samples; NOTICE and the READMEs credit the new source, and the "Duck footsteps" setting describes the servo whine
- A test pins that every servo take the audio adapter can pick is shipped, since a missing take only logs a warning
macOS (arm64/x64 dmg+zip) and Windows (x64/arm64 NSIS).
Full Changelog: v0.2.0...v0.2.1
Duck on Desk v0.2.0
v0.2.0
Adds Reachy Mini as a second robot: a physical Reachy Mini on the LAN, a 3D virtual one on the desktop, or both mirrored.
- Physical Reachy Mini support: mDNS or manual host discovery, auto-connect, and a "Robot" section in Settings and the tray menu
- Robot wake/sleep is driven by the semantic state machine and calls the official daemon moves (
wake_up,goto_sleep), with motor torque re-armed on wake and no retry of an unacknowledged physical move - Robot volume is mapped to the hardware's dB-linear scale, fixing a bug where any non-maximum volume made the robot inaudible; the official cue sounds are uploaded to the robot
- State-driven physical expressions replay the official recorded emotion and dance libraries through a hardware-screened whitelist, plus continuous liveliness (breathing, antenna sway, head glances with lagged body-yaw follow)
- "Robot motion" menu toggle turns physical expressions and liveliness off without disconnecting
- Virtual Reachy Mini in the 3D renderer, with the official rig, Draco decoder and Apache-2.0 cue sounds bundled offline; the renderer mirrors the physical robot's reported pose when one is connected
- Renderer now reads the current pet prefs on every reload, so switching robots never restores stale settings
- Source layout split into
src/state/,src/shell/,src/robots/duck/andsrc/robots/reachy/, with a guard test that every__dirnamepath literal undersrc/resolves - New guides:
docs/guides/reachy-mini.mdand the 2026-09-06 Reachy control audit; NOTICE credits the Pollen Robotics assets
macOS (arm64/x64 dmg+zip) and Windows (x64/arm64 NSIS).
Full Changelog: v0.1.1...v0.2.0