Repository navigation
TexMotion v1.3.0: WHAM+MediaPipe 3D Fusion & UI Overhaul
TexMotion v1.3.0 Release Notes
日本語 (Japanese)
1. WHAM と MediaPipe による 3次元姿勢統合パイプライン
動画からの姿勢抽出において、カメラ座標系での大域移動推定に強い WHAM と、画像平面上での関節点精度が高い MediaPipe を統合した推論パイプラインを導入した。
- ハイブリッド 3D 融合処理: WHAM による深度方向の位置推定と、MediaPipe による2次元座標および信頼度スコアを重み付け結合し、関節の3次元回転を再構成する。
- 脚部交差と遮蔽の耐性向上: ダンスや方向転換の際に脚が前後に重なる場面でも、左右の足の取り違えや深度反転による姿勢の崩れを抑制する。
- GPU による推論高速化: PyTorch バックエンドによる GPU 推論をサポートし、長尺動画における抽出時間を短縮した。
2. 可変アスペクト比プレビューとオーバーレイ表示
入力動画のアスペクト比に応じた描画処理を刷新した。
- 比率維持スケーリング: スマートフォンの縦型動画(9:16)、正方形(1:1)、横長動画のそれぞれに対し、余白(レターボックス)を自動計算して RenderTexture を生成することで、画像の歪みや端部の切り取りを防止する。
- 2D/3D 並列オーバーレイ: 検出された関節ランドマークを重畳描画した元映像と、推定姿勢を反映した 3D アバターを並列表示し、トラッキング結果を直感的に照合できる。
3. タイムラインエディタの操作性向上
タイムライン上の再生制御と描画安定性を強化した。
- 低速再生の拡張: 再生速度の設定幅を 0.1x から 2.0x まで広げ、0.1x 刻みでのスロー再生に対応した。
- HUD コントロールの追加: タイムラインヘッダーおよびプレビュー画面の HUD 上に、再生・一時停止を切り替えるボタンを配置した。
- 描画例外の解消: スクラビング操作中に発生していた null 参照および GUILayout 状態の不整合を修正した。
4. 設定画面の3タブ再編
設定ウィンドウの構造を見直し、用途別に3つのサブタブへ整理した。
- 一般 (General): 表示言語、ログ出力、UI 設定、更新確認を配置。
- モーション生成 (Text-to-Motion): Kimodo 拡散モデル、生成パラメータ、Hugging Face からのモデルダウンロード機能を配置。
- 動画モーション (Video-to-Motion): 姿勢推定バックエンド(WHAM, MediaPipe, RTMPose, HMR2)の選択、Python 仮想環境の状態確認、モデルアセット管理を配置。
- タブバーの固定表示: スクロール時にもタブ切り替えボタンが隠れない固定ヘッダー構造に変更した。
English
1. WHAM + MediaPipe 3D Pose Fusion Pipeline
Integrated WHAM and MediaPipe into a joint estimation pipeline, pairing WHAM's global trajectory estimation with MediaPipe's high-precision 2D planar landmarks:
- Hybrid 3D Fusion: Blends WHAM depth estimations with MediaPipe 2D coordinates and confidence weights to reconstruct accurate 3D joint rotations.
- Occlusion & Leg Crossing Robustness: Suppresses limb-flipping and depth-inversion artifacts when legs cross during turning, dancing, or rapid steps.
- GPU Inference Acceleration: Adds PyTorch-based GPU inference paths to accelerate extraction throughput on long video clips.
2. Dynamic Aspect-Ratio Viewport & Overlay Synchronization
Refactored preview rendering to preserve source video geometry:
- Aspect-Ratio Fitting: Dynamically computes letterboxing offsets for vertical mobile clips (9:16), square feeds (1:1), and wide formats, preventing dimensional distortion and unintended frame cropping.
- Synchronized 2D/3D Comparison: Displays detected 2D skeletal landmarks directly against the real-time 3D avatar viewport for immediate tracking evaluation.
3. Timeline Editor Navigation Enhancements
Refined playback manipulation and layout stability:
- Expanded Low-Speed Playback: Adds variable playback speeds ranging from 0.1x to 2.0x in 0.1x increments.
- HUD Playback Toggles: Places dedicated Play/Pause controls directly on the viewport HUD and the timeline header.
- Layout Exception Fixes: Eliminates null-reference exceptions and GUILayout mismatch state errors encountered during rapid timeline scrubbing.
4. Settings Interface Reorganization
Rebuilt the Settings interface into three dedicated workflow views:
- General: Language selection, logging verbosity, UI behaviors, and release check updates.
- Text-to-Motion: Kimodo diffusion model configs, inference parameters, and Hugging Face asset downloads.
- Video-to-Motion: Pose estimation engine selection (WHAM, MediaPipe, RTMPose, HMR2), Python environment status, and dependency management.
- Pinned Tab Header: Fixes the category bar at the top of the window to maintain navigation access during vertical scrolling.
📦 Installation
Download TexMotion.unitypackage below and import it into your Unity project (Unity 2022.3.x recommended, VRChat Avatar project supported).
Or install directly via Unity Package Manager using Git URL:
https://github.com/k0ta0uchi/TexMotion.git