Skip to content

Repository files navigation

Knowledge-embedded Transformer for 3D Human Pose Estimation

Installation instructions

  • Python 3.8
conda create --no-default-packages -n smpl python=3.8
conda activate smpl

packages

conda install pytorch==1.12.0 torchvision==0.13.0 torchaudio==0.12.0 cudatoolkit=11.3 -c pytorch
pip install "git+https://github.com/polmorenoc/opendr.git#subdirectory=opendr"
pip install "git+https://github.com/daniilidis-group/neural_renderer"
  • other packages listed in requirements.txt
pip install -r requirements.txt

necessary files

mesh_downsampling.npz & DensePose UV data

  • Run the following script to fetch mesh_downsampling.npz & DensePose UV data from other repositories.
bash fetch_data.sh

SMPL model files

Fetch preprocessed data from SPIN.

Fetch final_fits data from SPIN. [important note: using EFT fits for training is much better. Compatible npz files are available here]

Run demo code

python3 demo.py --checkpoint=data/pretrained_model/model_checkpoint.pt --img_file examples/COCO_val2014_000000019667.jpg

Run evaluation code

3DPW

Run the evaluation code. Using --dataset to specify the evaluation dataset.

# Example usage:
# 3DPW
python3 eval.py --checkpoint=data/pretrained_model/model_checkpoint.pt --dataset=3dpw --log_freq=20

Run training code

To perform training, we need to collect preprocessed files of training datasets first. This paper uses EFT fits. Compatible data is available here.

The preprocessed labels have the same format as SPIN and can be retrieved from here. Please refer to SPIN for more details about data preprocessing.

Similar to PyMAF, this model is trained on Human3.6M at the first stage and then trained on the mixture of both 2D and 3D datasets at the second stage. Example usage:

# training on COCO
CUDA_VISIBLE_DEVICES=0 python3 train.py --regressor pymaf_net --single_dataset --misc TRAIN.BATCH_SIZE 64
# training on mixed datasets
CUDA_VISIBLE_DEVICES=0 python3 train.py --regressor pymaf_net --pretrained_checkpoint path/to/checkpoint_file.pt --misc TRAIN.BATCH_SIZE 64

Running the above commands will use Human3.6M or mixed datasets for training, respectively. We can monitor the training process by setting up a TensorBoard at the directory ./logs.

Acknowledgments

The code is developed upon the following projects. Many thanks to their contributions.

Citation

If you find our code or paper helps, please consider citing:

@article{
  title={Knowledge-Embedded Transformer for 3-D Human Pose Estimation},
  author={Shu Chen, Ying He},
  journal={IEEE Transactions on Instrumentation and Measurement, 74 (2025), Article 5031811},
  year={2025}
}

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages