DiffuPose: Monocular 3D Human Pose Estimation via Denoising Diffusion Probabilistic Model

Choi, Jeongjun; Shim, Dongseok; Kim, H. Jin

Computer Science > Computer Vision and Pattern Recognition

arXiv:2212.02796 (cs)

[Submitted on 6 Dec 2022 (v1), last revised 3 Aug 2023 (this version, v3)]

Title:DiffuPose: Monocular 3D Human Pose Estimation via Denoising Diffusion Probabilistic Model

Authors:Jeongjun Choi, Dongseok Shim, H. Jin Kim

View PDF

Abstract:Thanks to the development of 2D keypoint detectors, monocular 3D human pose estimation (HPE) via 2D-to-3D uplifting approaches have achieved remarkable improvements. Still, monocular 3D HPE is a challenging problem due to the inherent depth ambiguities and occlusions. To handle this problem, many previous works exploit temporal information to mitigate such difficulties. However, there are many real-world applications where frame sequences are not accessible. This paper focuses on reconstructing a 3D pose from a single 2D keypoint detection. Rather than exploiting temporal information, we alleviate the depth ambiguity by generating multiple 3D pose candidates which can be mapped to an identical 2D keypoint. We build a novel diffusion-based framework to effectively sample diverse 3D poses from an off-the-shelf 2D detector. By considering the correlation between human joints by replacing the conventional denoising U-Net with graph convolutional network, our approach accomplishes further performance improvements. We evaluate our method on the widely adopted Human3.6M and HumanEva-I datasets. Comprehensive experiments are conducted to prove the efficacy of the proposed method, and they confirm that our model outperforms state-of-the-art multi-hypothesis 3D HPE methods.

Comments:	Accepted to IROS 2023. First two authors contributed equally
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2212.02796 [cs.CV]
	(or arXiv:2212.02796v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2212.02796

Submission history

From: Jeongjun Choi [view email]
[v1] Tue, 6 Dec 2022 07:22:20 UTC (1,696 KB)
[v2] Fri, 9 Dec 2022 03:15:54 UTC (1,696 KB)
[v3] Thu, 3 Aug 2023 09:03:30 UTC (1,586 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:DiffuPose: Monocular 3D Human Pose Estimation via Denoising Diffusion Probabilistic Model

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:DiffuPose: Monocular 3D Human Pose Estimation via Denoising Diffusion Probabilistic Model

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators