English

Model ReleasesFoundationPose

FoundationPose: A Transformer-based Model for 6-DoF Object Pose Estimation and Tracking

FoundationPose is a foundation model designed for 6-DoF (Degrees of Freedom) object pose estimation and tracking. The model allows for instant application to novel objects at test-time, provided a CAD model of the object is available, eliminating the need for fine-tuning.

The architecture is Transformer-based and consists of two separate networks: a refinement net and a score net. The refinement network extracts feature maps from two RGBD input branches using a shared CNN encoder to predict translation and rotation updates. The score net employs a pose ranking encoder to utilize global image context for comparison.

The model was trained on large-scale 3D databases, including Objaverse and Google Scanned Objects (GSO). FoundationPose is compatible with NVIDIA hardware, including NVIDIA Jetson devices, and is optimized for TensorRT. It is designed for deployment on edge devices and can be integrated with the NVIDIA Triton Inference Server for efficient AI application pipelines.

Sources

  1. nvidia/foundationpose (HF: NVIDIA, 2026-09-15)
  2. GitHub