English

Model ReleasesNvidiaRT-DETR Hand Detection

NVIDIA Releases RT-DETR Hand Detection v1.0 for Real-Time Hand Localization

NVIDIA has released RT-DETR Hand Detection v1.0, a model that adapts the Real-Time Detection Transformer (RT-DETR) architecture to detect and localize hands directly in RGB images. The model combines a convolutional backbone with an efficient hybrid encoder and a TransformerDecoder head to predict 2D hand bounding boxes and class confidence scores in a single end-to-end pass.

Developed as part of RT-DETRv2-S, the model contains approximately 20 million parameters. It is designed for real-time human-computer interaction (HCI) and egocentric computer vision pipelines. The resulting 2D hand bounding boxes serve as precise regions-of-interest (ROIs) for downstream models, such as EgoHand, to predict 3D hand keypoints or mesh.

The model is trained on a mix of synthetic and real-world hand motion image and video data. It is optimized to run on NVIDIA GPU-accelerated systems, leveraging hardware such as NVIDIA Lovelace microarchitecture to achieve faster training and inference times compared to CPU-only solutions.

NVIDIA states that this model is for demonstration purposes and is not intended for production usage. It is released under the NVIDIA Software and Model Evaluation License.

Sources

  1. nvidia/RT-DETR-Hand-Detector-v1 (HF: NVIDIA, 2026-09-15)
  2. NVIDIA Software and Model Evaluation License