NVIDIA has released RT-DETR Hand Detection v1.0, a model that adapts the Real-Time Detection Transformer (RT-DETR) architecture to detect and localize hands directly in RGB images. The model combines a convolutional backbone with an efficient hybrid encoder and a TransformerDecoder head to predict 2D hand bounding boxes and class confidence scores in a single end-to-end pass.
Model ReleasesNvidiaRT-DETR Hand Detection
NVIDIA Releases RT-DETR Hand Detection v1.0 for Real-Time Hand Localization
Developed as part of RT-DETRv2-S, the model contains approximately 20 million parameters. It is designed for real-time human-computer interaction (HCI) and egocentric computer vision pipelines. The resulting 2D hand bounding boxes serve as precise regions-of-interest (ROIs) for downstream models, such as EgoHand, to predict 3D hand keypoints or mesh.
The model is trained on a mix of synthetic and real-world hand motion image and video data. It is optimized to run on NVIDIA GPU-accelerated systems, leveraging hardware such as NVIDIA Lovelace microarchitecture to achieve faster training and inference times compared to CPU-only solutions.
NVIDIA states that this model is for demonstration purposes and is not intended for production usage. It is released under the NVIDIA Software and Model Evaluation License.
Sources
- nvidia/RT-DETR-Hand-Detector-v1 (HF: NVIDIA, 2026-09-15)
- NVIDIA Software and Model Evaluation License