DepthART | Scaling Foundation Monocular Depth to Tiny ModelsDepth Anything Rethought for Tiny Models<br>DepthART<br>ACMMM 2026<br>Scaling Foundation Monocular Depth to Tiny Models<br>Read paper Code🤗 Models
6.0M / 8.0M Relative-S / Metric-S parameters
0.92 ms DepthART-S 224 · A6000 TensorRT FP16
1088.9 DepthART-S 224 · A6000 FPS · TensorRT FP16
247.8 / 15.2 Orin NX TensorRT FP16 / Nano FP32 FPS
0.971 δ1 DepthART-L · NYUD v2 zero-shot
Videos from ADVIO datasets, captured on iPhoneRelative Depth in the Wild<br>DepthART processes handheld portrait video while preserving fine boundaries and accurate depth.<br>Train StationIndoorOfficeIndoorStreetOutdoorCommercial DistrictOutdoor
RGB<br>DepthART-L · Relative
Algorithm OverviewDepthART<br>Recent geometric foundation models have advanced monocular depth estimation, yet their benefits remain limited for tiny models. We present DepthART , a compact model designed for robust on-device depth estimation across diverse scenes. To address dataset-specific overfitting and unstable metric adaptation under camera shifts, DepthART combines bias-resistant data sampling with camera-conditioned fine-tuning that preserves the distilled encoder while adapting metric scale using camera intrinsics. These designs improve both cross-dataset generalization and metric depth prediction in capacity-constrained models.<br>Bias-resistant distillation Rebalance a 44M multi-source corpus, then distill DepthAnything v2-L.
Camera-conditioned fine-tuning Freeze the trunk and adapt scale using camera prompts and a multi-query head.
Multi-platform DeploymentAccuracy vs Speed
PlatformRTX A6000Jetson Orin NXJetson Nano 4GBSoon<br>Power modeMAXN20W15W10W<br>Compute precisionPyTorch FP32PyTorch AMPTensorRT FP32TensorRT FP16<br>Depth taskRelative depthMetric depth<br>DatasetNYUD v2KITTI<br>RTX A6000 PyTorch FP32Relative · NYUD v2
Loading benchmark data…<br>0.9000.9250.9500.9751.000Model latency (ms, log scale) →NYUD v2 δ1 ↑<br>Selected profileNo matching benchmark rows.<br>DepthART 224 × 224DepthART 448 × 448
DepthART scale: S B LComparison methods use affine-invariant δ1.
Self-collected ImageRelative Depth<br>Scene 01<br>01 / 23←→
Input · RGBTeacher · DepthAnything v2-L<br>Drag to compareMiDaS v3.1 · LeViTMiDaS v3.1 · Swin2-T
MiDaS v3.1 · LeViTDepthART-L
0102030405060708091011121314151617181920212223
Web VideoRelative Depth<br>Busy city street<br>01 / 6←→
Input · RGB videoPrediction · DepthART-L<br>❚❚0:000:00Synced playback<br>01Busy city street02Flower03Friends at a party04Parrot05University campus06Portrait
Robot deploymentMetric Point Cloud
DepthART-Metric-L Fine-tuned on NYUD v2 only
Cafe · Scene 1 Metric 3D reconstruction<br>▶Play scene<br>Cafe · Scene 1Cafe · Scene 2Corridor · Scene 1Corridor · Scene 2Corridor · Scene 4Home · Scene 2Home · Scene 3Home · Scene 4Market · Scene 1Market · Scene 2Market · Scene 3Office · Scene 3Office · Scene 4Office · Scene 6
Cite this workDepthART<br>Accepted to ACM Multimedia 2026. The final citation and public model links will be added with the camera-ready release.<br>Read on arXivGitHub repository
@inproceedings{depthart2026,<br>title = {DepthART: Scaling Foundation Monocular Depth to Tiny Models},<br>author = {Feng Xue and Wu Chen and Mingshuai Zhao and Guofeng Zhong and Anlong Ming and Haozhe Wang and Dianqiao Lei and Zhaowen Lin and Haiyang Zhang and Nicu Sebe},<br>booktitle = {ACM Multimedia},<br>year = {2026}