Sankar Ganesh, Madhu Vanthi (2025) Enhancing Monocular Depth Estimation with Self-Supervised Learning and Transformer-Based Architectures. Masters thesis, Dublin, National College of Ireland.
Preview |
PDF (Master of Science)
Download (147kB) | Preview |
Preview |
PDF (Configuration Manual)
Download (354kB) | Preview |
Abstract
This paper presents a self-supervised learning framework for monocular depth estimation in dynamic scenes using transformers. Accurate depth estimation in dynamic regions remains challenging due to spatial-temporal inconsistency and scene clutter. While recent approaches have made progress by separating static and dynamic regions and understanding object-level motion cues, architectures often rely on convolutional neural networks, which can limit the ability to capture long-term dependencies. This paper proposes a transformer- based architecture that models well for both static and dynamic scenes using global attention mechanisms. This approach starts off with an object-centric module for estimating the depth of moving regions under a rigid motion assumption, along with a scale alignment strategy to resolve scale inconsistencies across regions. The pseudo-labels generated from this step are then used for self-training a depth estimation network in an end-to-end manner. Experiments conducted on Cityscapes dataset demonstrate that this method proves to be better than other CNN-based counterparts in terms of performance, particularly in dynamic environments, highlighting the effectiveness of transformer-based representations for this task.
Actions (login required)
![]() |
View Item |
Tools
Tools