Authors :
Reedhi Shukla; Sampath Kumar P.; Kamini J.; Narender B.
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/rm3b9k9n
Scribd :
https://tinyurl.com/3vbpkaxe
DOI :
https://doi.org/10.38124/ijisrt/26jul153
Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.
Abstract :
The Segment Anything Model (SAM) has shown very good generalisation across many image types, but when we applied it directly to drone imagery, we found it does not work well — the top-down geometry, spatial scale, and object shapes in aerial data are quite different from the natural images SAM was trained on. Based on this, we designed SAMD (Segment Anything Model for Drone data), which combines a U-Net with a ResNet-34 encoder and SAM into a single pipeline. The U-Net is fine-tuned on drone imagery and produces pixel-wise masks for buildings and trees; these masks are then used as spatial prompts for SAM, which refines the boundaries and converts everything into georeferenced vector shapefiles. Building heights are derived from DSM-DMT differencing, and the final output is a 3D city model rendered in CesiumJS. We tested SAMD on drone imagery of Kanchipuram city, Tamil Nadu, India. Building extraction reached 85% F1 and tree delineation reached 80% F1. SAMD performed better than standalone SAM on drone data and is also slightly better than U-Net alone on boundary precision.
Keywords :
SAM; U-Net; ResNet-34; Drone Imagery; UAV Data; Building Extraction; Tree Extraction, Urban Mapping.
References :
- Ronneberger, O., Fischer, P. and Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. MICCAI 2015, LNCS 9351, 234–241.
- Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 4015-4026)..
- Song, J., Zhu, A.-X., & Zhu, Y. (2023). Transformer-Based Semantic Segmentation for Extraction of Building Footprints from Very-High-Resolution Images. Sensors, 23(11), 5166. https://doi.org/10.3390/s23115166.
- Weinstein, B. G., Marconi, S., Bohlman, S., Zare, A., & White, E. (2019). Individual Tree-Crown Detection in RGB Imagery Using Semi-Supervised Deep Learning Neural Networks. Remote Sensing, 11(11), 1309. https://doi.org/10.3390/rs11111309.
- Keyan Chen , Chenyang Liu , Hao Chen , Haotian Zhang , Wenyuan Li , Zhengxia Zou ,Zhenwei Shi (2024), "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation Based on Visual Foundation Model," in IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1-17, 2024, doi: 10.1109/TGRS.2024.3356074.
- Wu, Q., & Osco, L. (2023). samgeo: A Python package for segmenting geospatial data with the Segment Anything Model (SAM). Journal of Open Source Software, 8(89), 5663. https://doi.org/10.21105/joss.05663.
- Zhao, X., et al. (2023). GeoSAM: Fine-tuning SAM for geographic remote sensing image segmentation. arXiv:2308.09634.
- Gerke, M. and Nex, F. (2020). 3D building reconstruction from UAV images. In: Handbook of UAV Applications (Springer).
- Bittner, K., et al. (2018). Building footprint extraction from VHR remote sensing images combined with normalized DSMs. IEEE JSTARS, 11(8), 2615–2629.
The Segment Anything Model (SAM) has shown very good generalisation across many image types, but when we applied it directly to drone imagery, we found it does not work well — the top-down geometry, spatial scale, and object shapes in aerial data are quite different from the natural images SAM was trained on. Based on this, we designed SAMD (Segment Anything Model for Drone data), which combines a U-Net with a ResNet-34 encoder and SAM into a single pipeline. The U-Net is fine-tuned on drone imagery and produces pixel-wise masks for buildings and trees; these masks are then used as spatial prompts for SAM, which refines the boundaries and converts everything into georeferenced vector shapefiles. Building heights are derived from DSM-DMT differencing, and the final output is a 3D city model rendered in CesiumJS. We tested SAMD on drone imagery of Kanchipuram city, Tamil Nadu, India. Building extraction reached 85% F1 and tree delineation reached 80% F1. SAMD performed better than standalone SAM on drone data and is also slightly better than U-Net alone on boundary precision.
Keywords :
SAM; U-Net; ResNet-34; Drone Imagery; UAV Data; Building Extraction; Tree Extraction, Urban Mapping.