Abstract:To address the challenges of limited accuracy of conventional visual measurement methods for interior doors and windows in buildings, as well as the high computational complexity of laser point cloud data processing, a vision-based measurement method is proposed. This approach integrates visual and laser technologies and employs deep learning models to optimize and streamline the image processing pipeline. Multi-view images with depth information are captured through reliance on a self-developed measuring equipment. Geometric correction and image stitching are performed using depth information and pose data to construct panoramas of walls containing the target doors and windows. To enhance measurement accuracy while reducing computational load, the framework utilizes a YOLOv11-CUA (YOLOv11n-C3k2_UIB-ADown) for object detection, combined with the Line Segment Detector (LSD) algorithm to achieve precise measurement within the target regions. The improved model of YOLOv11 includes the introduction of the C3K2_UIB module to strengthen multi-scale feature extraction and improve detection of small doors and windows. The ADown module employs a lightweight design to reduce computational overhead and replaces conventional downsampling convolutions to strengthen feature representation in key regions, thereby reducing information loss in low-contrast scenarios. Evaluated on a dedicated door-window dataset, the system achieves a detection accuracy of 97.5%, a mAP?? of 97.2%, and contains only 2.0M parameters. Compared with the baseline YOLOv11n, the proposed model shows an improvement of 4.7% in accuracy, 4.8% in mAP??, and a reduction of 0.6M parameters. The resulting dimensional measurements exhibit an average error of 5.63 mm, with maximum deviations controlled within 10 mm, satisfying relevant building measurement standards. This vision-laser fusion method for measuring interior doors and windows in buildings provides an efficient automated measurement solution for intelligent construction.