OpenStreetMap (OSM) is an open-source geospatial database widely utilized in a broad range of research and application domains.
However, due to its reliance on voluntary user contributions, OSM data often exhibits spatial and temporal inconsistencie...
OpenStreetMap (OSM) is an open-source geospatial database widely utilized in a broad range of research and application domains.
However, due to its reliance on voluntary user contributions, OSM data often exhibits spatial and temporal inconsistencies, particularly in its reflection of urban building-level metadata. The absence of detailed semantic attributes—such as building names and commercial functions—poses practical limitations for fine-grained urban analysis, simulation, and digital twin applications. To address these challenges, this study proposes a modular framework that automatically enriches OSM building attributes using street-level imagery from commercial map services. The system determines optimal locations and parameters for Street View API calls based on building geometry and road vectors, detects signage through object detection models,extracts text using optical character recognition (OCR), and
classifies semantic attributes with the assistance of a large language model (LLM). The framework is designed to be flexible, allowing users to interchange vision, OCR, and classification components based on local conditions or research goals. Although comprehensive experiments are reserved for future work, the proposed framework is expected to enhance the semantic completeness and geographic fidelity of OSM data. Its extensibility and automation potential offer a promising tool for next-generation urban data augmentation pipelines.