Planning public transit in newly developing urban districts is challenging because conventional methods require origin-destination demand matrices that do not exist where existing service does not yet capture local demand. This thesis develops a metho...
Planning public transit in newly developing urban districts is challenging because conventional methods require origin-destination demand matrices that do not exist where existing service does not yet capture local demand. This thesis develops a methodology that infers spatial demand structure from publicly available taxi trip records and translates it into bus stop locations and a route. Unlike conventional transit planning methods, the approach requires neither prior demand volumes nor predefined candidate stop sites. The approach has two phases. In Phase 1, taxi origin-destination points are spatially filtered and clustered using HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise), which identifies stable demand concentrations at varying densities while explicitly filtering noise, addressing the unknown cluster count, heterogeneous density, and data quality challenges inherent in mobility data. A two-stage clustering design applies fine-grained parameters for intra-district stop placement and coarser parameters for transfer point identification. In Phase 2, cluster centroids are projected onto the road network and connected via the Christofides-Serdyukov algorithm, which provides a 3/2-approximation guarantee for the undirected metric TSP relaxation; the final loop is then reconciled against directed road constraints and reported empirically for feasible length and cycle time. Applied to the Daegu Medical R&D District (1.19 km² study-area polygon), a transit-underserved development area in South Korea with only one existing bus route covering 33% of identified demand clusters. From 454,563 taxi records, 701 in-zone demand points are extracted and clustered into 15 stops with a 95.3% clustering rate. The resulting 6.67 km loop route achieves complete cluster coverage, reduces the area-weighted average walking distance from 238 m (existing service) to 135 m, and executes in under 30 seconds on a standard workstation. Parameter sensitivity tests show a stable region (ε ∈ [0.03, 0.05], k ∈ [3, 4]), bootstrap resampling stability (12.7 ± 0.9 clusters with 35 m mean stop displacement under 20% data removal), and noise injection resilience (cluster structure preserved at 20% random noise). Comparison against k-means, DBSCAN, and uniform grid baselines confirms HDBSCAN's unique combination of automatic cluster determination, noise rejection, and density- adaptive placement. The result is a reproducible, end-to-end pipeline for transit planning in data- scarce environments, positioning density-based clustering as a demand structure inference stage that precedes formal facility location optimization.