Research Series

Urban Spatial Intelligence with Large Language Models

This research series explores how large language models can acquire, evaluate, and reliably use urban spatial knowledge, spanning systematic benchmarking, spatial cognition, multimodal understanding, and factuality alignment.

Benchmark urban capabilities

CityBench establishes a systematic and scalable evaluation environment for examining how large language and vision-language models perform across diverse urban tasks. By connecting heterogeneous city data with interactive simulation, it reveals both the emerging capabilities and persistent limitations of general-purpose models in urban perception, reasoning, and decision-making.

CityBench framework connecting multi-source city data, urban simulation, and perception and decision-making tasks
CityBench integrates CityData, CitySimu, and eight urban tasks across 13 cities worldwide.

Learn urban spatial cognition

CityGPT explores how urban spatial knowledge and reasoning capabilities can be incorporated into large language models. It combines urban instruction data, self-weighted fine-tuning, and systematic evaluation to build models that understand city-scale spatial contexts while preserving their general-purpose capabilities.

CityGPT framework with CityInstruction, self-weighted fine-tuning, and CityEval
CityGPT combines CityInstruction, self-weighted fine-tuning, and CityEval to strengthen urban spatial cognition.

Extend to multimodal understanding

UrbanLLaVA extends urban spatial intelligence from language-based reasoning to multimodal understanding. Through unified urban instruction data and multi-stage training, it enables a single model to interpret and reason across heterogeneous visual and spatial contexts at both local and city scales.

UrbanLLaVA data, training, and evaluation framework for multimodal urban understanding
UrbanLLaVA unifies multimodal urban data, staged training, and evaluation from location to city scale.

Align geospatial factuality

This work studies the reliability of geospatial knowledge encoded in large language models. It introduces a structured benchmark for identifying geospatial hallucinations and a dynamic factuality alignment method that improves the factual consistency and trustworthiness of spatial reasoning.

Geospatial hallucination framework with GeoHaluBench and DynamicKTO factuality alignment
GeoHaluBench detects and attributes geospatial knowledge errors, while DynamicKTO improves model factuality.

Research Outlook

Together, these works move from measuring urban capabilities to building spatially grounded, multimodal, and more trustworthy large language models for open and dynamic environments.

Explore the broader research program