Research Series
Urban Spatial Intelligence with Large Language Models
This research series explores how large language models can acquire, evaluate, and reliably use urban spatial knowledge, spanning systematic benchmarking, spatial cognition, multimodal understanding, and factuality alignment.
01Benchmark
CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
KDD 2025 D&B CCF A / Core A*
Benchmark urban capabilities
CityBench establishes a systematic and scalable evaluation environment for examining how large language and vision-language models perform across diverse urban tasks. By connecting heterogeneous city data with interactive simulation, it reveals both the emerging capabilities and persistent limitations of general-purpose models in urban perception, reasoning, and decision-making.
02Learn
CityGPT: Empowering Urban Spatial Cognition of Large Language Models
KDD 2025 Research CCF A / Core A*
Learn urban spatial cognition
CityGPT explores how urban spatial knowledge and reasoning capabilities can be incorporated into large language models. It combines urban instruction data, self-weighted fine-tuning, and systematic evaluation to build models that understand city-scale spatial contexts while preserving their general-purpose capabilities.
03Understand
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
ICCV 2025 CCF A / Core A*
Extend to multimodal understanding
UrbanLLaVA extends urban spatial intelligence from language-based reasoning to multimodal understanding. Through unified urban instruction data and multi-stage training, it enables a single model to interpret and reason across heterogeneous visual and spatial contexts at both local and city scales.
04Align
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
EMNLP 2025 Findings CCF B / Core A
Align geospatial factuality
This work studies the reliability of geospatial knowledge encoded in large language models. It introduces a structured benchmark for identifying geospatial hallucinations and a dynamic factuality alignment method that improves the factual consistency and trustworthiness of spatial reasoning.
Research Outlook
Together, these works move from measuring urban capabilities to building spatially grounded, multimodal, and more trustworthy large language models for open and dynamic environments.
Explore the broader research program