Research Series
Urban Multimodal Reasoning with Vision-Language Models
This research series advances urban multimodal reasoning from socioeconomic understanding toward interpretable inference, spatio-temporal reasoning, and agentic intelligence for cities as they evolve.
01Benchmark
CityLens: Benchmarking Large Vision-Language Models for Urban Socioeconomic Sensing
ICLR 2026 CCF A / Core A*
Evaluate urban socioeconomic understanding
CityLens establishes a large-scale benchmark for evaluating how vision-language models infer urban socioeconomic conditions from satellite and street-view imagery. Across 17 cities, 11 tasks, and six domains, it exposes both the potential and current limitations of general-purpose models for visually grounded urban sensing.
02Reason
CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning
ACM MM 2026 CCF A / Core A*
Learn interpretable socioeconomic reasoning
CityRiSE moves beyond direct prediction by training a vision-language model to reason over meaningful urban visual cues. Its reinforcement-learning framework combines perceptual urban reasoning data with task-aligned keyword and regression rewards, improving prediction, interpretability, and transfer to unseen cities and indicators.
03Dynamics
UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
KDD 2026 D&B CCF A / Core A*
Reason across space and time
UrbanWell extends multimodal urban sensing from static socioeconomic prediction to spatio-temporal wellbeing analytics. It aligns satellite and street-view observations with 19 indicators across 38 cities and multiple years, enabling systematic evaluation of single-year estimation, multi-year forecasting, and temporal trend reasoning.
04Agentic
UrbanAtlas: Agentic Spatio-Temporal Reasoning for Dynamic Urban Environments
Ongoing Work
Build agentic urban reasoning
UrbanAtlas addresses dynamic urban sensing by jointly reasoning over locations, temporal evolution, visual perspectives, and socioeconomic signals. The 8B agentic vision-language model learns to select virtual tools and produce structured intermediate trajectories under a dynamic process-aware reward. Across 12 evaluation settings, it reaches 0.608 accuracy, improving over Qwen3-VL-8B by 35.9%, while transferring to unseen cities and external urban benchmarks.
Multi-view temporal observations
Dynamic urban reasoning
Research Outlook
Together, these works develop a path from measuring what vision-language models understand about cities to building structured, transferable, temporally aware, and agentic urban reasoning systems.
Explore the broader research program