Research Series

Urban Multimodal Reasoning with Vision-Language Models

This research series advances urban multimodal reasoning from socioeconomic understanding toward interpretable inference, spatio-temporal reasoning, and agentic intelligence for cities as they evolve.

Evaluate urban socioeconomic understanding

CityLens establishes a large-scale benchmark for evaluating how vision-language models infer urban socioeconomic conditions from satellite and street-view imagery. Across 17 cities, 11 tasks, and six domains, it exposes both the potential and current limitations of general-purpose models for visually grounded urban sensing.

CityLens benchmark covering socioeconomic sensing tasks from satellite and street-view imagery across global cities
CityLens evaluates socioeconomic sensing across 17 cities, 11 indicators, 17 models, and three evaluation paradigms.

Learn interpretable socioeconomic reasoning

CityRiSE moves beyond direct prediction by training a vision-language model to reason over meaningful urban visual cues. Its reinforcement-learning framework combines perceptual urban reasoning data with task-aligned keyword and regression rewards, improving prediction, interpretability, and transfer to unseen cities and indicators.

CityRiSE dataset and reinforcement-learning framework for urban socioeconomic reasoning
CityRiSE integrates socioeconomic, perceptual urban, and general visual reasoning data with a reward-guided training pipeline.

Reason across space and time

UrbanWell extends multimodal urban sensing from static socioeconomic prediction to spatio-temporal wellbeing analytics. It aligns satellite and street-view observations with 19 indicators across 38 cities and multiple years, enabling systematic evaluation of single-year estimation, multi-year forecasting, and temporal trend reasoning.

UrbanWell framework with multi-year satellite and street-view imagery, wellbeing indicators, and temporal reasoning tasks
UrbanWell spans 38 cities, 19 wellbeing indicators, and imagery from 2012 to 2024 across three spatial and temporal task paradigms.

04Agentic

UrbanAtlas: Agentic Spatio-Temporal Reasoning for Dynamic Urban Environments

Ongoing Work

Build agentic urban reasoning

UrbanAtlas addresses dynamic urban sensing by jointly reasoning over locations, temporal evolution, visual perspectives, and socioeconomic signals. The 8B agentic vision-language model learns to select virtual tools and produce structured intermediate trajectories under a dynamic process-aware reward. Across 12 evaluation settings, it reaches 0.608 accuracy, improving over Qwen3-VL-8B by 35.9%, while transferring to unseen cities and external urban benchmarks.

UrbanAtlas connects multi-view temporal observations with virtual-tool reasoning and process-aware learning for dynamic urban sensing.

Research Outlook

Together, these works develop a path from measuring what vision-language models understand about cities to building structured, transferable, temporally aware, and agentic urban reasoning systems.

Explore the broader research program