Seg & Predict

Optimizing urban perception — integrating image segmentation and machine learning in London.

Type
Research
Team
Wenshuo Zhang, Lewen Zhang, Robin Song
Models
PSPNet · Random Forest · XGBoost · SAC
Platform
Flask + React
Seg & Predict — Optimizing urban perception — integrating image segmentation and machine learning in London.

Project overview

This research explores the relationship between street-view imagery and crime rates in London using computer vision and machine learning. By analyzing the visual characteristics of urban environments, we look for patterns that correlate with criminal activity and develop predictive models for urban safety assessment.

The project combines image segmentation, regression analysis and reinforcement learning into an automated framework for predicting crime rates from street-view data — offering insight for urban planning, policy-making and community safety.

Research question

Q

Core question

How can we quantify the correlation between street views and crime rates?

Can visual features of urban environments predict areas with higher crime incidence?

Aim

Research goal

Develop automated tools that predict crime rates from street-view imagery.

Create actionable insight for urban planning and safety improvement.

Street views of London framing the research question
Urban safety and the street view.

Methodology

  1. 01

    Data collection

    Crime data and street-view images from London.

  2. 02

    Segmentation

    PSPNet trained on ADE20k for scene parsing.

  3. 03

    Regression

    Random Forest and XGBoost for prediction.

  4. 04

    Optimization

    SAC reinforcement learning to improve low-scoring scenes.

Research methodology workflow diagram
Methodology workflow.

Image segmentation

Network

PSPNet

Pyramid Scene Parsing Network, with a Pyramid Pooling Module for semantic segmentation.

Backbone

ResNet-50

A 50-layer CNN with residual connections that avoid vanishing gradients.

Dataset

ADE20k

20,000+ images across 150 semantic categories for comprehensive scene understanding.

Original street view next to its semantic segmentation
Original and segmented street view.

Regression analysis

Rejected

Linear regression

R²
0.348
Adjusted R²
0.343

Rejected due to low performance.

Accepted

Random Forest

R²
0.451
Variance explained
44.32%

Accepted for low-to-middle crime rates.

Accepted

XGBoost

R²
0.451
Distribution
Even spread

Accepted for higher crime rates.

Regression model performance chart
Model performance I.
Regression model performance chart
Model performance II.
Regression model performance chart
Model performance III.

Reinforcement learning optimization

Agent

SAC (Soft Actor-Critic)

  • An actor network generates actions that modify the image.
  • A dual-Q network evaluates the value of those actions.
  • Together they guide optimization toward lower predicted crime.

Reward

Reward mechanism

Score improvement
Sigmoid function
Colour-ratio reward
Weighted by correlation
Trend reward
5-step regression
Soft Actor-Critic architecture diagram
SAC architecture.
Reinforcement learning optimization results
Optimization process.

Interactive platform

A web platform wraps the whole pipeline: upload a street view, see it segmented in real time, get a predicted crime score and let the agent suggest improvements. Open the live demo ↗

Features

Platform features

  • Upload street-view images
  • Real-time image segmentation
  • Crime-rate prediction
  • AI-powered optimization

Stack

Technical stack

Backend
Flask · PyTorch · OpenCV
Frontend
React · live, interactive visualization
Screenshot of the interactive platform interface
Platform interface.
Platform prediction and optimization view
Prediction and optimization view.

Results & evaluation

Outcomes

Key achievements

  • Correlated street-view features with crime rates
  • Developed an automated workflow for crime analysis
  • Introduced new reward mechanisms for reinforcement learning
  • Built a scalable platform for real-time analysis

Next steps

Limitations

  • Limited database scale and diversity
  • Segmentation precision needs improvement
  • Overall R² values remain below 0.5
  • The reinforcement learning stage needs further tuning

WeChat

WeChat QR code for Robin Song

Scan with WeChat to add me.