RL-Building Generator

Agent-based reinforcement learning to increase housing density in London.

Type
Research
Site
Waltham Forest, London
Method
Multi-agent reinforcement learning (SAC)
Training
120,000 steps per run
RL-Building Generator — Agent-based reinforcement learning to increase housing density in London.

Project overview

This research addresses London’s housing crisis through agent-based reinforcement learning. With house prices up 130% between 2005 and 2023 and the city meeting only a fraction of its housing targets, the study proposes an AI-driven way to increase urban density while maintaining livability.

The project combines machine learning with urban planning principles, using multi-agent systems to find effective densification strategies. Through site digitalization and agent behavior modeling, it builds automated tools that help planners, architects and policymakers make data-driven decisions for sustainable urban development.

Housing crisis in London

130%
Rise in house prices between January 2005 and January 2023.
<50%
Of the original housing expansion plan delivered in 2023.

London has the highest rents in the country, with acute pressure in boroughs such as Waltham Forest — which has the fourth-highest overcrowding rate in Outer London, with 18% of homes overcrowded. The project targets these areas for intelligent densification.

Charts of London house prices and housing delivery
London’s housing crisis in data.

Methodology

  1. 01

    Site digitalization

    2D/3D mapping, land-use analysis and building indexing.

  2. 02

    Agent modeling

    A multi-agent system with ground agents and roof agents.

  3. 03

    RL training

    120,000-step training runs with reward optimization.

  4. 04

    Optimization

    More density while maintaining livability.

The approach couples comprehensive site analysis with agent-based modeling. Digitalization captures land-use patterns, building functions, solar exposure and spatial relationships, using quadtrees to identify empty space.

The multi-agent system uses two agent types — ground agents for horizontal expansion and roof agents for vertical densification — each trained through extensive simulation to maximize housing density while preserving environmental quality and regulatory compliance.

Workflow diagram of the AI methodology
Methodology workflow.

Site digitalization

Data processing of the site
Data processing.
Digitalized site layers
Digitalized site.

Layer 01

Land-use analysis

Land-use map
  • Residential areas identified
  • Commercial zones mapped
  • Green space preserved

Layer 02

Building index

Building index map
  • Height and density mapping
  • Function classification

Layer 03

Solar analysis

Ground solar radiation map
  • Ground radiation mapping
  • Shadow impact assessment

The agent

Overview of the agent in its voxel environment
Agent overview.

Action

Move

Action: Move
  • 6 directions
  • Up, down, front, back, left, right

Action

Occupy

Action: Occupy
  • 2 options
  • Occupied or not

Observation

Coordinate

Observation: Coordinate
  • 6 directions
  • Up, down, front, back, left, right

Observation

Neighbor

Observation: Neighbor
  • 25 + 8 + 1 positions
  • Also checks cell types

Reward functions

Each rule shapes where an agent is encouraged — or discouraged — to place the next voxel.

Reward 01

Neighbors underneath

Reward rule: Neighbors underneath
  • AddReward(0.04f * underbuildingCell)
  • AddReward(2.0) — exactly above a cube
  • AddReward(-2.0) — not exactly above a cube

Reward 02

Neighbors surrounding

Reward rule: Neighbors surrounding
  • AddReward(1.0)
  • AddReward(-1.0)
  • The first cube always gets the reward.

Reward 03

Footprint in a residential area

Reward rule: Footprint in a residential area
  • AddReward(1.0)
  • AddReward(-1.0)

Reward 04

Footprint away from buildings

Reward rule: Footprint away from buildings
  • AddReward(distance - 2.0)

Reward 05

Agent on the ground

Reward rule: Agent on the ground
  • AddReward(0.4 - (0.1 * floor))
  • Only for the first six cubes.

Reward 06

Bounding-box compactness

Reward rule: Bounding-box compactness
  • Compactness = volumeRatio − diagonalRatio
  • AddReward(Compactness > 0)
  • AddReward(Compactness < 0)

Reward 07

Mean solar index

  • MeanSunIndex = Sum(sun_map) / (width * length)
  • SolarDelta = (MeanSunIndex − InitialMeanSunIndex) * 50 (negative)
  • A ray is cast from each ground cell toward the sun; if nothing blocks it, that cell of the sun map gains +1.
  • Sun vectors: hourly sunlight on the winter solstice in London.
Mean solar index map
Mean solar index.

Training results

Training result on an orthogonal plot
Orthogonal plot.
Training result on a non-orthogonal plot
Non-orthogonal plot.
Training result of the roof agent
Roof agent.

Before / after

Orthogonal plot

Orthogonal plot — before training
Before training
Orthogonal plot — after training
After training

Before / after

Non-orthogonal plot

Non-orthogonal plot — before training
Before training
Non-orthogonal plot — after training
After training

Before / after

Plot without greenery

Plot without greenery — before training
Before training
Plot without greenery — after training
After training

Setup

Training parameters

Training steps
120,000 per run
Training duration
~1.2 hours
Site dimensions
400 × 400 × 40 m
Voxel resolution
8 × 8 × 8

Goals

Optimization targets

  • Maximize housing density
  • Improve sunlight access
  • Maintain plot-ratio compliance
  • Preserve green space

Case study: Waltham Forest

Site selection criteria

  • Mixed land use covering residential, commercial and green space
  • Well-connected transport networks and road junctions
  • A site large enough (400 × 400 × 40 m) for comprehensive analysis
  • Representative of London’s broader housing challenges
Urban analysis of Waltham Forest
Waltham Forest site analysis.

Application & optimization

Generated densification applied to the site
Application on site.
Optimization process analytics
Optimization process.

Rewards

Agent performance

Agent reward curves during training
Ground agents
257.48 (smoothed)
Roof agents
−648.56 (optimizing)

Training duration: 1.2 hours per 120,000 steps.

Outcomes

Optimization results

Optimization results chart
Density
Achieved
Sunlight access
Improved
Regulatory compliance
Maintained
Green space
Protected

Compared with traditional planning

Process stageTraditional planningAI-driven approach
Site analysisManual assessment of physical characteristicsAutomated plot digitalization with 2D/3D mapping
Contextual studiesSurrounding architecture analyzed by handComprehensive building index and distance mapping
Program developmentUse mix defined from stakeholder inputAutomated land-use optimization
Initial massingPreliminary diagrams considering sun and windIntegrated ground solar radiation mapping
Training / iterationManual refinement based on feedback120,000-step automated training runs
Final productionA human designer creates the final designAI-generated, optimized urban layout

WeChat

WeChat QR code for Robin Song

Scan with WeChat to add me.