RL-Building Generator
Agent-based reinforcement learning to increase housing density in London.
- Type
- Research
- Site
- Waltham Forest, London
- Method
- Multi-agent reinforcement learning (SAC)
- Training
- 120,000 steps per run
Project overview
This research addresses London’s housing crisis through agent-based reinforcement learning. With house prices up 130% between 2005 and 2023 and the city meeting only a fraction of its housing targets, the study proposes an AI-driven way to increase urban density while maintaining livability.
The project combines machine learning with urban planning principles, using multi-agent systems to find effective densification strategies. Through site digitalization and agent behavior modeling, it builds automated tools that help planners, architects and policymakers make data-driven decisions for sustainable urban development.
Housing crisis in London
- 130%
- Rise in house prices between January 2005 and January 2023.
- <50%
- Of the original housing expansion plan delivered in 2023.
London has the highest rents in the country, with acute pressure in boroughs such as Waltham Forest — which has the fourth-highest overcrowding rate in Outer London, with 18% of homes overcrowded. The project targets these areas for intelligent densification.
Methodology
- 01
Site digitalization
2D/3D mapping, land-use analysis and building indexing.
- 02
Agent modeling
A multi-agent system with ground agents and roof agents.
- 03
RL training
120,000-step training runs with reward optimization.
- 04
Optimization
More density while maintaining livability.
The approach couples comprehensive site analysis with agent-based modeling. Digitalization captures land-use patterns, building functions, solar exposure and spatial relationships, using quadtrees to identify empty space.
The multi-agent system uses two agent types — ground agents for horizontal expansion and roof agents for vertical densification — each trained through extensive simulation to maximize housing density while preserving environmental quality and regulatory compliance.
Site digitalization
Layer 01
Land-use analysis
- Residential areas identified
- Commercial zones mapped
- Green space preserved
Layer 02
Building index
- Height and density mapping
- Function classification
Layer 03
Solar analysis
- Ground radiation mapping
- Shadow impact assessment
The agent
Action
Move
- 6 directions
- Up, down, front, back, left, right
Action
Occupy
- 2 options
- Occupied or not
Observation
Coordinate
- 6 directions
- Up, down, front, back, left, right
Observation
Neighbor
- 25 + 8 + 1 positions
- Also checks cell types
Reward functions
Each rule shapes where an agent is encouraged — or discouraged — to place the next voxel.
Reward 01
Neighbors underneath
AddReward(0.04f * underbuildingCell)AddReward(2.0) — exactly above a cubeAddReward(-2.0) — not exactly above a cube
Reward 02
Neighbors surrounding
AddReward(1.0)AddReward(-1.0)- The first cube always gets the reward.
Reward 03
Footprint in a residential area
AddReward(1.0)AddReward(-1.0)
Reward 04
Footprint away from buildings
AddReward(distance - 2.0)
Reward 05
Agent on the ground
AddReward(0.4 - (0.1 * floor))- Only for the first six cubes.
Reward 06
Bounding-box compactness
Compactness = volumeRatio − diagonalRatioAddReward(Compactness > 0)AddReward(Compactness < 0)
Reward 07
Mean solar index
MeanSunIndex = Sum(sun_map) / (width * length)SolarDelta = (MeanSunIndex − InitialMeanSunIndex) * 50(negative)- A ray is cast from each ground cell toward the sun; if nothing blocks it, that cell of the sun map gains +1.
- Sun vectors: hourly sunlight on the winter solstice in London.
Training results
Before / after
Orthogonal plot
Before / after
Non-orthogonal plot
Before / after
Plot without greenery


Setup
Training parameters
- Training steps
- 120,000 per run
- Training duration
- ~1.2 hours
- Site dimensions
- 400 × 400 × 40 m
- Voxel resolution
- 8 × 8 × 8
Goals
Optimization targets
- Maximize housing density
- Improve sunlight access
- Maintain plot-ratio compliance
- Preserve green space
Case study: Waltham Forest
Site selection criteria
- Mixed land use covering residential, commercial and green space
- Well-connected transport networks and road junctions
- A site large enough (400 × 400 × 40 m) for comprehensive analysis
- Representative of London’s broader housing challenges
Application & optimization
Rewards
Agent performance
- Ground agents
- 257.48 (smoothed)
- Roof agents
- −648.56 (optimizing)
Training duration: 1.2 hours per 120,000 steps.
Outcomes
Optimization results
- Density
- Achieved
- Sunlight access
- Improved
- Regulatory compliance
- Maintained
- Green space
- Protected
Compared with traditional planning
| Process stage | Traditional planning | AI-driven approach |
|---|---|---|
| Site analysis | Manual assessment of physical characteristics | Automated plot digitalization with 2D/3D mapping |
| Contextual studies | Surrounding architecture analyzed by hand | Comprehensive building index and distance mapping |
| Program development | Use mix defined from stakeholder input | Automated land-use optimization |
| Initial massing | Preliminary diagrams considering sun and wind | Integrated ground solar radiation mapping |
| Training / iteration | Manual refinement based on feedback | 120,000-step automated training runs |
| Final production | A human designer creates the final design | AI-generated, optimized urban layout |