dc-l2-lagos-snap-normalize
Map (reference ← swipe → agent)
0:00
I need to get our legacy zoning layer ready for the state portal. lagos_zoning_legacy.gpkg has inconsistent zoning_class values (mixed casing, spelling variations, and some blank entries) plus vertex precision issues that cause tiny gaps and overlaps between adjacent parcels. Can you clean up the class names to consistent canonical TitleCase values, drop any rows with blank or whitespace-only classes, fix the vertex precision problems, and produce one merged geometry per canonical zoning class? Please write the result to zoning_aggregated.gpkg, with each row carrying two columns: zoning_class and a numeric area_m2.
0:00
System info
{
"adapter": "openrouter",
"model": "google/gemma-4-26b-a4b-it",
"base_url": "https://openrouter.ai/api/v1"
}
0:01
Bash
ls -F
lagos_zoning_legacy.gpkg
0:03
Write
/work/inspect.py
content · 16 lines · py
import geopandas as gpd
def inspect_data(file_path):
gdf = gpd.read_file(file_path)
print("Columns:", gdf.columns.tolist())
print("First 5 rows of zoning_class:")
print(gdf['zoning_class'].head())
print("\nUnique values in zoning_class:")
print(gdf['zoning_class'].unique())
print("\nCRS:", gdf.crs)
print("\nGeometry type:", gdf.geometry.type.unique())
print("\nShape info:", gdf.shape)
if __name__ == "__main__":
inspect_data('lagos_zoning_legacy.gpkg')
Wrote /work/inspect.py (498 chars).
0:04
Bash
python inspect.py
Traceback (most recent call last):
File "/work/inspect.py", line 1, in <module>
import geopandas as gpd
File "/usr/local/lib/python3.14/site-packages/geopandas/__init__.py", line 3, in <module>
from geopandas.geoseries import GeoSeries
File "/usr/local/lib/python3.14/site-packages/geopandas/geoseries.py", line 7, in <module>
import numpy as np
File "/usr/local/lib/python3.14/site-packages/numpy/__init__.py", line 112, in <module>
from numpy.__config__ import show_config
File "/usr/local/lib/python3.14/site-packages/numpy/__config__.py", line 4, in <module>
from numpy._core._multiarray_umath import (
File "/usr/local/lib/python3.14/site-packages/numpy/_core/__init__.py", line 24, in <module>
from . import multiarray
File "/usr/local/lib/python3.14/site-packages/numpy/_core/multiarray.py", line 115, in <module>
@array_function_from_c_func_and_dispatcher(_multiarray_umath.empty_like)
File "/usr/local/lib/python3.14/site-packages/numpy/_core/overrides.py", line 185, in decorator
return array_function_dispatch(
File "/usr/local/lib/python3.14/site-packages/numpy/_core/overrides.py", line 161, in decorator
doc = inspect.cleandoc(dispatcher.__doc__)
AttributeError: module 'inspect' has no attribute 'cleandoc' (consider renaming '/work/inspect.py' since it has the same name as the standard library module named 'inspect' and prevents importing that standard library module)
0:05
Bash
rm /work/inspect.py
(no output)
0:07
Write
/work/data_inspect.py
content · 16 lines · py
import geopandas as gpd
def inspect_data(file_path):
gdf = gpd.read_file(file_path)
print("Columns:", gdf.columns.tolist())
print("First 5 rows of zoning_class:")
print(gdf['zoning_class'].head())
print("\nUnique values in zoning_class:")
print(gdf['zoning_class'].unique())
print("\nCRS:", gdf.crs)
print("\nGeometry type:", gdf.geometry.type.unique())
print("\nShape info:", gdf.shape)
if __name__ == "__main__":
inspect_data('lagos_zoning_legacy.gpkg')
Wrote /work/data_inspect.py (498 chars).
0:08
Bash
python data_inspect.py
Columns: ['parcel_id', 'lga_source', 'zoning_class', 'area_m2', 'geometry']
First 5 rows of zoning_class:
0 RESIDENTIAL
1 residential
2 Resi.
3 resi.
4 RESI.
Name: zoning_class, dtype: str
Unique values in zoning_class:
<ArrowStringArray>
[ 'RESIDENTIAL', 'residential', 'Resi.', 'resi.',
'RESI.', 'Residential', 'Comm.', 'comm.',
'COMM.', 'Commercial', 'COMMERCIAL', 'commercial',
'Indus.', 'indus.', 'INDUS.', 'Industrial',
'INDUSTRIAL', 'industrial', 'AGRI.', 'Agricultural',
'AGRICULTURAL', 'agricultural', 'Agri.', 'agri.',
'', ' ', nan, '\t']
Length: 28, dtype: str
CRS: EPSG:26331
Geometry type: <ArrowStringArray>
['Polygon']
Length: 1, dtype: str
Shape info: (10080, 5)
0:17
Write
/work/solve.py
content · 88 lines · py
import geopandas as gpd
import pandas as pd
from shapely.geometry import MultiPolygon, Polygon
from shapely import snap, simplify
def clean_zoning_class(val):
if pd.isna(val):
return None
s = str(val).strip().lower()
if not s:
return None
# Mapping logic based on inspection
if s in ['residential', 'resi.']:
return 'Residential'
elif s in ['commercial', 'comm.']:
return 'Commercial'
elif s in ['industrial', 'indus.']:
return 'Industrial'
elif s in ['agricultural', 'agri.']:
return 'Agricultural'
else:
# Fallback to Title Case if it doesn't match known patterns
return s.title()
def solve():
print("Loading data...")
gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
print("Cleaning zoning_class...")
# 1. Clean class names
gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
# 2. Drop rows with None/NaN/empty classes
gdf = gdf.dropna(subset=['zoning_class'])
gdf = gdf[gdf['zoning_class'].str.strip() != '']
print(f"Remaining rows after cleaning classes: {len(gdf)}")
print("Unique classes:", gdf['zoning_class'].unique())
# 3. Fix vertex precision problems (gaps and overlaps)
# A common technique is to snap geometries to a grid or to each other.
# Since we want to merge them anyway, we can simplify/snap first.
# Given they are adjacent, snapping to a grid is often effective.
print("Fixing vertex precision...")
# Snapping to a grid of 1cm (0.01m) if EPSG:26331 (metric)
# EPSG:26331 is likely a UTM projection.
gdf['geometry'] = gdf['geometry'].apply(lambda geom: simplify(snap(geom, geom, 0.01), 0.01))
# Alternatively, a more robust way to fix gaps/overlaps in adjacent polygons
# is to use a buffer(0) or a small negative/positive buffer,
# but for "merging" later, let's try to snap them.
# Given the instruction, we'll use a small snap.
# 4. Group by zoning_class and merge geometries
print("Merging geometries per class...")
# Ensure we have a metric CRS for area calculation
# The inspect showed EPSG:26331 which is metric.
results = []
for z_class, group in gdf.groupby('zoning_class'):
print(f"Processing class: {z_class}")
# Dissolve geometries
merged_geom = group.dissolve(by='zoning_class')['geometry'].iloc[0]
# Calculate area
area = merged_geom.area
results.append({
'zoning_class': z_class,
'area_m2': area,
'geometry': merged_geom
})
# Create the final GeoDataFrame
final_gdf = gpd.GeoDataFrame(results, crs=gdf.crs)
# Final check on geometry type (dissolve might produce MultiPolygons)
# The prompt asks for "one merged geometry per canonical zoning class".
print(f"Writing to zoning_aggregated.gpkg...")
final_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
print("Done!")
if __name__ == "__main__":
solve()
Wrote /work/solve.py (3072 chars).
0:18
Bash
python solve.py
Loading data... Cleaning zoning_class... Remaining rows after cleaning classes: 10030 Unique classes: <ArrowStringArray> ['Residential', 'Commercial', 'Industrial', 'Agricultural'] Length: 4, dtype: str Fixing vertex precision... Merging geometries per class... Processing class: Agricultural Processing class: Commercial Processing class: Industrial Processing class: Residential Writing to zoning_aggregated.gpkg... Done!
0:31
Bash
python -c "import geopandas as gpd; gdf = gpd.read_file('zoning_aggregated.gpkg'...
zoning_class ... geometry 0 Agricultural ... MULTIPOLYGON (((540500 720500, 540500 720510, ... 1 Commercial ... MULTIPOLYGON (((540510 720030, 540510 720030, ... 2 Industrial ... MULTIPOLYGON (((540000 720530, 540000 720540, ... 3 Residential ... MULTIPOLYGON (((540010 720000, 540000 720000, ... [4 rows x 3 columns] Index(['zoning_class', 'area_m2', 'geometry'], dtype='str') EPSG:26331
0:47
Write
/work/solve.py
content · 105 lines · py
import geopandas as gpd
import pandas as pd
from shapely.geometry import MultiPolygon, Polygon
from shapely import snap, simplify
def clean_zoning_class(val):
if pd.isna(val):
return None
s = str(val).strip().lower()
if not s:
return None
# Mapping logic based on inspection
if s in ['residential', 'resi.']:
return 'Residential'
elif s in ['commercial', 'comm.']:
return 'Commercial'
elif s in ['industrial', 'indus.']:
return 'Industrial'
elif s in ['agricultural', 'agri.']:
return 'Agricultural'
else:
# Fallback to Title Case if it doesn't match known patterns
return s.title()
def solve():
print("Loading data...")
gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
print("Cleaning zoning_class...")
# 1. Clean class names
gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
# 2. Drop rows with None/NaN/empty classes
gdf = gdf.dropna(subset=['zoning_class'])
gdf = gdf[gdf['zoning_class'].str.strip() != '']
print(f"Remaining rows after cleaning classes: {len(gdf)}")
print("Unique classes:", gdf['zoning_class'].unique())
# 3. Fix vertex precision problems (gaps and overlaps)
print("Fixing vertex precision...")
# Use a small snap tolerance to align vertices
# Since we are dissolving later, the most important thing is to avoid
# tiny gaps or slivers that might create tiny "islands" or holes.
# A good approach for adjacent polygons is to snap everything to a grid.
# First, round the coordinates to a certain precision to reduce vertex clutter
def round_coords(geom, precision=0.01):
if geom.is_empty:
return geom
# We use a trick: simplify with a very small tolerance or just round
# For most GIS workflows, snapping to a grid is done via:
return simplify(geom, tolerance=0.001) # Very small tolerance to clean up noise
# Let's try snapping to a grid by rounding coordinates.
# This is often more robust for "merging" than snap(geom, geom).
def snap_to_grid(geom, precision=0.01):
if geom.is_empty:
return geom
# Using simplify with a tiny tolerance often cleans up precision issues
return simplify(geom, tolerance=precision)
# Apply a small snap/simplify to clean up the topology
# We use 0.01m as the tolerance for EPSG:26331
gdf['geometry'] = gdf['geometry'].apply(lambda g: simplify(g, 0.001))
# 4. Group by zoning_class and merge geometries
print("Merging geometries per class...")
# To handle overlaps and gaps more effectively during dissolution:
# We can use the 'dissolve' method which performs a union.
# To mitigate the impact of gaps/overlaps, we can also use a tiny buffer.
results = []
for z_class, group in gdf.groupby('zoning_class'):
print(f"Processing class: {z_class}")
# Dissolve (union) the geometries in the group
# We dissolve by 'zoning_class'
dissolved = group.dissolve(by='zoning_class')
merged_geom = dissolved.geometry.iloc[0]
# Calculate area
area = merged_geom.area
results.append({
'zoning_class': z_class,
'area_m2': area,
'geometry': merged_geom
})
# Create the final GeoDataFrame
final_gdf = gpd.GeoDataFrame(results, crs=gdf.crs)
# Ensure columns are in requested order
final_gdf = final_gdf[['zoning_class', 'area_m2', 'geometry']]
print(f"Writing to zoning_aggregated.gpkg...")
final_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
print("Done!")
if __name__ == "__main__":
solve()
Wrote /work/solve.py (3779 chars).
0:48
Bash
python solve.py
Loading data... Cleaning zoning_class... Remaining rows after cleaning classes: 10030 Unique classes: <ArrowStringArray> ['Residential', 'Commercial', 'Industrial', 'Agricultural'] Length: 4, dtype: str Fixing vertex precision... Merging geometries per class... Processing class: Agricultural Processing class: Commercial Processing class: Industrial Processing class: Residential Writing to zoning_aggregated.gpkg... Done!
0:52
Bash
python -c "import geopandas as gpd; gdf = gpd.read_file('zoning_aggregated.gpkg'...
zoning_class ... geometry 0 Agricultural ... MULTIPOLYGON (((540500 720500, 540500 720510, ... 1 Commercial ... MULTIPOLYGON (((540510 720030, 540510 720030, ... 2 Industrial ... MULTIPOLYGON (((540000 720530, 540000 720540, ... 3 Residential ... MULTIPOLYGON (((540010 720000, 540000 720000, ... [4 rows x 3 columns] Index(['zoning_class', 'area_m2', 'geometry'], dtype='str') EPSG:26331 zoning_class str area_m2 float64 geometry geometry dtype: object
0:58