dc-l2-lagos-snap-normalize
Map (reference ← swipe → agent)
0:00
I need to get our legacy zoning layer ready for the state portal. lagos_zoning_legacy.gpkg has inconsistent zoning_class values (mixed casing, spelling variations, and some blank entries) plus vertex precision issues that cause tiny gaps and overlaps between adjacent parcels. Can you clean up the class names to consistent canonical TitleCase values, drop any rows with blank or whitespace-only classes, fix the vertex precision problems, and produce one merged geometry per canonical zoning class? Please write the result to zoning_aggregated.gpkg, with each row carrying two columns: zoning_class and a numeric area_m2.
0:00
System info
{
"adapter": "openrouter",
"model": "google/gemma-4-26b-a4b-it",
"base_url": "https://openrouter.ai/api/v1"
}
0:02
Bash
ls -F
lagos_zoning_legacy.gpkg
0:06
Write
/work/inspect_data.py
content · 13 lines · py
import geopandas as gpd
def inspect_data(file_path):
gdf = gpd.read_file(file_path)
print("Columns:", gdf.columns.tolist())
print("First 5 rows:\n", gdf.head())
print("\nUnique zoning_class values:\n", gdf['zoning_class'].unique())
print("\nShape info:", gdf.shape)
print("\nCRS:", gdf.crs)
if __name__ == "__main__":
inspect_data('lagos_zoning_legacy.gpkg')
Wrote /work/inspect_data.py (389 chars).
0:08
Bash
python inspect_data.py
Columns: ['parcel_id', 'lga_source', 'zoning_class', 'area_m2', 'geometry']
First 5 rows:
parcel_id ... geometry
0 1 ... POLYGON ((540000 720000, 540010 720000, 540010...
1 2 ... POLYGON ((540010 720000, 540020 720000, 540020...
2 3 ... POLYGON ((540020 720000, 540030 720000, 540030...
3 4 ... POLYGON ((540030 720000, 540040 720000, 540040...
4 5 ... POLYGON ((540040 720000, 540050 720000, 540050...
[5 rows x 5 columns]
Unique zoning_class values:
<ArrowStringArray>
[ 'RESIDENTIAL', 'residential', 'Resi.', 'resi.',
'RESI.', 'Residential', 'Comm.', 'comm.',
'COMM.', 'Commercial', 'COMMERCIAL', 'commercial',
'Indus.', 'indus.', 'INDUS.', 'Industrial',
'INDUSTRIAL', 'industrial', 'AGRI.', 'Agricultural',
'AGRICULTURAL', 'agricultural', 'Agri.', 'agri.',
'', ' ', nan, '\t']
Length: 28, dtype: str
Shape info: (10080, 5)
CRS: EPSG:26331
1:07
Write
/work/solve.py
content · 99 lines · py
import geopandas as gpd
import pandas as pd
from shapely.geometry import MultiPolygon, Polygon
from shapely.ops import unary_union
def clean_zoning_class(val):
if pd.isna(val):
return None
s = str(val).strip().lower()
if not s:
return None
# Mapping dictionary for cleaning up abbreviations and variations
mapping = {
'resi.': 'Residential',
'residential': 'Residential',
'comm.': 'Commercial',
'commercial': 'Commercial',
'indus.': 'Industrial',
'industrial': 'Industrial',
'agri.': 'Agricultural',
'agricultural': 'Agricultural'
}
# First try the mapping
if s in mapping:
return mapping[s]
# Fallback: title case if not in mapping (though mapping covers all seen so far)
return s.title()
def solve():
print("Loading data...")
gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
print("Cleaning zoning classes...")
# Remove rows with null or whitespace-only classes
gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
gdf = gdf.dropna(subset=['zoning_class'])
gdf = gdf[gdf['zoning_class'].str.strip() != '']
# Ensure casing is consistent (Title Case) as per mapping
# Actually, the mapping above already returns Title Case.
print(f"Remaining unique classes: {gdf['zoning_class'].unique()}")
print("Fixing vertex precision and aggregating...")
# To fix vertex precision issues (gaps/overlaps), we use a small buffer/snap approach.
# However, the prompt asks to "fix the vertex precision problems, and produce one merged geometry per canonical zoning class".
# A common way to fix topology issues when merging is to use unary_union on the geometries of each class.
# unary_union effectively dissolves boundaries and handles overlaps.
aggregated_data = []
for z_class, group in gdf.groupby('zoning_class'):
print(f"Processing class: {z_class}")
# Fix precision by rounding coordinates to a reasonable precision (e.g., 2 decimal places)
# This helps align vertices that are nearly identical but slightly off.
group = group.copy()
group['geometry'] = group['geometry'].apply(lambda geom: lambda x: round(x, 2) if isinstance(x, float) else x) # This is not quite right for geometry.
# Correct way to round geometry coordinates:
def round_coords(geom, precision=2):
if geom.is_empty:
return geom
# A simpler way to handle precision for topological consistency is to use buffer(0) or snap
# But rounding is requested by "fix vertex precision".
# Let's use a trick: buffer(0) to fix invalidities, then unary_union.
return geom
# Using unary_union is the standard way to dissolve and resolve overlaps/gaps.
# To handle the "precision" part, we'll snap to a grid or round.
# Let's try snapping to a 0.01m grid.
merged_geom = unary_union(group.geometry)
# Calculate area in m2
area = merged_geom.area
aggregated_data.append({
'zoning_class': z_class,
'area_m2': area,
'geometry': merged_geom
})
print("Creating final GeoDataFrame...")
result_gdf = gpd.GeoDataFrame(aggregated_data, crs=gdf.crs)
# Reorder columns to ensure zoning_class and area_m2 are present as requested
# (and geometry is the geometry column)
result_gdf = result_gdf[['zoning_class', 'area_m2', 'geometry']]
print(f"Saving to zoning_aggregated.gpkg...")
result_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
print("Done!")
if __name__ == "__main__":
solve()
Wrote /work/solve.py (3805 chars).
1:59
Write
/work/solve.py
content · 104 lines · py
import geopandas as gpd
import pandas as pd
from shapely.ops import unary_union
import numpy as np
def clean_zoning_class(val):
if pd.isna(val):
return None
s = str(val).strip().lower()
if not s:
return None
# Mapping dictionary for cleaning up abbreviations and variations
mapping = {
'resi.': 'Residential',
'residential': 'Residential',
'comm.': 'Commercial',
'commercial': 'Commercial',
'indus.': 'Industrial',
'industrial': 'Industrial',
'agri.': 'Agricultural',
'agricultural': 'Agricultural'
}
# Try the mapping
if s in mapping:
return mapping[s]
# Fallback: Title Case
return s.title()
def snap_to_grid(geom, precision=0.01):
"""
Snap geometry coordinates to a grid to resolve precision issues.
"""
if geom.is_empty:
return geom
# A simple way to snap is to use the buffer(0) trick or snap to a grid
# For many GIS workflows, buffer(0) fixes topology, but snapping coordinates is more robust for gaps.
# Here we use a simple rounding of coordinates via a transformation or just buffer.
# Let's use the 'snap' idea by using a small buffer to merge things that are close.
return geom.buffer(0)
def solve():
print("Loading data...")
gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
print("Cleaning zoning classes...")
gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
gdf = gdf.dropna(subset=['zoning_class'])
gdf = gdf[gdf['zoning_class'].str.strip() != '']
print(f"Remaining unique classes: {gdf['zoning_class'].unique()}")
# To fix vertex precision (gaps/overlaps), we will:
# 1. Round coordinates slightly to align near-identical vertices.
# 2. Dissolve geometries.
# Note: Rounding coordinates directly on geometry is tricky in shapely.
# A robust way is to use a small buffer to close gaps and then unary_union.
# But the request says "fix the vertex precision problems".
# Let's try to snap to a 1cm grid (0.01) using a trick or simply using buffer(0)
# which often fixes precision-related topology issues.
aggregated_data = []
for z_class, group in gdf.groupby('zoning_class'):
print(f"Processing class: {z_class}")
# Use unary_union to dissolve. To handle small gaps, we can apply a tiny buffer.
# But let's try unary_union first. If there are gaps/overlaps, unary_union
# on the geometries themselves is the best way to 'dissolve'.
# To truly address vertex precision, we can round the geometries.
# We'll use a small buffer to bridge gaps.
# A common technique: buffer(dist) then buffer(-dist).
# Or just buffer(0).
# Let's do: unary_union on geometries that have been buffered slightly
# to close gaps, but that might be too aggressive.
# Let's just use unary_union; it handles overlaps well.
merged_geom = unary_union(group.geometry)
# Calculate area
area = merged_geom.area
aggregated_data.append({
'zoning_class': z_class,
'area_m2': area,
'geometry': merged_geom
})
print("Creating final GeoDataFrame...")
result_gdf = gpd.GeoDataFrame(aggregated_data, crs=gdf.crs)
result_gdf = result_gdf[['zoning_class', 'area_m2', 'geometry']]
print(f"Saving to zoning_aggregated.gpkg...")
result_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
print("Done!")
if __name__ == "__main__":
solve()
Wrote /work/solve.py (3676 chars).
2:34
Write
/work/solve.py
content · 104 lines · py
import geopandas as gpd
import pandas as pd
from shapely.ops import unary_union
import shapely
def clean_zoning_class(val):
if pd.isna(val):
return None
s = str(val).strip().lower()
if not s:
return None
mapping = {
'resi.': 'Residential',
'residential': 'Residential',
'comm.': 'Commercial',
'commercial': 'Commercial',
'indus.': 'Industrial',
'industrial': 'Industrial',
'agri.': 'Agricultural',
'agricultural': 'Agricultural'
}
if s in mapping:
return mapping[s]
return s.title()
def solve():
print("Loading data...")
gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
print("Cleaning zoning classes...")
gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
gdf = gdf.dropna(subset=['zoning_class'])
gdf = gdf[gdf['zoning_class'].str.strip() != '']
print(f"Unique classes: {gdf['zoning_class'].unique()}")
# To address "vertex precision issues that cause tiny gaps and overlaps"
# We can snap all geometries to a grid.
# A common way in shapely is to use a small buffer to close gaps,
# but "fix precision" often means rounding the coordinates.
# Let's use a coordinate rounding approach for the geometries.
def round_geometry(geom, decimals=2):
if geom.is_empty:
return geom
# Using the coordinate precision trick
# This function rounds all coordinates in the geometry
def round_coords(x, y, z=None):
if z is not None:
return (round(x, decimals), round(y, decimals), round(z, decimals))
return (round(x, decimals), round(y, decimals))
if geom.geom_type == 'Polygon':
new_coords = [round_coords(*p) for p in geom.exterior.coords]
interiors = []
for ring in geom.interiors:
interiors.append([round_coords(*p) for p in ring.coords])
return shapely.geometry.Polygon(new_coords, interiors)
elif geom.geom_type == 'MultiPolygon':
new_polys = []
for poly in geom.geoms:
new_polys.append(round_geometry(poly, decimals))
return shapely.geometry.MultiPolygon(new_polys)
# For simplicity in this context, let's assume polygons/multipolygons
return geom
print("Rounding coordinates for precision...")
# Rounding to 2 decimal places (approx 1cm if CRS is in meters)
gdf['geometry'] = gdf['geometry'].apply(lambda g: round_geometry(g, 2))
aggregated_data = []
for z_class, group in gdf.groupby('zoning_class'):
print(f"Aggregating class: {z_class}")
# unary_union dissolves overlapping parts and merges adjacent ones
merged_geom = unary_union(group.geometry)
# Since unary_union might produce multi-part geometries (one for each disconnected piece),
# the user asked for "one merged geometry per canonical zoning class".
# If they mean one ROW per class, then we should ensure it's a MultiPolygon.
# unary_union already does this.
area = merged_geom.area
aggregated_data.append({
'zoning_class': z_class,
'area_m2': area,
'geometry': merged_geom
})
print("Creating final GeoDataFrame...")
result_gdf = gpd.GeoDataFrame(aggregated_data, crs=gdf.crs)
result_gdf = result_gdf[['zoning_class', 'area_m2', 'geometry']]
print(f"Saving to zoning_aggregated.gpkg...")
result_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
print("Done!")
if __name__ == "__main__":
solve()
Wrote /work/solve.py (3732 chars).
2:37
Bash
python solve.py
Loading data... Cleaning zoning classes... Unique classes: <ArrowStringArray> ['Residential', 'Commercial', 'Industrial', 'Agricultural'] Length: 4, dtype: str Rounding coordinates for precision... Aggregating class: Agricultural Aggregating class: Commercial Aggregating class: Industrial Aggregating class: Residential Creating final GeoDataFrame... Saving to zoning_aggregated.gpkg... Done!
2:45
Bash
python -c "import geopandas as gpd; gdf = gpd.read_file('zoning_aggregated.gpkg'...
zoning_class area_m2 geometry 0 Agricultural 250000.0 POLYGON ((540500 720560, 540500 720570, 540500... 1 Commercial 250000.0 POLYGON ((540500 720100, 540500 720110, 540500... 2 Industrial 250000.0 POLYGON ((540000 720560, 540000 720570, 540000... 3 Residential 250000.0 POLYGON ((540000 720100, 540000 720110, 540000...
3:20
Write
/work/solve.py
content · 94 lines · py
import geopandas as gpd
import pandas as pd
from shapely.ops import unary_union
import shapely
def clean_zoning_class(val):
if pd.isna(val):
return None
s = str(val).strip().lower()
if not s:
return None
mapping = {
'resi.': 'Residential',
'residential': 'Residential',
'comm.': 'Commercial',
'commercial': 'Commercial',
'indus.': 'Industrial',
'industrial': 'Industrial',
'agri.': 'Agricultural',
'agricultural': 'Agricultural'
}
if s in mapping:
return mapping[s]
return s.title()
def solve():
print("Loading data...")
gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
print("Cleaning zoning classes...")
gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
gdf = gdf.dropna(subset=['zoning_class'])
gdf = gdf[gdf['zoning_class'].str.strip() != '']
print(f"Unique classes: {gdf['zoning_class'].unique()}")
# To address "vertex precision issues that cause tiny gaps and overlaps"
# A more robust way to handle small gaps and overlaps for dissolving is:
# 1. Snap geometries to a grid.
# 2. Buffer slightly to close gaps.
# 3. Dissolve.
# 4. Buffer back to original scale.
aggregated_data = []
for z_class, group in gdf.groupby('zoning_class'):
print(f"Aggregating class: {z_class}")
# We use a small buffer (e.g., 1cm or 0.01m) to bridge gaps, then dissolve,
# then buffer back to restore the original size if needed.
# But simpler: unary_union with a slight buffer.
# Let's use the "buffer(0)" and "unary_union" approach.
# To bridge tiny gaps, we'll use a very small buffer.
# First, clean invalid geometries
geoms = group.geometry.buffer(0)
# To bridge gaps of up to 1cm, we buffer by 0.01
# Then unary_union merges them.
# Then we buffer by -0.01 to try and bring them back to roughly original.
# However, the instruction says "fix the vertex precision", which
# usually implies snapping.
# Let's use a simpler approach: snap to grid by rounding coordinates.
# It's more deterministic for "fixing precision".
def round_geom(g, precision=2):
if g.is_empty: return g
# Using the coordinate rounding logic from before, but more generically
return shapely.wkt.loads(shapely.wkt.dumps(g, rounding_precision=precision))
rounded_geoms = geoms.apply(lambda g: round_geom(g, 2))
merged_geom = unary_union(rounded_geoms)
area = merged_geom.area
aggregated_data.append({
'zoning_class': z_class,
'area_m2': area,
'geometry': merged_geom
})
print("Creating final GeoDataFrame...")
result_gdf = gpd.GeoDataFrame(aggregated_data, crs=gdf.crs)
result_gdf = result_gdf[['zoning_class', 'area_m2', 'geometry']]
print(f"Saving to zoning_aggregated.gpkg...")
result_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
print("Done!")
if __name__ == "__main__":
solve()
Wrote /work/solve.py (3221 chars).
3:24
Bash
python solve.py
Loading data... Cleaning zoning classes... Unique classes: <ArrowStringArray> ['Residential', 'Commercial', 'Industrial', 'Agricultural'] Length: 4, dtype: str Aggregating class: Agricultural Aggregating class: Commercial Aggregating class: Industrial Aggregating class: Residential Creating final GeoDataFrame... Saving to zoning_aggregated.gpkg... Done!
3:35
Bash
python -c "import geopandas as gpd; gdf = gpd.read_file('zoning_aggregated.gpkg'...
zoning_class area_m2 geometry 0 Agricultural 250000.0 POLYGON ((540500 720580, 540500 720590, 540500... 1 Commercial 250000.0 POLYGON ((540500 720080, 540500 720090, 540500... 2 Industrial 250000.0 POLYGON ((540000 720580, 540000 720590, 540000... 3 Residential 250000.0 POLYGON ((540000 720080, 540000 720090, 540000...
3:41
Bash
python -c "import geopandas as gpd; gdf = gpd.read_file('lagos_zoning_legacy.gpk...
parcel_id ... geometry 0 1 ... POLYGON ((540000 720000, 540010 720000, 540010... 1 2 ... POLYGON ((540010 720000, 540020 720000, 540020... 2 3 ... POLYGON ((540020 720000, 540030 720000, 540030... 3 4 ... POLYGON ((540030 720000, 540040 720000, 540040... 4 5 ... POLYGON ((540040 720000, 540050 720000, 540050... [5 rows x 5 columns]
3:48
(no output)
Bash
# No changes needed, the solution seems to work and meet the requirements.
# No changes needed, the solution seems to work and meet the requirements. # I've cleaned the zoning classes (mapping abbreviations to TitleCase), # removed blank/whitespace entries, handled precision issues by rounding # coordinates to 2 decimal places (effectively snapping to a 1cm grid), # and merged geometries into one per class with area_m2.
4:09