dc-l2-lagos-snap-normalize

Status: done
Score: 75
Duration: 1:00
Cost: 0.52¢
Model: google/gemma-4-26b-a4b-it

Map (reference ← swipe → agent)

0:00
I need to get our legacy zoning layer ready for the state portal. lagos_zoning_legacy.gpkg has inconsistent zoning_class values (mixed casing, spelling variations, and some blank entries) plus vertex precision issues that cause tiny gaps and overlaps between adjacent parcels. Can you clean up the class names to consistent canonical TitleCase values, drop any rows with blank or whitespace-only classes, fix the vertex precision problems, and produce one merged geometry per canonical zoning class? Please write the result to zoning_aggregated.gpkg, with each row carrying two columns: zoning_class and a numeric area_m2.
0:00
System info
{
  "adapter": "openrouter",
  "model": "google/gemma-4-26b-a4b-it",
  "base_url": "https://openrouter.ai/api/v1"
}
0:01
Bash ls -F
lagos_zoning_legacy.gpkg
0:03
Write /work/inspect.py
content · 16 lines · py
import geopandas as gpd

def inspect_data(file_path):
    gdf = gpd.read_file(file_path)
    print("Columns:", gdf.columns.tolist())
    print("First 5 rows of zoning_class:")
    print(gdf['zoning_class'].head())
    print("\nUnique values in zoning_class:")
    print(gdf['zoning_class'].unique())
    print("\nCRS:", gdf.crs)
    print("\nGeometry type:", gdf.geometry.type.unique())
    print("\nShape info:", gdf.shape)

if __name__ == "__main__":
    inspect_data('lagos_zoning_legacy.gpkg')
Wrote /work/inspect.py (498 chars).
0:04
Bash python inspect.py
Traceback (most recent call last):
  File "/work/inspect.py", line 1, in <module>
    import geopandas as gpd
  File "/usr/local/lib/python3.14/site-packages/geopandas/__init__.py", line 3, in <module>
    from geopandas.geoseries import GeoSeries
  File "/usr/local/lib/python3.14/site-packages/geopandas/geoseries.py", line 7, in <module>
    import numpy as np
  File "/usr/local/lib/python3.14/site-packages/numpy/__init__.py", line 112, in <module>
    from numpy.__config__ import show_config
  File "/usr/local/lib/python3.14/site-packages/numpy/__config__.py", line 4, in <module>
    from numpy._core._multiarray_umath import (
  File "/usr/local/lib/python3.14/site-packages/numpy/_core/__init__.py", line 24, in <module>
    from . import multiarray
  File "/usr/local/lib/python3.14/site-packages/numpy/_core/multiarray.py", line 115, in <module>
    @array_function_from_c_func_and_dispatcher(_multiarray_umath.empty_like)
  File "/usr/local/lib/python3.14/site-packages/numpy/_core/overrides.py", line 185, in decorator
    return array_function_dispatch(
  File "/usr/local/lib/python3.14/site-packages/numpy/_core/overrides.py", line 161, in decorator
    doc = inspect.cleandoc(dispatcher.__doc__)
AttributeError: module 'inspect' has no attribute 'cleandoc' (consider renaming '/work/inspect.py' since it has the same name as the standard library module named 'inspect' and prevents importing that standard library module)
0:05
Bash rm /work/inspect.py
(no output)
0:07
Write /work/data_inspect.py
content · 16 lines · py
import geopandas as gpd

def inspect_data(file_path):
    gdf = gpd.read_file(file_path)
    print("Columns:", gdf.columns.tolist())
    print("First 5 rows of zoning_class:")
    print(gdf['zoning_class'].head())
    print("\nUnique values in zoning_class:")
    print(gdf['zoning_class'].unique())
    print("\nCRS:", gdf.crs)
    print("\nGeometry type:", gdf.geometry.type.unique())
    print("\nShape info:", gdf.shape)

if __name__ == "__main__":
    inspect_data('lagos_zoning_legacy.gpkg')
Wrote /work/data_inspect.py (498 chars).
0:08
Bash python data_inspect.py
Columns: ['parcel_id', 'lga_source', 'zoning_class', 'area_m2', 'geometry']
First 5 rows of zoning_class:
0    RESIDENTIAL
1    residential
2          Resi.
3          resi.
4          RESI.
Name: zoning_class, dtype: str

Unique values in zoning_class:
<ArrowStringArray>
[ 'RESIDENTIAL',  'residential',        'Resi.',        'resi.',
        'RESI.',  'Residential',        'Comm.',        'comm.',
        'COMM.',   'Commercial',   'COMMERCIAL',   'commercial',
       'Indus.',       'indus.',       'INDUS.',   'Industrial',
   'INDUSTRIAL',   'industrial',        'AGRI.', 'Agricultural',
 'AGRICULTURAL', 'agricultural',        'Agri.',        'agri.',
             '',          '   ',            nan,           '\t']
Length: 28, dtype: str

CRS: EPSG:26331

Geometry type: <ArrowStringArray>
['Polygon']
Length: 1, dtype: str

Shape info: (10080, 5)
0:17
Write /work/solve.py
content · 88 lines · py
import geopandas as gpd
import pandas as pd
from shapely.geometry import MultiPolygon, Polygon
from shapely import snap, simplify

def clean_zoning_class(val):
    if pd.isna(val):
        return None
    s = str(val).strip().lower()
    if not s:
        return None
    
    # Mapping logic based on inspection
    if s in ['residential', 'resi.']:
        return 'Residential'
    elif s in ['commercial', 'comm.']:
        return 'Commercial'
    elif s in ['industrial', 'indus.']:
        return 'Industrial'
    elif s in ['agricultural', 'agri.']:
        return 'Agricultural'
    else:
        # Fallback to Title Case if it doesn't match known patterns
        return s.title()

def solve():
    print("Loading data...")
    gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
    
    print("Cleaning zoning_class...")
    # 1. Clean class names
    gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
    
    # 2. Drop rows with None/NaN/empty classes
    gdf = gdf.dropna(subset=['zoning_class'])
    gdf = gdf[gdf['zoning_class'].str.strip() != '']
    
    print(f"Remaining rows after cleaning classes: {len(gdf)}")
    print("Unique classes:", gdf['zoning_class'].unique())

    # 3. Fix vertex precision problems (gaps and overlaps)
    # A common technique is to snap geometries to a grid or to each other.
    # Since we want to merge them anyway, we can simplify/snap first.
    # Given they are adjacent, snapping to a grid is often effective.
    print("Fixing vertex precision...")
    # Snapping to a grid of 1cm (0.01m) if EPSG:26331 (metric)
    # EPSG:26331 is likely a UTM projection.
    gdf['geometry'] = gdf['geometry'].apply(lambda geom: simplify(snap(geom, geom, 0.01), 0.01))
    
    # Alternatively, a more robust way to fix gaps/overlaps in adjacent polygons 
    # is to use a buffer(0) or a small negative/positive buffer, 
    # but for "merging" later, let's try to snap them.
    # Given the instruction, we'll use a small snap.
    
    # 4. Group by zoning_class and merge geometries
    print("Merging geometries per class...")
    
    # Ensure we have a metric CRS for area calculation
    # The inspect showed EPSG:26331 which is metric.
    
    results = []
    for z_class, group in gdf.groupby('zoning_class'):
        print(f"Processing class: {z_class}")
        # Dissolve geometries
        merged_geom = group.dissolve(by='zoning_class')['geometry'].iloc[0]
        
        # Calculate area
        area = merged_geom.area
        
        results.append({
            'zoning_class': z_class,
            'area_m2': area,
            'geometry': merged_geom
        })
    
    # Create the final GeoDataFrame
    final_gdf = gpd.GeoDataFrame(results, crs=gdf.crs)
    
    # Final check on geometry type (dissolve might produce MultiPolygons)
    # The prompt asks for "one merged geometry per canonical zoning class".
    
    print(f"Writing to zoning_aggregated.gpkg...")
    final_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
    print("Done!")

if __name__ == "__main__":
    solve()
Wrote /work/solve.py (3072 chars).
0:18
Bash python solve.py
Loading data...
Cleaning zoning_class...
Remaining rows after cleaning classes: 10030
Unique classes: <ArrowStringArray>
['Residential', 'Commercial', 'Industrial', 'Agricultural']
Length: 4, dtype: str
Fixing vertex precision...
Merging geometries per class...
Processing class: Agricultural
Processing class: Commercial
Processing class: Industrial
Processing class: Residential
Writing to zoning_aggregated.gpkg...
Done!
0:31
Bash python -c "import geopandas as gpd; gdf = gpd.read_file('zoning_aggregated.gpkg'...
   zoning_class  ...                                           geometry
0  Agricultural  ...  MULTIPOLYGON (((540500 720500, 540500 720510, ...
1    Commercial  ...  MULTIPOLYGON (((540510 720030, 540510 720030, ...
2    Industrial  ...  MULTIPOLYGON (((540000 720530, 540000 720540, ...
3   Residential  ...  MULTIPOLYGON (((540010 720000, 540000 720000, ...

[4 rows x 3 columns]
Index(['zoning_class', 'area_m2', 'geometry'], dtype='str')
EPSG:26331
0:47
Write /work/solve.py
content · 105 lines · py
import geopandas as gpd
import pandas as pd
from shapely.geometry import MultiPolygon, Polygon
from shapely import snap, simplify

def clean_zoning_class(val):
    if pd.isna(val):
        return None
    s = str(val).strip().lower()
    if not s:
        return None
    
    # Mapping logic based on inspection
    if s in ['residential', 'resi.']:
        return 'Residential'
    elif s in ['commercial', 'comm.']:
        return 'Commercial'
    elif s in ['industrial', 'indus.']:
        return 'Industrial'
    elif s in ['agricultural', 'agri.']:
        return 'Agricultural'
    else:
        # Fallback to Title Case if it doesn't match known patterns
        return s.title()

def solve():
    print("Loading data...")
    gdf = gpd.read_file('lagos_zoning_legacy.gpkg')
    
    print("Cleaning zoning_class...")
    # 1. Clean class names
    gdf['zoning_class'] = gdf['zoning_class'].apply(clean_zoning_class)
    
    # 2. Drop rows with None/NaN/empty classes
    gdf = gdf.dropna(subset=['zoning_class'])
    gdf = gdf[gdf['zoning_class'].str.strip() != '']
    
    print(f"Remaining rows after cleaning classes: {len(gdf)}")
    print("Unique classes:", gdf['zoning_class'].unique())

    # 3. Fix vertex precision problems (gaps and overlaps)
    print("Fixing vertex precision...")
    # Use a small snap tolerance to align vertices
    # Since we are dissolving later, the most important thing is to avoid
    # tiny gaps or slivers that might create tiny "islands" or holes.
    # A good approach for adjacent polygons is to snap everything to a grid.
    
    # First, round the coordinates to a certain precision to reduce vertex clutter
    def round_coords(geom, precision=0.01):
        if geom.is_empty:
            return geom
        # We use a trick: simplify with a very small tolerance or just round
        # For most GIS workflows, snapping to a grid is done via:
        return simplify(geom, tolerance=0.001) # Very small tolerance to clean up noise

    # Let's try snapping to a grid by rounding coordinates. 
    # This is often more robust for "merging" than snap(geom, geom).
    def snap_to_grid(geom, precision=0.01):
        if geom.is_empty:
            return geom
        # Using simplify with a tiny tolerance often cleans up precision issues
        return simplify(geom, tolerance=precision)

    # Apply a small snap/simplify to clean up the topology
    # We use 0.01m as the tolerance for EPSG:26331
    gdf['geometry'] = gdf['geometry'].apply(lambda g: simplify(g, 0.001))

    # 4. Group by zoning_class and merge geometries
    print("Merging geometries per class...")
    
    # To handle overlaps and gaps more effectively during dissolution:
    # We can use the 'dissolve' method which performs a union.
    # To mitigate the impact of gaps/overlaps, we can also use a tiny buffer.
    
    results = []
    for z_class, group in gdf.groupby('zoning_class'):
        print(f"Processing class: {z_class}")
        
        # Dissolve (union) the geometries in the group
        # We dissolve by 'zoning_class'
        dissolved = group.dissolve(by='zoning_class')
        merged_geom = dissolved.geometry.iloc[0]
        
        # Calculate area
        area = merged_geom.area
        
        results.append({
            'zoning_class': z_class,
            'area_m2': area,
            'geometry': merged_geom
        })
    
    # Create the final GeoDataFrame
    final_gdf = gpd.GeoDataFrame(results, crs=gdf.crs)
    
    # Ensure columns are in requested order
    final_gdf = final_gdf[['zoning_class', 'area_m2', 'geometry']]
    
    print(f"Writing to zoning_aggregated.gpkg...")
    final_gdf.to_file('zoning_aggregated.gpkg', driver='GPKG')
    print("Done!")

if __name__ == "__main__":
    solve()
Wrote /work/solve.py (3779 chars).
0:48
Bash python solve.py
Loading data...
Cleaning zoning_class...
Remaining rows after cleaning classes: 10030
Unique classes: <ArrowStringArray>
['Residential', 'Commercial', 'Industrial', 'Agricultural']
Length: 4, dtype: str
Fixing vertex precision...
Merging geometries per class...
Processing class: Agricultural
Processing class: Commercial
Processing class: Industrial
Processing class: Residential
Writing to zoning_aggregated.gpkg...
Done!
0:52
Bash python -c "import geopandas as gpd; gdf = gpd.read_file('zoning_aggregated.gpkg'...
   zoning_class  ...                                           geometry
0  Agricultural  ...  MULTIPOLYGON (((540500 720500, 540500 720510, ...
1    Commercial  ...  MULTIPOLYGON (((540510 720030, 540510 720030, ...
2    Industrial  ...  MULTIPOLYGON (((540000 720530, 540000 720540, ...
3   Residential  ...  MULTIPOLYGON (((540010 720000, 540000 720000, ...

[4 rows x 3 columns]
Index(['zoning_class', 'area_m2', 'geometry'], dtype='str')
EPSG:26331
zoning_class         str
area_m2          float64
geometry        geometry
dtype: object
0:58