dd-l2-bangkok-multicrs-audit

Status: done
Score: 88
Duration: 2:17
Cost: 0.34¢
Model: google/gemma-4-26b-a4b-it

Map (reference ← swipe → agent)

0:00
A contractor just shipped `bangkok_contractor_delivery.gpkg` and I need a defect cite-sheet before I sign off. Please walk every layer in the file and write one row per layer to `crs_audit.csv` with these columns: layer_name, declared_crs (formatted as `EPSG:NNNN`), geometry_type, feature_count, sample_x and sample_y for a sample coordinate taken from the layer as it sits on disk, and encoding_detected. For encoding_detected, look at the text attributes for anomalies and report either `latin1-mojibake` or `utf-8`.
0:00
System info
{
  "adapter": "openrouter",
  "model": "google/gemma-4-26b-a4b-it",
  "base_url": "https://openrouter.ai/api/v1"
}
0:07
Bash ls -F
ls -F
bangkok_contractor_delivery.gpkg
0:44
Write /work/solve.py
content · 86 lines · py
import geopandas as gpd
import pandas as pd
import fiona
from shapely.geometry import Point

def check_encoding(gdf):
    """
    Check if the text attributes contain mojibake by trying to encode/decode.
    A simple way is to see if common utf-8 strings look weird or if we can 
    detect non-utf8 patterns. However, since geopandas reads it, it's already decoded.
    To check for 'latin1-mojibake', we can check if characters like 'é' appear
    where 'é' should be.
    """
    # Check all string columns
    str_cols = gdf.select_dtypes(include=['object']).columns
    for col in str_cols:
        for val in gdf[col].dropna().astype(str):
            # A common sign of mojibake (UTF-8 interpreted as Latin-1) 
            # is sequences like à followed by something.
            # Let's look for the presence of the replacement character or common mojibake patterns.
            # In a real scenario, this is heuristic.
            if 'Ã' in val:
                return 'latin1-mojibake'
    return 'utf-8'

def audit_gpkg(file_path):
    layers = fiona.listlayers(file_path)
    audit_results = []

    for layer_name in layers:
        gdf = gpd.read_file(file_path, layer=layer_name)
        
        # CRS
        crs = gdf.crs
        if crs is None:
            declared_crs = "None"
        else:
            # Format as EPSG:NNNN
            try:
                declared_crs = f"EPSG:{crs.to_epsg()}"
            except:
                declared_crs = str(crs)

        # Geometry Type
        if not gdf.empty:
            geometry_type = gdf.geometry.type.iloc[0]
        else:
            geometry_type = "Unknown"

        # Feature Count
        feature_count = len(gdf)

        # Sample coordinates
        sample_x, sample_y = None, None
        if feature_count > 0:
            # Take the first geometry's centroid or point
            first_geom = gdf.geometry.iloc[0]
            if first_geom.geom_type == 'Point':
                sample_x, sample_y = first_geom.x, first_geom.y
            elif first_geom.geom_type in ['LineString', 'MultiLineString', 'Polygon', 'MultiPolygon']:
                centroid = first_geom.centroid
                sample_x, sample_y = centroid.x, centroid.y
            else:
                # fallback
                sample_x, sample_y = None, None

        # Encoding detection
        encoding = check_encoding(gdf)

        audit_results.append({
            'layer_name': layer_name,
            'declared_crs': declared_crs,
            'geometry_type': geometry_type,
            'feature_count': feature_count,
            'sample_x': sample_x,
            'sample_y': sample_y,
            'encoding_detected': encoding
        })

    df_audit = pd.DataFrame(audit_results)
    df_audit.to_csv('crs_audit.csv', index=False)
    print("Audit complete. Saved to crs_audit.csv")

if __name__ == "__main__":
    audit_gpkg('bangkok_contractor_delivery.gpkg')
Wrote /work/solve.py (2928 chars).
0:51
Bash python solve.py
python solve.py
Audit complete. Saved to crs_audit.csv
/work/solve.py:15: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  str_cols = gdf.select_dtypes(include=['object']).columns
/work/solve.py:15: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  str_cols = gdf.select_dtypes(include=['object']).columns
/work/solve.py:15: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  str_cols = gdf.select_dtypes(include=['object']).columns
1:02
Bash cat crs_audit.csv
cat crs_audit.csv
layer_name,declared_crs,geometry_type,feature_count,sample_x,sample_y,encoding_detected
parcels,EPSG:24047,Polygon,4000,657706.4571303247,1509274.702150785,utf-8
roads,EPSG:32647,LineString,5000,656996.3371533038,1509487.963830215,utf-8
markets,EPSG:4326,Point,1000,100.4533,13.655000000000001,utf-8
1:48
Write /work/solve.py
content · 90 lines · py
import geopandas as gpd
import pandas as pd
import fiona
from shapely.geometry import Point

def check_encoding(gdf):
    """
    Check if the text attributes contain mojibake by trying to encode/decode.
    A simple way is to see if common utf-8 strings look weird or if we can 
    detect non-utf8 patterns. However, since geopandas reads it, it's already decoded.
    To check for 'latin1-mojibake', we can check if characters like 'é' appear
    where 'é' should be.
    """
    # Check all string columns
    str_cols = gdf.select_dtypes(include=['object']).columns
    for col in str_cols:
        for val in gdf[col].dropna().astype(str):
            # A common sign of mojibake (UTF-8 interpreted as Latin-1) 
            # is sequences like à followed by something.
            # Let's look for the presence of the replacement character or common mojibake patterns.
            # In a real scenario, this is heuristic.
            if 'Ã' in val:
                return 'latin1-mojibake'
    return 'utf-8'

def audit_gpkg(file_path):
    layers = fiona.listlayers(file_path)
    audit_results = []

    for layer_name in layers:
        gdf = gpd.read_file(file_path, layer=layer_name)
        
        # CRS
        crs = gdf.crs
        if crs is None:
            declared_crs = "None"
        else:
            # Format as EPSG:NNNN
            try:
                # If the CRS has an EPSG code, use it
                epsg = crs.to_epsg()
                if epsg:
                    declared_crs = f"EPSG:{epsg}"
                else:
                    declared_crs = str(crs)
            except:
                declared_crs = str(crs)

        # Geometry Type
        if not gdf.empty:
            # Get the actual geometry type of the first feature
            # (might be MultiPolygon, etc.)
            geometry_type = gdf.geometry.iloc[0].geom_type
        else:
            geometry_type = "Unknown"

        # Feature Count
        feature_count = len(gdf)

        # Sample coordinates
        sample_x, sample_y = None, None
        if feature_count > 0:
            # Take the first geometry's centroid or point
            first_geom = gdf.geometry.iloc[0]
            if first_geom.geom_type == 'Point':
                sample_x, sample_y = first_geom.x, first_geom.y
            else:
                centroid = first_geom.centroid
                sample_x, sample_y = centroid.x, centroid.y

        # Encoding detection
        encoding = check_encoding(gdf)

        audit_results.append({
            'layer_name': layer_name,
            'declared_crs': declared_crs,
            'geometry_type': geometry_type,
            'feature_count': feature_count,
            'sample_x': sample_x,
            'sample_y': sample_y,
            'encoding_detected': encoding
        })

    df_audit = pd.DataFrame(audit_results)
    df_audit.to_csv('crs_audit.csv', index=False)
    print("Audit complete. Saved to crs_audit.csv")

if __name__ == "__main__":
    audit_gpkg('bangkok_contractor_delivery.gpkg')
Wrote /work/solve.py (3040 chars).
1:50
Bash python solve.py
python solve.py
Audit complete. Saved to crs_audit.csv
/work/solve.py:15: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  str_cols = gdf.select_dtypes(include=['object']).columns
/work/solve.py:15: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  str_cols = gdf.select_dtypes(include=['object']).columns
/work/solve.py:15: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  str_cols = gdf.select_dtypes(include=['object']).columns
1:54
Bash cat crs_audit.csv
cat crs_audit.csv
layer_name,declared_crs,geometry_type,feature_count,sample_x,sample_y,encoding_detected
parcels,EPSG:24047,Polygon,4000,657706.4571303247,1509274.702150785,utf-8
roads,EPSG:32647,LineString,5000,656996.3371533038,1509487.963830215,utf-8
markets,EPSG:4326,Point,1000,100.4533,13.655000000000001,utf-8
2:16