dd-l2-bangkok-multicrs-audit

Status: done
Score: 100
Duration: 1:02
Cost: 0.60¢
Model: google/gemma-4-26b-a4b-it

Map (reference ← swipe → agent)

0:00
A contractor just shipped `bangkok_contractor_delivery.gpkg` and I need a defect cite-sheet before I sign off. Please walk every layer in the file and write one row per layer to `crs_audit.csv` with these columns: layer_name, declared_crs (formatted as `EPSG:NNNN`), geometry_type, feature_count, sample_x and sample_y for a sample coordinate taken from the layer as it sits on disk, and encoding_detected. For encoding_detected, look at the text attributes for anomalies and report either `latin1-mojibake` or `utf-8`.
0:00
System info
{
  "adapter": "openrouter",
  "model": "google/gemma-4-26b-a4b-it",
  "base_url": "https://openrouter.ai/api/v1"
}
0:02
Bash ls -l /work/bangkok_contractor_delivery.gpkg
-rw-r--r-- 1 agent agent 2703360 Jun 18 10:26 /work/bangkok_contractor_delivery.gpkg
0:19
Write /work/solve.py
content · 101 lines · py
import geopandas as gpd
import pandas as pd
import fiona
from shapely.geometry import Point

def detect_encoding(df):
    """
    Heuristic to check if text columns contain mojibake.
    We check if common UTF-8 characters are represented as multiple latin1 chars.
    However, a simpler way is to check if the data can be encoded to utf-8 
    and if it contains common mojibake patterns.
    For this task, we'll check if any string column contains characters that 
    look like latin1-mojibake (e.g., Ã, ©, etc. where they shouldn't be).
    Actually, the prompt asks to report 'latin1-mojibake' or 'utf-8'.
    """
    text_cols = df.select_dtypes(include=['object']).columns
    if len(text_cols) == 0:
        return 'utf-8'
    
    for col in text_cols:
        # Sample some values
        samples = df[col].dropna().astype(str).head(100).tolist()
        for s in samples:
            # A common sign of mojibake is the presence of characters like 'Ã' 
            # followed by other symbols that result from UTF-8 interpreted as latin1.
            # We'll check if we can encode it as utf-8 (it should always be possible for strings in Python)
            # and then see if it "looks" like it was incorrectly decoded.
            # A simple check: if it contains sequences like 'é' instead of 'é'
            # This is a bit naive but works for many cases.
            try:
                # Try to see if it's valid utf-8
                s.encode('utf-8')
                
                # Heuristic: if we see "Ã" followed by non-ascii, it's often mojibake
                # (where a single utf-8 char was read as multiple latin1 chars)
                # We search for the pattern of common mojibake.
                import re
                if re.search(r'[ÃÂÊÎÒÛ][ÂÃÊÎÒÛ]', s):
                    return 'latin1-mojibake'
            except UnicodeEncodeError:
                continue
    return 'utf-8'

def audit_gpkg(file_path, output_csv):
    layers = fiona.listlayers(file_path)
    audit_results = []

    for layer in layers:
        gdf = gpd.read_file(file_path, layer=layer)
        
        # CRS
        crs = gdf.crs
        epsg_code = f"EPSG:{crs.to_epsg()}" if crs and crs.to_epsg() else "Unknown"
        
        # Geometry Type
        if not gdf.empty:
            geom_type = gdf.geometry.type.iloc[0]
        else:
            geom_type = "None"
            
        # Feature Count
        feature_count = len(gdf)
        
        # Sample X and Y
        sample_x = None
        sample_y = None
        if not gdf.empty:
            sample_geom = gdf.geometry.iloc[0]
            if sample_geom.geom_type == 'Point':
                sample_x = sample_geom.x
                sample_y = sample_geom.y
            elif sample_geom.geom_type in ['LineString', 'MultiLineString']:
                # Use centroid or start point
                centroid = sample_geom.centroid
                sample_x = centroid.x
                sample_y = centroid.y
            elif sample_geom.geom_type in ['Polygon', 'MultiPolygon']:
                centroid = sample_geom.centroid
                sample_x = centroid.x
                sample_y = centroid.y

        # Encoding
        encoding = detect_encoding(gdf)
        
        audit_results.append({
            'layer_name': layer,
            'declared_crs': epsg_code,
            'geometry_type': geom_type,
            'feature_count': feature_count,
            'sample_x': sample_x,
            'sample_y': sample_y,
            'encoding_detected': encoding
        })

    df_audit = pd.DataFrame(audit_results)
    df_audit.to_csv(output_csv, index=False)
    print(f"Audit completed. Results saved to {output_csv}")

if __name__ == "__main__":
    audit_gpkg('/work/bangkok_contractor_delivery.gpkg', '/work/crs_audit.csv')
Wrote /work/solve.py (3825 chars).
0:20
Bash python /work/solve.py
Audit completed. Results saved to /work/crs_audit.csv
/work/solve.py:16: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = df.select_dtypes(include=['object']).columns
/work/solve.py:16: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = df.select_dtypes(include=['object']).columns
/work/solve.py:16: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = df.select_dtypes(include=['object']).columns
0:22
Bash cat /work/crs_audit.csv
layer_name,declared_crs,geometry_type,feature_count,sample_x,sample_y,encoding_detected
parcels,EPSG:24047,Polygon,4000,657706.4571303247,1509274.702150785,utf-8
roads,EPSG:32647,LineString,5000,656996.3371533038,1509487.963830215,utf-8
markets,EPSG:4326,Point,1000,100.4533,13.655000000000001,utf-8
0:31
Bash # I'll check the content of the layers to see if I missed any mojibake.
# I'll check the content of the layers to see if I missed any mojibake.
# I'll create a script to inspect the strings in each layer more closely.
cat << 'EOF' > /work/inspect_data.py
import geopandas as gpd
import fiona

def check_mojibake(gdf):
    text_cols = gdf.select_dtypes(include=['object']).columns
    for col in text_cols:
        for val in gdf[col].dropna().astype(str).head(500):
            # Check for patterns like à followed by non-ascii
            # A common mojibake for 'é' (UTF-8: C3 A9) in latin1 is 'é'
            # We can look for the sequence of characters that are commonly part of mojibake
            # If we find any "Ã" that is followed by something that is not part of a 
            # normal word but is part of a multibyte sequence represented as latin1.
            # For simplicity, we'll use the same heuristic as before but more aggressively.
            import re
            if re.search(r'[ÃÂÊÎÒÛ][ÂÃÊÎÒÛ]', val):
                return True
    return False

layers = fiona.listlayers('/work/bangkok_contractor_delivery.gpkg')
for layer in layers:
    gdf = gpd.read_file('/work/bangkok_contractor_delivery.gpkg', layer=layer)
    print(f"Layer: {layer}")
    if check_mojibake(gdf):
        print(f"  !!! Detected mojibake in {layer}")
    else:
        print(f"  Clean")
    # Print a few rows of object columns to be sure
    obj_cols = gdf.select_dtypes(include=['object']).columns
    if not obj_cols.empty:
        print(f"  Sample text: {gdf[obj_cols[0]].head(3).tolist()}")

EOF
python /work/inspect_data.py
Layer: parcels
  Clean
  Sample text: ['P00000', 'P00001', 'P00002']
Layer: roads
  Clean
  Sample text: ['R00000', 'R00001', 'R00002']
Layer: markets
  Clean
  Sample text: ['M00000', 'M00001', 'M00002']
/work/inspect_data.py:5: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = gdf.select_dtypes(include=['object']).columns
/work/inspect_data.py:28: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  obj_cols = gdf.select_dtypes(include=['object']).columns
/work/inspect_data.py:5: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = gdf.select_dtypes(include=['object']).columns
/work/inspect_data.py:28: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  obj_cols = gdf.select_dtypes(include=['object']).columns
/work/inspect_data.py:5: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = gdf.select_dtypes(include=['object']).columns
/work/inspect_data.py:28: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  obj_cols = gdf.select_dtypes(include=['object']).columns
0:36
Bash # Let's check if any layer has actual text content that is not just IDs.
# Let's check if any layer has actual text content that is not just IDs.
# I'll modify the inspection script to show more columns and content.
cat << 'EOF' > /work/inspect_data.py
import geopandas as gpd
import fiona

layers = fiona.listlayers('/work/bangkok_contractor_delivery.gpkg')
for layer in layers:
    gdf = gpd.read_file('/work/bangkok_contractor_delivery.gpkg', layer=layer)
    print(f"--- Layer: {layer} ---")
    print(f"Columns: {gdf.columns.tolist()}")
    # Show first 5 rows and all columns
    print(gdf.head())
    
    # Specifically check for strings that look like they might be encoded incorrectly
    text_cols = gdf.select_dtypes(include=['object']).columns
    for col in text_cols:
        print(f"Sample from {col}: {gdf[col].head(5).tolist()}")
EOF
python /work/inspect_data.py
--- Layer: parcels ---
Columns: ['id', 'name_th', 'owner_class', 'geometry']
       id  ...                                           geometry
0  P00000  ...  POLYGON ((657421.516 1509162.37, 657992.722 15...
1  P00001  ...  POLYGON ((658135.524 1509166.644, 658706.731 1...
2  P00002  ...  POLYGON ((658849.533 1509170.937, 659420.743 1...
3  P00003  ...  POLYGON ((659563.545 1509175.249, 660134.756 1...
4  P00004  ...  POLYGON ((660277.558 1509179.581, 660848.771 1...

[5 rows x 4 columns]
Sample from id: ['P00000', 'P00001', 'P00002', 'P00003', 'P00004']
Sample from name_th: ['à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 à¸\x9aาà¸\x87รัà¸\x81', 'à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 à¸\x9bà¸\x97ุมวัà¸\x99', 'à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 วัà¸\x92à¸\x99า', 'à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 à¸\x84ลà¸\xadà¸\x87à¹\x80à¸\x95ย', 'à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 à¸\x94ิà¸\x99à¹\x81à¸\x94à¸\x87']
Sample from owner_class: ['private', 'state', 'religious', 'private', 'state']
--- Layer: roads ---
Columns: ['id', 'name', 'highway', 'geometry']
       id  ...                                           geometry
0  R00000  ...  LINESTRING (656835.685 1509487.003, 657156.989...
1  R00001  ...  LINESTRING (657192.69 1509489.138, 657190.305 ...
2  R00002  ...  LINESTRING (657549.695 1509491.278, 657871 150...
3  R00003  ...  LINESTRING (657906.701 1509493.423, 657904.305...
4  R00004  ...  LINESTRING (658263.707 1509495.572, 658585.013...

[5 rows x 4 columns]
Sample from id: ['R00000', 'R00001', 'R00002', 'R00003', 'R00004']
Sample from name: ['à¸\x96à¸\x99à¸\x99สุà¸\x82ุมวิà¸\x97', 'à¸\x96à¸\x99à¸\x99à¸\x9eระราม 4', 'à¸\x96à¸\x99à¸\x99à¸\x9eระราม 9', 'à¸\x96à¸\x99à¸\x99รัà¸\x8aà¸\x94าà¸\xa0ิà¹\x80ษà¸\x81', 'à¸\x96à¸\x99à¸\x99à¹\x80à¸\x9eà¸\x8aรà¸\x9aุรี']
Sample from highway: ['primary', 'secondary', 'tertiary', 'residential', 'service']
--- Layer: markets ---
Columns: ['id', 'name_th', 'kind', 'geometry']
       id         name_th   kind                 geometry
0  M00000  ตลาดนัดจตุจักร  fresh  POINT (100.4533 13.655)
1  M00001      ตลาดอ.ต.ก.    wet  POINT (100.4599 13.655)
2  M00002     ตลาดคลองเตย  night  POINT (100.4665 13.655)
3  M00003    ตลาดสะพานขาว  fresh  POINT (100.4731 13.655)
4  M00004      ตลาดบางรัก    wet  POINT (100.4797 13.655)
Sample from id: ['M00000', 'M00001', 'M00002', 'M00003', 'M00004']
Sample from name_th: ['ตลาดนัดจตุจักร', 'ตลาดอ.ต.ก.', 'ตลาดคลองเตย', 'ตลาดสะพานขาว', 'ตลาดบางรัก']
Sample from kind: ['fresh', 'wet', 'night', 'fresh', 'wet']
/work/inspect_data.py:13: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = gdf.select_dtypes(include=['object']).columns
/work/inspect_data.py:13: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = gdf.select_dtypes(include=['object']).columns
/work/inspect_data.py:13: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = gdf.select_dtypes(include=['object']).columns
0:49
Write /work/solve.py
content · 83 lines · py
import geopandas as gpd
import pandas as pd
import fiona
import re

def detect_encoding(df):
    """
    Heuristic to check if text columns contain mojibake.
    """
    text_cols = df.select_dtypes(include=['object']).columns
    if len(text_cols) == 0:
        return 'utf-8'
    
    for col in text_cols:
        samples = df[col].dropna().astype(str).head(100).tolist()
        for s in samples:
            # The inspection showed things like 'à¹\x81à¸\x9bà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99'
            # These are typical of UTF-8 bytes being interpreted as latin1.
            # 'à¹' is a common start for Thai characters in mojibake.
            # A simpler way: if the string contains characters in the range \x80-\x9f 
            # or other high-bit characters that look like they were meant to be multi-byte.
            # We'll look for the 'à' followed by non-ascii characters pattern.
            if re.search(r'à[^\x00-\x7f]', s):
                return 'latin1-mojibake'
    return 'utf-8'

def audit_gpkg(file_path, output_csv):
    layers = fiona.listlayers(file_path)
    audit_results = []

    for layer in layers:
        gdf = gpd.read_file(file_path, layer=layer)
        
        # CRS
        crs = gdf.crs
        epsg_code = f"EPSG:{crs.to_epsg()}" if crs and crs.to_epsg() else "Unknown"
        
        # Geometry Type
        if not gdf.empty:
            geom_type = gdf.geometry.type.iloc[0]
        else:
            geom_type = "None"
            
        # Feature Count
        feature_count = len(gdf)
        
        # Sample X and Y
        sample_x = None
        sample_y = None
        if not gdf.empty:
            sample_geom = gdf.geometry.iloc[0]
            if sample_geom.geom_type == 'Point':
                sample_x = sample_geom.x
                sample_y = sample_geom.y
            elif sample_geom.geom_type in ['LineString', 'MultiLineString']:
                centroid = sample_geom.centroid
                sample_x = centroid.x
                sample_y = centroid.y
            elif sample_geom.geom_type in ['Polygon', 'MultiPolygon']:
                centroid = sample_geom.centroid
                sample_x = centroid.x
                sample_y = centroid.y

        # Encoding
        encoding = detect_encoding(gdf)
        
        audit_results.append({
            'layer_name': layer,
            'declared_crs': epsg_code,
            'geometry_type': geom_type,
            'feature_count': feature_count,
            'sample_x': sample_x,
            'sample_y': sample_y,
            'encoding_detected': encoding
        })

    df_audit = pd.DataFrame(audit_results)
    df_audit.to_csv(output_csv, index=False)
    print(f"Audit completed. Results saved to {output_csv}")

if __name__ == "__main__":
    audit_gpkg('/work/bangkok_contractor_delivery.gpkg', '/work/crs_audit.csv')
Wrote /work/solve.py (2863 chars).
0:50
Bash python /work/solve.py
Audit completed. Results saved to /work/crs_audit.csv
/work/solve.py:10: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = df.select_dtypes(include=['object']).columns
/work/solve.py:10: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = df.select_dtypes(include=['object']).columns
/work/solve.py:10: Pandas4Warning: For backward compatibility, 'str' dtypes are included by select_dtypes when 'object' dtype is specified. This behavior is deprecated and will be removed in a future version. Explicitly pass 'str' to `include` to select them, or to `exclude` to remove them and silence this warning.
See https://pandas.pydata.org/docs/user_guide/migration-3-strings.html#string-migration-select-dtypes for details on how to write code that works with pandas 2 and 3.
  text_cols = df.select_dtypes(include=['object']).columns
0:53
Bash cat /work/crs_audit.csv
layer_name,declared_crs,geometry_type,feature_count,sample_x,sample_y,encoding_detected
parcels,EPSG:24047,Polygon,4000,657706.4571303247,1509274.702150785,latin1-mojibake
roads,EPSG:32647,LineString,5000,656996.3371533038,1509487.963830215,latin1-mojibake
markets,EPSG:4326,Point,1000,100.4533,13.655000000000001,utf-8
1:00