dd-l2-bangkok-multicrs-audit

Status: done
Score: 100
Duration: 1:32
Cost: 0.33¢
Model: deepseek/deepseek-v4-flash

Map (reference ← swipe → agent)

0:00
A contractor just shipped `bangkok_contractor_delivery.gpkg` and I need a defect cite-sheet before I sign off. Please walk every layer in the file and write one row per layer to `crs_audit.csv` with these columns: layer_name, declared_crs (formatted as `EPSG:NNNN`), geometry_type, feature_count, sample_x and sample_y for a sample coordinate taken from the layer as it sits on disk, and encoding_detected. For encoding_detected, look at the text attributes for anomalies and report either `latin1-mojibake` or `utf-8`.
0:00
System info
{
  "adapter": "openrouter",
  "model": "deepseek/deepseek-v4-flash",
  "base_url": "https://openrouter.ai/api/v1"
}
0:03
Let me start by examining the GeoPackage file to understand its structure.
0:03
Bash cd /work && ls -la bangkok_contractor_delivery.gpkg
-rw-r--r-- 1 agent agent 2703360 Jun 18 23:23 bangkok_contractor_delivery.gpkg
0:10
Let me explore this GeoPackage file to understand its layers and structure.
0:10
Bash cd /work && python3 -c "
cd /work && python3 -c "
import geopandas as gpd
import sqlite3

# List all layers in the GeoPackage
conn = sqlite3.connect('bangkok_contractor_delivery.gpkg')
tables = conn.execute(\"SELECT name FROM sqlite_master WHERE type='table' AND name NOT LIKE 'gpkg_%' AND name NOT LIKE 'sqlite_%' ORDER BY name\").fetchall()
print('Tables:', [t[0] for t in tables])

# Also check gpkg_contents
contents = conn.execute('SELECT * FROM gpkg_contents').fetchall()
print('gpkg_contents:')
for row in contents:
    print(row)

# Check geometry columns
geom_cols = conn.execute('SELECT * FROM gpkg_geometry_columns').fetchall()
print('gpkg_geometry_columns:')
for row in geom_cols:
    print(row)

# Check spatial ref sys
srs = conn.execute('SELECT * FROM gpkg_spatial_ref_sys').fetchall()
print('gpkg_spatial_ref_sys:')
for row in srs:
    print(row)

conn.close()
"
Tables: ['markets', 'parcels', 'roads', 'rtree_markets_geom', 'rtree_markets_geom_node', 'rtree_markets_geom_parent', 'rtree_markets_geom_rowid', 'rtree_parcels_geom', 'rtree_parcels_geom_node', 'rtree_parcels_geom_parent', 'rtree_parcels_geom_rowid', 'rtree_roads_geom', 'rtree_roads_geom_node', 'rtree_roads_geom_parent', 'rtree_roads_geom_rowid']
gpkg_contents:
('parcels', 'features', 'parcels', '', '2026-05-08T11:52:38.974Z', 657288.7771434318, 1509162.370185999, 692981.5473445851, 1531471.526142154, 24047)
('roads', 'features', 'roads', '', '2026-05-08T11:52:39.063Z', 656705.2518471001, 1509487.003180475, 692181.5905333688, 1531805.742956211, 32647)
('markets', 'features', 'markets', '', '2026-05-08T11:52:39.104Z', 100.4533, 13.655, 100.7767, 13.845, 4326)
gpkg_geometry_columns:
('parcels', 'geom', 'POLYGON', 24047, 0, 0)
('roads', 'geom', 'LINESTRING', 32647, 0, 0)
('markets', 'geom', 'POINT', 4326, 0, 0)
gpkg_spatial_ref_sys:
('Undefined Cartesian SRS', -1, 'NONE', -1, 'undefined', 'undefined Cartesian coordinate reference system')
('Undefined geographic SRS', 0, 'NONE', 0, 'undefined', 'undefined geographic coordinate reference system')
('WGS 84 geodetic', 4326, 'EPSG', 4326, 'GEOGCS["WGS 84",DATUM["WGS_1984",SPHEROID["WGS 84",6378137,298.257223563,AUTHORITY["EPSG","7030"]],AUTHORITY["EPSG","6326"]],PRIMEM["Greenwich",0,AUTHORITY["EPSG","8901"]],UNIT["degree",0.0174532925199433,AUTHORITY["EPSG","9122"]],AXIS["Latitude",NORTH],AXIS["Longitude",EAST],AUTHORITY["EPSG","4326"]]', 'longitude/latitude coordinates in decimal degrees on the WGS 84 spheroid')
('Indian 1975 / UTM zone 47N', 24047, 'EPSG', 24047, 'PROJCS["Indian 1975 / UTM zone 47N",GEOGCS["Indian 1975",DATUM["Indian_1975",SPHEROID["Everest 1830 (1937 Adjustment)",6377276.345,300.8017,AUTHORITY["EPSG","7015"]],AUTHORITY["EPSG","6240"]],PRIMEM["Greenwich",0,AUTHORITY["EPSG","8901"]],UNIT["degree",0.0174532925199433,AUTHORITY["EPSG","9122"]],AUTHORITY["EPSG","4240"]],PROJECTION["Transverse_Mercator"],PARAMETER["latitude_of_origin",0],PARAMETER["central_meridian",99],PARAMETER["scale_factor",0.9996],PARAMETER["false_easting",500000],PARAMETER["false_northing",0],UNIT["metre",1,AUTHORITY["EPSG","9001"]],AXIS["Easting",EAST],AXIS["Northing",NORTH],AUTHORITY["EPSG","24047"]]', None)
('WGS 84 / UTM zone 47N', 32647, 'EPSG', 32647, 'PROJCS["WGS 84 / UTM zone 47N",GEOGCS["WGS 84",DATUM["WGS_1984",SPHEROID["WGS 84",6378137,298.257223563,AUTHORITY["EPSG","7030"]],AUTHORITY["EPSG","6326"]],PRIMEM["Greenwich",0,AUTHORITY["EPSG","8901"]],UNIT["degree",0.0174532925199433,AUTHORITY["EPSG","9122"]],AUTHORITY["EPSG","4326"]],PROJECTION["Transverse_Mercator"],PARAMETER["latitude_of_origin",0],PARAMETER["central_meridian",99],PARAMETER["scale_factor",0.9996],PARAMETER["false_easting",500000],PARAMETER["false_northing",0],UNIT["metre",1,AUTHORITY["EPSG","9001"]],AXIS["Easting",EAST],AXIS["Northing",NORTH],AUTHORITY["EPSG","32647"]]', None)
0:17
Interesting! So we have 3 data layers:

1. `parcels` - POLYGON, SRID 24047 (Indian 1975 / UTM zone 47N)
2. `roads` - LINESTRING, SRID 32647 (WGS 84 / UTM zone 47N)
3. `markets` - POINT, SRID 4326 (WGS 84)

Now I need to examine the data in each layer to:
1. Get feature_count
2. Get a sample coordinate (from the geometry column as it sits on disk)
3. Check text attributes for encoding anomalies (latin1-mojibake vs utf-8)

Let me write a comprehensive Python script.
0:17
Bash cd /work && python3 -c "
cd /work && python3 -c "
import geopandas as gpd
import sqlite3

# Inspect each layer
layers = ['parcels', 'roads', 'markets']
for layer in layers:
    gdf = gpd.read_file('bangkok_contractor_delivery.gpkg', layer=layer)
    print(f'=== {layer} ===')
    print(f'Shape: {gdf.shape}')
    print(f'Columns: {list(gdf.columns)}')
    print(f'CRS: {gdf.crs}')
    print(f'Geometry type: {gdf.geometry.type.unique()}')
    print(f'First few rows:')
    print(gdf.head(3).to_string())
    print()
    # Check non-geometry columns for text
    for col in gdf.columns:
        if col != 'geom' and col != 'geometry':
            print(f'Column {col}: dtype={gdf[col].dtype}')
            if gdf[col].dtype == 'object':
                vals = gdf[col].dropna().head(20).tolist()
                print(f'  Sample values: {vals}')
    print()
"
=== parcels ===
Shape: (4000, 4)
Columns: ['id', 'name_th', 'owner_class', 'geometry']
CRS: EPSG:24047
Geometry type: <ArrowStringArray>
['Polygon']
Length: 1, dtype: str
First few rows:
       id                                               name_th owner_class                                                                                                                            geometry
0  P00000     แปลงที่ดิน บางรัก     private    POLYGON ((657421.516 1509162.37, 657992.722 1509165.787, 657991.396 1509387.035, 657420.195 1509383.617, 657421.516 1509162.37))
1  P00001  แปลงที่ดิน ปทุมวัน       state  POLYGON ((658135.524 1509166.644, 658706.731 1509170.077, 658705.399 1509391.324, 658134.196 1509387.891, 658135.524 1509166.644))
2  P00002        แปลงที่ดิน วัฒนา   religious    POLYGON ((658849.533 1509170.937, 659420.743 1509174.385, 659419.404 1509395.633, 658848.2 1509392.185, 658849.533 1509170.937))

Column id: dtype=str
Column name_th: dtype=str
Column owner_class: dtype=str

=== roads ===
Shape: (5000, 4)
Columns: ['id', 'name', 'highway', 'geometry']
CRS: EPSG:32647
Geometry type: <ArrowStringArray>
['LineString']
Length: 1, dtype: str
First few rows:
       id                               name    highway                                                     geometry
0  R00000  ถนนสุขุมวิท    primary  LINESTRING (656835.685 1509487.003, 657156.989 1509488.924)
1  R00001      ถนนพระราม 4  secondary   LINESTRING (657192.69 1509489.138, 657190.305 1509887.384)
2  R00002      ถนนพระราม 9   tertiary      LINESTRING (657549.695 1509491.278, 657871 1509493.208)

Column id: dtype=str
Column name: dtype=str
Column highway: dtype=str

=== markets ===
Shape: (1000, 4)
Columns: ['id', 'name_th', 'kind', 'geometry']
CRS: EPSG:4326
Geometry type: <ArrowStringArray>
['Point']
Length: 1, dtype: str
First few rows:
       id         name_th   kind                 geometry
0  M00000  ตลาดนัดจตุจักร  fresh  POINT (100.4533 13.655)
1  M00001      ตลาดอ.ต.ก.    wet  POINT (100.4599 13.655)
2  M00002     ตลาดคลองเตย  night  POINT (100.4665 13.655)

Column id: dtype=str
Column name_th: dtype=str
Column kind: dtype=str
0:25
Interesting! Let me analyze what I see:

1. **parcels** - CRS: EPSG:24047, Polygon, 4000 features
   - The `name_th` column shows garbled text like `แปลงที่ดิน บางรัก` - this looks like Thai text that was encoded as UTF-8 but then decoded as latin1 (latin1-mojibake). The correct Thai would be something like "แปลงที่ดิน บางรัก" etc.
   
2. **roads** - CRS: EPSG:32647, LineString, 5000 features
   - `name` column shows similar garbled text: `ถนนสุขุมวิท` - also latin1-mojibake for Thai text like "ถนนสุขุมวิท"

3. **markets** - CRS: EPSG:4326, Point, 1000 features
   - `name_th` shows proper Thai characters: `ตลาดนัดจตุจักร`, `ตลาดอ.ต.ก.`, `ตลาดคลองเตย` - this is UTF-8.

So:
- parcels: encoding_detected = `latin1-mojibake`
- roads: encoding_detected = `latin1-mojibake`
- markets: encoding_detected = `utf-8`

Now for sample_x and sample_y - I need a sample coordinate from each layer as it sits on disk.

For parcels (Polygon): I can get the first geometry's centroid coordinates
For roads (LineString): I can get the first geometry's start point coordinates
For markets (Point): I can get the first geometry's coordinates

Let me get the sample coordinates properly.
0:25
Bash cd /work && python3 -c "
cd /work && python3 -c "
import geopandas as gpd
import numpy as np

layers = ['parcels', 'roads', 'markets']
for layer in layers:
    gdf = gpd.read_file('bangkok_contractor_delivery.gpkg', layer=layer)
    geom_col = gdf.geometry.name
    
    # Get first geometry
    first_geom = gdf.iloc[0][geom_col]
    
    if first_geom.geom_type == 'Point':
        sample_x, sample_y = first_geom.x, first_geom.y
    elif first_geom.geom_type == 'Polygon':
        sample_x, sample_y = first_geom.exterior.xy[0][0], first_geom.exterior.xy[1][0]
    elif first_geom.geom_type == 'LineString':
        sample_x, sample_y = first_geom.xy[0][0], first_geom.xy[1][0]
    else:
        sample_x, sample_y = None, None
    
    print(f'{layer}: geom_type={first_geom.geom_type}, sample_x={sample_x}, sample_y={sample_y}, crs={gdf.crs}')
"
parcels: geom_type=Polygon, sample_x=657421.5158199039, sample_y=1509162.3701859994, crs=EPSG:24047
roads: geom_type=LineString, sample_x=656835.6850029442, sample_y=1509487.0031804747, crs=EPSG:32647
markets: geom_type=Point, sample_x=100.4533, sample_y=13.655000000000001, crs=EPSG:4326
0:37
Now I need to check for encoding anomalies. The `name_th` column in `parcels` and `roads` looks like it could be mojibake - Thai characters displayed as garbage. Let me check more carefully.

Looking at the output:
- `parcels` name_th: `แปลงที่ดิน บางรัก` - This looks like UTF-8 bytes interpreted as Latin-1 (mojibake)
- `roads` name: `ถนนสุขุมวิท` - same pattern
- `markets` name_th: `ตลาดนัดจตุจักร` - This looks like proper Thai characters (UTF-8)

Let me verify by trying to decode the bytes. If the bytes are actually UTF-8 Thai characters interpreted as Latin-1, then encoding them back to Latin-1 bytes and decoding as UTF-8 would give proper Thai.

Let me check more carefully.
0:37
Bash cd /work && python3 -c "
cd /work && python3 -c "
import geopandas as gpd

# Check encoding - test for mojibake
layers = ['parcels', 'roads', 'markets']
for layer in layers:
    gdf = gpd.read_file('bangkok_contractor_delivery.gpkg', layer=layer)
    print(f'=== {layer} ===')
    for col in gdf.columns:
        if gdf[col].dtype == 'object' and col not in ['geometry', 'geom']:
            sample = str(gdf[col].iloc[0])
            print(f'  {col}: {repr(sample)}')
            
            # Try to detect mojibake: if it looks like Latin-1 rendering of UTF-8 bytes
            # Try encoding as latin-1 and decoding as utf-8
            try:
                # If the string appears to be latin1-mojibake:
                # The characters we see are actually latin-1 characters
                # which are the raw UTF-8 bytes being misinterpreted
                raw_bytes = sample.encode('latin-1')
                decoded = raw_bytes.decode('utf-8')
                print(f'    -> latin1->utf8 decode: {repr(decoded)}')
            except:
                print(f'    -> latin1->utf8 decode: FAILED')
            
            # Try the reverse: is it UTF-8 that can be double-decoded?
            try:
                raw_bytes = sample.encode('utf-8')
                decoded = raw_bytes.decode('latin-1')
                print(f'    -> utf8->latin1 decode: {repr(decoded)}')
            except:
                print(f'    -> utf8->latin1 decode: FAILED')
    print()
"
=== parcels ===

=== roads ===

=== markets ===
0:45
The columns show as having no values printed because they're using ArrowStringArray or something. Let me check more carefully.

Let me look at the raw bytes from the geopackage to understand the encoding.
0:45
Bash cd /work && python3 -c "
cd /work && python3 -c "
import sqlite3

conn = sqlite3.connect('bangkok_contractor_delivery.gpkg')

# Check parcels table raw data
rows = conn.execute('SELECT name_th FROM parcels LIMIT 3').fetchall()
for r in rows:
    val = r[0]
    print(repr(val))
    print('  bytes:', val.encode('latin-1', errors='replace'))
    print()

# Check roads
rows = conn.execute('SELECT name FROM roads LIMIT 3').fetchall()
for r in rows:
    val = r[0]
    print(repr(val))
    print()

# Check markets  
rows = conn.execute('SELECT name_th FROM markets LIMIT 3').fetchall()
for r in rows:
    val = r[0]
    print(repr(val))
    print()

conn.close()
"
'à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 à¸\x9aาà¸\x87รัà¸\x81'
  bytes: b'\xe0\xb9\x81\xe0\xb8\x9b\xe0\xb8\xa5\xe0\xb8\x87\xe0\xb8\x97\xe0\xb8\xb5\xe0\xb9\x88\xe0\xb8\x94\xe0\xb8\xb4\xe0\xb8\x99 \xe0\xb8\x9a\xe0\xb8\xb2\xe0\xb8\x87\xe0\xb8\xa3\xe0\xb8\xb1\xe0\xb8\x81'

'à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 à¸\x9bà¸\x97ุมวัà¸\x99'
  bytes: b'\xe0\xb9\x81\xe0\xb8\x9b\xe0\xb8\xa5\xe0\xb8\x87\xe0\xb8\x97\xe0\xb8\xb5\xe0\xb9\x88\xe0\xb8\x94\xe0\xb8\xb4\xe0\xb8\x99 \xe0\xb8\x9b\xe0\xb8\x97\xe0\xb8\xb8\xe0\xb8\xa1\xe0\xb8\xa7\xe0\xb8\xb1\xe0\xb8\x99'

'à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99 วัà¸\x92à¸\x99า'
  bytes: b'\xe0\xb9\x81\xe0\xb8\x9b\xe0\xb8\xa5\xe0\xb8\x87\xe0\xb8\x97\xe0\xb8\xb5\xe0\xb9\x88\xe0\xb8\x94\xe0\xb8\xb4\xe0\xb8\x99 \xe0\xb8\xa7\xe0\xb8\xb1\xe0\xb8\x92\xe0\xb8\x99\xe0\xb8\xb2'

'à¸\x96à¸\x99à¸\x99สุà¸\x82ุมวิà¸\x97'

'à¸\x96à¸\x99à¸\x99à¸\x9eระราม 4'

'à¸\x96à¸\x99à¸\x99à¸\x9eระราม 9'

'ตลาดนัดจตุจักร'

'ตลาดอ.ต.ก.'

'ตลาดคลองเตย'
0:56
Now I can see clearly:

1. **parcels** - `name_th` values show garbled Thai text like `à¹\x81à¸\x9bลà¸\x87à¸\x97ีà¹\x88à¸\x94ิà¸\x99` - This is typical "mojibake" where UTF-8 bytes were decoded as Latin-1. When I encode these characters back to Latin-1 bytes, I get `b'\xe0\xb9\x81\xe0\xb8\x9b...'` which are valid UTF-8 bytes for Thai characters. So this is `latin1-mojibake`.

2. **roads** - Similarly, `à¸\x96à¸\x99à¸\x99สุà¸\x82ุมวิà¸\x97` - same pattern, also `latin1-mojibake`.

3. **markets** - `ตลาดนัดจตุจักร` - This is proper Thai text (UTF-8). So this is `utf-8`.

Wait, let me double-check. The raw bytes for the parcels name_th: when I do `repr()` it shows `'à¹\x81à¸\x9bล...'`. The actual UTF-8 bytes are `b'\xe0\xb9\x81\xe0\xb8\x9b...'` which when decoded as Latin-1 would give the characters I'm seeing.

So:
- parcels: `latin1-mojibake` (the Thai text was stored as UTF-8 bytes but interpreted as Latin-1)
- roads: `latin1-mojibake` (same)
- markets: `utf-8` (proper Thai text visible)

Let me verify the markets by looking at the raw bytes.
0:56
Bash cd /work && python3 -c "
cd /work && python3 -c "
import sqlite3

conn = sqlite3.connect('bangkok_contractor_delivery.gpkg')

# Check markets raw
rows = conn.execute('SELECT name_th FROM markets LIMIT 3').fetchall()
for r in rows:
    val = r[0]
    print(repr(val))
    print('  bytes:', val.encode('utf-8'))
    print()

conn.close()
"
'ตลาดนัดจตุจักร'
  bytes: b'\xe0\xb8\x95\xe0\xb8\xa5\xe0\xb8\xb2\xe0\xb8\x94\xe0\xb8\x99\xe0\xb8\xb1\xe0\xb8\x94\xe0\xb8\x88\xe0\xb8\x95\xe0\xb8\xb8\xe0\xb8\x88\xe0\xb8\xb1\xe0\xb8\x81\xe0\xb8\xa3'

'ตลาดอ.ต.ก.'
  bytes: b'\xe0\xb8\x95\xe0\xb8\xa5\xe0\xb8\xb2\xe0\xb8\x94\xe0\xb8\xad.\xe0\xb8\x95.\xe0\xb8\x81.'

'ตลาดคลองเตย'
  bytes: b'\xe0\xb8\x95\xe0\xb8\xa5\xe0\xb8\xb2\xe0\xb8\x94\xe0\xb8\x84\xe0\xb8\xa5\xe0\xb8\xad\xe0\xb8\x87\xe0\xb9\x80\xe0\xb8\x95\xe0\xb8\xa2'
1:05
Let me analyze the encoding situation:

1. **parcels** - The text looks like mojibake. When I look at the raw bytes, the Python string displays `à¹\x81à¸\x9bล...` which are the Latin-1 rendering of UTF-8 bytes. Let me check: if I encode the string as latin-1, I get `b'\xe0\xb9\x81\xe0\xb8\x9b...'` which are valid UTF-8 bytes. Decoding those as UTF-8 should give proper Thai text.

Let me verify: `à¹\x81à¸\x9b` - the Python repr shows `à¹\x81` which is `à` (0xE0) followed by `¹` (0xB9) followed by `\x81` (0x81) - that's `E0 B9 81` which is the UTF-8 encoding of a Thai character.

So parcels data is latin1-mojibake (UTF-8 bytes misinterpreted as Latin-1).

2. **roads** - Same pattern with `à¸\x96à¸\x99à¸\x99ส...` - looks like mojibake.

3. **markets** - The text is proper Thai characters (`ตลาดนัดจตุจักร`) - this is correct UTF-8.

So:
- parcels: `latin1-mojibake`
- roads: `latin1-mojibake`
- markets: `utf-8`

Let me also verify the feature counts more precisely.
1:05
Bash cd /work && python3 -c "
cd /work && python3 -c "
import sqlite3

conn = sqlite3.connect('bangkok_contractor_delivery.gpkg')

# Verify feature counts
for t in ['parcels', 'roads', 'markets']:
    cnt = conn.execute(f'SELECT COUNT(*) FROM {t}').fetchone()[0]
    print(f'{t}: {cnt} features')

conn.close()
"
parcels: 4000 features
roads: 5000 features
markets: 1000 features
1:15
Now I have all the information I need. Let me analyze the encoding situation:

1. **markets** - The Thai text displays correctly (e.g., "ตลาดนัดจตุจักร"). This is valid UTF-8 Thai characters. So encoding is `utf-8`.

2. **parcels** - The text looks like "แปลงที่ดิน บางรัก" which are Latin-1 characters that are actually UTF-8 byte sequences being interpreted as Latin-1. When I look at the bytes: `b'\xe0\xb9\x81\xe0\xb8\x9b...'` - these are valid UTF-8 bytes for Thai characters. The string shown as Python repr is `'à¹\x81à¸\x9b...'` which is the Latin-1 interpretation. So this is classic "mojibake" - the data was written as UTF-8 bytes but read as Latin-1. So encoding is `latin1-mojibake`.

3. **roads** - Same pattern as parcels. The text "ถนนสุขุมวิท" is Latin-1 mojibake. When we try to decode the Latin-1 bytes of that string as UTF-8, we'd get proper Thai. So encoding is `latin1-mojibake`.

Wait, actually I need to think more carefully about what "encoding_detected" means.

The task says: "For encoding_detected, look at the text attributes for anomalies and report either `latin1-mojibake` or `utf-8`."

So if the text attributes contain Latin-1 mojibake (i.e., they're actually UTF-8 bytes misinterpreted as Latin-1), report `latin1-mojibake`. If they're proper UTF-8, report `utf-8`.

- **parcels** (name_th): contains mojibake → `latin1-mojibake`
- **roads** (name): contains mojibake → `latin1-mojibake`  
- **markets** (name_th): contains proper Thai UTF-8 → `utf-8`

Let me also double check by trying to decode the parcels text properly.

For parcels: the raw string is `'à¹\x81à¸\x9bลà¸\x87...'` - if I encode this as latin-1, I get the raw bytes `\xe0\xb9\x81\xe0\xb8\x9b...` which is valid UTF-8. Decoding that as UTF-8 gives proper Thai.

For markets: the raw string is `'ตลาดนัดจตุจักร'` which is valid UTF-8 Thai text. No mojibake.

So the CSV should have:
- parcels: EPSG:24047, POLYGON, 4000, sample_x=657421.5158199039, sample_y=1509162.370185999, latin1-mojibake
- roads: EPSG:32647, LINESTRING, 5000, sample_x=656835.6850029442, sample_y=1509487.0031804747, latin1-mojibake
- markets: EPSG:4326, POINT, 1000, sample_x=100.4533, sample_y=13.655000000000001, utf-8

Let me format the CRS as EPSG:NNNN properly.

For parcels: srs_id=24047 → EPSG:24047
For roads: srs_id=32647 → EPSG:32647
For markets: srs_id=4326 → EPSG:4326

Let me create the CSV now.
1:15
1:15
Write /work/crs_audit.csv
content · 4 lines · csv
layer_name,declared_crs,geometry_type,feature_count,sample_x,sample_y,encoding_detected
parcels,EPSG:24047,POLYGON,4000,657421.5158199039,1509162.370185999,latin1-mojibake
roads,EPSG:32647,LINESTRING,5000,656835.6850029442,1509487.0031804747,latin1-mojibake
markets,EPSG:4326,POINT,1000,100.4533,13.655000000000001,utf-8
Wrote /work/crs_audit.csv (320 chars).
1:21
Let me verify the CSV file looks correct.
1:21
Read /work/crs_audit.csv
     1	layer_name,declared_crs,geometry_type,feature_count,sample_x,sample_y,encoding_detected
     2	parcels,EPSG:24047,POLYGON,4000,657421.5158199039,1509162.370185999,latin1-mojibake
     3	roads,EPSG:32647,LINESTRING,5000,656835.6850029442,1509487.0031804747,latin1-mojibake
     4	markets,EPSG:4326,POINT,1000,100.4533,13.655000000000001,utf-8
1:30
The CSV looks correct. Let me verify my reasoning:

1. **parcels**: 
   - CRS: EPSG:24047 (Indian 1975 / UTM zone 47N) — from gpkg_contents srs_id=24047
   - Geometry: POLYGON
   - Feature count: 4000
   - Sample: first vertex of first polygon's exterior ring
   - Encoding: The text `แปลงที่ดิน` is clearly mojibake — Latin-1 rendering of UTF-8 encoded Thai text. The raw bytes `\xe0\xb9\x81\xe0\xb8\x9b...` are valid UTF-8 for Thai characters like แปลง, so when read as Latin-1, they produce the garbage we see.

2. **roads**:
   - CRS: EPSG:32647 (WGS 84 / UTM zone 47N) — from gpkg_contents srs_id=32647
   - Geometry: LINESTRING
   - Feature count: 5000
   - Sample: first coordinate of first linestring
   - Encoding: Same mojibake pattern (e.g., `ถนนสุขุมวิท` which is Latin-1 rendering of UTF-8 Thai)

3. **markets**:
   - CRS: EPSG:4326 (WGS 84)
   - Geometry: POINT
   - Feature count: 1000
   - Sample: (100.4533, 13.655)
   - Encoding: The text `ตลาดนัดจตุจักร` is valid Thai UTF-8, so encoding is `utf-8`

This looks complete and correct.
1:30