dd-l2-tokyo-overture-schools
Map (reference ← swipe → agent)
0:00
Pulling together a summer briefing on school-density disparities for children aged 8–14 across the 23 special wards. Out of `tokyo_places` I need every school relevant to that age range whose point sits inside the wards rectangle in `tokyo_23wards_bbox`. A colleague will visualise the result in R, so it has to land as `tokyo_schools.geojson`.
Each feature's properties must use these exact keys: `id`, `name`, `confidence`, `address_freeform`, `address_locality`, and `address_postcode`.
0:00
System info
{
"adapter": "openrouter",
"model": "google/gemma-4-26b-a4b-it",
"base_url": "https://openrouter.ai/api/v1"
}
0:01
Bash
ls -F
ls -F
tokyo_23wards_bbox.geojson tokyo_places.parquet
0:02
Bash
head -n 20 tokyo_places.parquet
head -n 20 tokyo_places.parquet
PAR1��A��9L�� �� �$ 0036772c-0c6a-4a84-a168-95f8f2097d0b$(�5b475a-c465-4fdc-b2b7-863f5b63cf5
(�e41387-1ceb-4dee-80a4-29705b5a09fd$P�1067b37-4a19-4791-90c0-6c2e99a8867a��1139e42-34cf-4306-a3e8-9f3b58bc769d
(�6b489-7500-4a2e-add9-2e2dec7416a5 (�2e8bea-a93a-4bbd-b837-c764035a756f$��401c46-7de8-48c8-b1c6-737cd4e3e7f(0f234d-3bd8-49!<5c5-ddaa294accfa x49387ab-9477-4f!h<494-7128d6b02da4
(|512a0-36a0-445a-b013-90ee8fa37c6
Ȁd74607-fc7b-4bdf-870c-ac1a41cb5fd-h@dbc010-9273-45c6-!�0-e0af9f8da717 x e60e7a-e7!�La36-b744-75572e9646b x�213296f-efba-49fb-819e-c312f3594a88P$2233dbd-04!L19b-82a3-16fc91fd62c)�230a4f7-bd01-4dd2-9be5-fc89171272c8 P�30a77a-32f4-417e-beef-b40ad42868fe$!�P23b4c70-6734-4230-833%"9711cd6c)h24f6!2a30!�De-85fc-1c0bb533889
(�60f788-6f93-4e74-bb27-99405728c499$x(84f2bd-6d51xD1-a65d-d831590e711-@097277b-bb75-4A�<b483-07c8fe2e6f8
Ȁ98c979-8955-4977-87f7-55cecd2191dP,9ad17-dc4b-4-b9_(70b8db2aac3)@$b96264-767A�H08-ae01-10c3bfe2860( cc6ef-89c(f8-b21dA 4ca78d152 P�ccacdc-2f5c-4763-8fbb-a7e56f73d14IX$2eb2153-e7A�f8e-aap(8eed7c04e12I�2eb�#4-3fe5-48f0-bd6�!dee5a81 x�f06ad5-eed6-4fe0-81ca-91f32b52fc79$!�310d4!c\ce2-4ce5-993e-6a8f7a36b1M��323c2d7-cae1-440e-96ab-e161d14d5045x329a�(921c-4e3d-a!'(46f9b12f413 ��32e69b2-d5b0-4b9c-9da6-8cfd5fad00ef P�32acc9-e4b2-435d-9978-69065605b836$�36fc5A{ d�xH06-964d-e8c9458e456)��33782f2-c4ab-4365-b896-ac8d7d7cda97
xd1!�d22bb-4217-af29-d524510fad3IX�33ed3fe-752c-42f5-af99-ca8217d13b�8<35097a2-796e-44cap858-01f7bab72d60 x�5c2ae8-0d3d-4446-8d59-8805196ff44��371cce8-acf8-464b-a554-9282b2af2210 P,7e149f-74ac-��D-ad96-14683d06125f (9d414e-b!�P4c42-a1a9-a8db965351bI�3d86dŐ da� 7�[8fd-41adfca416f8 P dba6de-a6!�L23f-85df-011d8eaca81i 3e6ebf��!@�h8a37-1e7d4baf3c9-�(ebb2e0-e152�PD3-b291-c68db64e8b6
��f66cb1-b090-46d3-aede-467c6a10f9aI�$4135b74-3e�X662a4d-0823e4c177acȈ43041e6-f9b4-4e94-b75f-ecf90c12bee8
( 36280-0ceA�Lbd-a823-0a982acadad4 (\4a6cf4-5f8b-4533-bca3-d6��d1f29)@�45f05bb-93ed-45f6-a8e9-185451a5d76b P9d3a8d <fa5-91c2-ece9a8c�N)�49e52ad-f222-4f46-9b80-65615c564290 P4a99a3c-2ab8-45Hf5d�w03e25260
Ȁcb2bd2-6196-4577-a25d-c962e6479e7)�4cf8f5�&0a��D2-ba72-0a6229eae2dI�44dbc83d-c0e9-4a<b1e5-8ac13b99ce7)� 5011168-3��4a��817c-5e43af01d76 �$5088ce9-5cA0L5b7-9745-6fd642686ab ��5261b57-5ef2-4a4f-ba44-34167a9442m�,52b3fac-1a5b� Dd-b456-850c67f910d
P|414ced-27c1-4ca1-99f6-3104286dcdm ,551bf3c-703f��D8-91a6-5bd8204ac45�(45675920-ea41-4a*9b�� 9d7d2bbcd��56adf6e-6833-4bcd-b760-2a7cef85fea2%�(58f60b1-5c6�D36-8b06-de47e18907�@�5ad0db9-8086-43f3-93c6-6d115bb676M�,5b1d280-23eeA�2-9�_$6d6f0b0299��<5b532aa-63c8-403H403-e91d046046d-�d54�e0a��H80-b70f-7a4699db0f6�85dceA+(da16-4828-8'(e136fc7c4e0I�$5e6541b-f5��Lddb-b275-bf5a1be7eb0�x,5fbab39-81d2
@9-a6b0-52766bd4fa͐ 62367eb-6>477a�4a5-2ebdf499b98I�@65557a4-c643-4861Z
,5-c8f0f4a8ac���6876095-809f-4d13-8458-8e49d6c50f83%� 6957dde-3� Pe30-a3cd-c5e26c2d3c50
(xc995b-d799-4873-965c-637202082e��6e7b3#
cP
e�2159-c0x 12d12��$6e94ee9-04A!@$8065-7e503� b�(6e98cdb-96b�Hf1-beaa-50d6e6ec0ce�\6ed9d11-1198-4f7b-9d99-da�0e406m�(6fda49c-6b3��62-b7e� 795462984 �<7013045-8d27-4ee�40f-d3664d66592i�713aeS`0fa-4ffb-9819-51f158ddd60 �$71df9c6-87A� d��8e66-92eade6d9b7i�$7498a69-f2�L7bf-9079-b91284112fe�74d06!� d(
Hd5b-8522-249e6dde0e(P5550a0-08cf-42d2-a2ff]52fd06f3)@75a896� ,71-41dc-8da9�'fe2613da��$75b16d0-9aP
d�x4496-b30e173c57 (75b5846-2c9��Hea-b399-63a1c38db47-h7ed42c-a��H7ed-a515-7f45a470cb�,787420a-2d97��Dd-9292-89bedd4267e
P 8b3246-f3!@ f� 8753-ad82370f303i�<7b03158-f344-443!�485-211b35322ed-� eb072e-1fa�H239-b37c-6ef49735de��78-e0f H3d-8ba8-e4d7cecb573I0�808a835-b6d4-4cb4-b9a7-ba8aa9e08e66ep81Ad-a459@D2-a872-23a9392623d)�81�4-56cp
D6f-a4f3-55b4121f97m�$814aebb-09hL578-937d-950b234ec4f�(,85d596d-8c9c�xb-b348��9257bf8cX86b52�� 7�L1af-813d-f476363dbc2I� 86cc5fa-1�4bc2-am(fde56bfad54
x<75895d-62a1-40d2�
00-f8d8789d1bc
PL849bd4-a767-4892-a41�� 7209a5b8f
( 9
�XX8f-49b1-aaaf-5eb8b19e1bM0$8ab07e1-c7�06ap405-d70e5ff9746I,8be3467-0015! c�1-8de7!Sdff- c1e396-ea�b5a-88�224�121(
(8de3215-c35�8!�8e52-18088b4a12e�h(8fb5d1f-b98� cJ8bb-60ba68cacfe3E0�907ccd8-e96f-4161-951d-bbc920b40c2b
(81cd7��X5-43e6-bb5e-12abf5e300b �90e�
-34H
Da2-ab8c-4a6048a0e2 $90f1eec-fe�c7f�
02-13a654f9af5)@91b2x
0556x
49a8-b54d894b5b� 9408f� cA� 5x
8b83-fdac1208d85
�5ccfb��x�#291-�\5cfe7c�� 98285b6-8��P4bc8-8465-d5a7b510c65I�$98e4f7c-a78 df1A�f-1AQcc94c� $9d09770-70��6d1-9e�$77e311b819
Pe5dddc-'4-4317-b944-b5bY688�� 9W c�
ADb93-a3b1-fd00851fb (9ed41ab-890!@$85-b9c5-c5�4faf10%� aAZ c�Fea b�$-14d4a3ce5�a269d8�e3�� e��0-d2b826� 6@a2e638� 8�H4a-a8b4-91b89ee08cdI�a73949d29-465f� 3Z8453007��@adb1002-9df9-44e50a-ead4612426f
P(e2f73b-f3ae��X0af-f023e13f030b1b�-5�421a�04b-588cc43f6c��$b3d7098-c9�@Hf61-a71f-7b17afe535�
$b568be6-5c�Heef-9c74-60ea775a98-� b6de450-0S4fK46a7-79ad873ec1�x$b72bd3f-b5��
8994-e98fde410c8��@b7701b9-ac2a-4520��(8-7ed69ce1bQX�b88c0d6-6145-4b62-a6f1-4818953716�
(bbe5389-94ca �� b6cfb82e5�Hbfc117b-3ae4-4641-8
(de656d68e2di�c019� d��L4b50-8a8e-182ef02bcf�W
$19-3712-44���@(8000ed66b95��c47a22���(L7a9-b7c5-df9ec82f953iH�c54d604-6985-4d09-abac-67c1a27ef685E�(c5c3602-ef8!@48�04-52dcec9cdc2I�c67e98cT461a-8af0-a6a045c8fff5 P6a66�79xD73-b756-76bcd41c40�$c74bb5c-c4x
�m<b798-ec71f6cf930)��cc549df-281c-40cb-b9fe-95bcb45e9ccc xd56�630p 7X,b-225fcdfb0e�(ce8dcc2-6f1
A:80ab-354a7c0e1c3�,ced6c17-cbd98@1-a2b0-1c3f8d91a4P�@8b-13c9-4cd4-9cd5�bd6e5f88 Pd088K
1cf6-4!r890c2-32c877641a8d2e3�,ff05-429b-a2j$c4bddfd2f6hd3fce2c��h a
,3-0afdbf3157 $db997f0-fe� Lcbc-907a-20f02de24cf�$dd9154f-1c�
f�6a4-fd!�d0da�de��9-}
464f-a�(0e85a85f2ea � e5ec4da-7= P45f8-a86f-c4ba8653cbf� e91fA�5a�818 (0-04939cbd5He9e87&k4429-8��$c0ea4562d58@ecf185a-d2fb-40f5��,4-220f2e70f4 ed59508-8�46ad-bc�$6e124467a4�ef721a�d0P 0�fa�
5413711
Pf033!�
�Df-8ada-c71a887c754)�$f14fcae-3e($de0-a707-ca`c8a475I�(f265409-2e8�64-9d24�c342c6b� f� 5��
L60a-8fdf-9d0dd39ac05
x5e8� 8�2e2-p0-0aaf9bfd5be1eH$f8095dc-73(Hc92-99ed-fb88030bf5f850923��b-486C089-353a57e635 8f9f1071-4fbd-4fhc0� a07194021
�`a2f8a3-bbad-4372-a7b7-b5f�e6d0�� fA474-eb46-418c-a
0:04
Write
/work/check_bbox.py
content · 7 lines · py
import geopandas as gpd
# Load the bbox
bbox = gpd.read_file('tokyo_23wards_bbox.geojson')
print(bbox.head())
print(bbox.columns)
Wrote /work/check_bbox.py (131 chars).
0:05
Bash
python check_bbox.py
python check_bbox.py
name geometry 0 Tokyo 23 Special Wards (bbox) POLYGON ((139.56 35.52, 139.91 35.52, 139.91 3... Index(['name', 'geometry'], dtype='str')
0:09
Write
/work/check_places.py
content · 9 lines · py
import pandas as pd
import geopandas as gpd
# Load the places
places = pd.read_parquet('tokyo_places.parquet')
print(places.head())
print(places.columns)
print(places.dtypes)
Wrote /work/check_places.py (176 chars).
0:10
Bash
python check_places.py
python check_places.py
id ... bucket
0 0036772c-0c6a-4a84-a168-95f8f2097d0b ... 0
1 005b475a-c465-4fdc-b2b7-863f5b63cf5b ... 0
2 00e41387-1ceb-4dee-80a4-29705b5a09fd ... 0
3 01067b37-4a19-4791-90c0-6c2e99a8867a ... 0
4 01139e42-34cf-4306-a3e8-9f3b58bc769d ... 0
[5 rows x 8 columns]
Index(['id', 'geometry', 'bbox', 'confidence', 'categories', 'names',
'addresses', 'bucket'],
dtype='str')
id str
geometry object
bbox object
confidence float64
categories object
names object
addresses object
bucket int32
dtype: object
0:20
Write
/work/check_data.py
content · 28 lines · py
import pandas as pd
import geopandas as gpd
from shapely.geometry import Point
# Load the bbox
bbox = gpd.read_file('tokyo_23wards_bbox.geojson')
bbox_geom = bbox.geometry.iloc[0]
# Load the places
places = pd.read_parquet('tokyo_places.parquet')
# Convert places to GeoDataFrame
# The 'geometry' column contains objects (likely shapely geometries or dicts)
# Let's check if they are already shapely geometries or need conversion
gdf_places = gpd.GeoDataFrame(places, geometry='geometry', crs="EPSG:4326")
print("First geometry type:", type(gdf_places.geometry.iloc[0]))
# Filter for schools (we'll define what 'school' means from categories)
# Need to see what's in 'categories'
print("Example categories:", gdf_places['categories'].iloc[0])
# For now, just filter by bbox
mask = gdf_places.within(bbox_geom)
gdf_filtered = gdf_places[mask].copy()
print("Filtered count:", len(gdf_filtered))
Wrote /work/check_data.py (900 chars).
0:22
Bash
python check_data.py
python check_data.py
Traceback (most recent call last):
File "/work/check_data.py", line 15, in <module>
gdf_places = gpd.GeoDataFrame(places, geometry='geometry', crs="EPSG:4326")
File "/usr/local/lib/python3.14/site-packages/geopandas/geodataframe.py", line 243, in __init__
self.set_geometry(geometry, inplace=True, crs=crs)
~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/geopandas/geodataframe.py", line 464, in set_geometry
level = _ensure_geometry(level, crs=crs)
File "/usr/local/lib/python3.14/site-packages/geopandas/geodataframe.py", line 68, in _ensure_geometry
out = from_shapely(np.asarray(data), crs=crs)
File "/usr/local/lib/python3.14/site-packages/geopandas/array.py", line 187, in from_shapely
raise TypeError(f"Input must be valid geometry objects: {geom}")
TypeError: Input must be valid geometry objects: b'\x01\x01\x00\x00\x00;\xdfO\x8d\x97wa@\xa8W\xca2\xc4\xd1A@'
0:29
Write
/work/inspect_content.py
content · 15 lines · py
import pandas as pd
import geopandas as gpd
from shapely.geometry import shape
# Load the places
places = pd.read_parquet('tokyo_places.parquet')
# Let's look at the structure of 'addresses' and 'names'
print("Addresses example:", places['addresses'].iloc[0])
print("Names example:", places['names'].iloc[0])
print("Categories example:", places['categories'].iloc[0])
# Let's check the geometry column more carefully
print("Geometry type of first row:", type(places.iloc[0]['geometry']))
Wrote /work/inspect_content.py (491 chars).
0:30
Bash
python inspect_content.py
python inspect_content.py
Addresses example: [{'country': 'JP', 'freeform': '高輪1-16-25', 'locality': '港区', 'postcode': '108-0074', 'region': None}]
Names example: {'common': None, 'primary': 'ちいさなピカソ(親と子の絵画・造形)', 'rules': None}
Categories example: {'alternate': array(['education', 'topic_concert_venue'], dtype=object), 'primary': 'school'}
Geometry type of first row: <class 'bytes'>
1:09
Write
/work/solve.py
content · 88 lines · py
import pandas as pd
import geopandas as gpd
from shapely import wkb
import json
# Load the places
places = pd.read_parquet('tokyo_places.parquet')
# The geometry is stored as bytes (WKB)
# Let's try to load it using wkb.loads
def load_wkb(wkb_bytes):
try:
return wkb.loads(wkb_bytes)
except:
return None
places['geometry'] = places['geometry'].apply(load_wkb)
# Drop rows where geometry couldn't be loaded
places = places.dropna(subset=['geometry'])
# Convert to GeoDataFrame
gdf_places = gpd.GeoDataFrame(places, geometry='geometry', crs="EPSG:4326")
# Load the bbox
bbox = gpd.read_file('tokyo_23wards_bbox.geojson')
bbox_geom = bbox.geometry.iloc[0]
# Filter for schools
# categories is a dict: {'alternate': ..., 'primary': 'school'}
# We need to check if 'school' is in categories.primary or categories.alternate
def is_school(cat_dict):
if not isinstance(cat_dict, dict):
return False
primary = cat_dict.get('primary', '')
if primary == 'school':
return True
alternate = cat_dict.get('alternate', [])
if 'school' in alternate:
return True
return False
gdf_places['is_school'] = gdf_places['categories'].apply(is_school)
gdf_schools = gdf_places[gdf_places['is_school']].copy()
# Filter by bbox
gdf_schools_in_bbox = gdf_schools[gdf_schools.within(bbox_geom)].copy()
# Now we need to handle the properties:
# id, name, confidence, address_freeform, address_locality, address_postcode
# names is a dict: {'common': ..., 'primary': ..., 'rules': ...}
# addresses is a list of dicts: [{'country': 'JP', 'freeform': '...', 'locality': '...', 'postcode': '...', 'region': ...}]
def extract_properties(row):
name_dict = row['names']
addr_list = row['addresses']
name = None
if isinstance(name_dict, dict):
name = name_dict.get('primary') or name_dict.get('common')
addr_freeform = None
addr_locality = None
addr_postcode = None
if isinstance(addr_list, list) and len(addr_list) > 0:
addr = addr_list[0]
addr_freeform = addr.get('freeform')
addr_locality = addr.get('locality')
addr_postcode = addr.get('postcode')
return pd.Series([name, addr_freeform, addr_locality, addr_postcode])
gdf_schools_in_bbox[['name_ext', 'addr_freeform', 'addr_locality', 'addr_postcode']] = gdf_schools_in_bbox.apply(extract_properties, axis=1)
# Final selection and renaming
# We need: id, name, confidence, address_freeform, address_locality, address_postcode
# Note: name_ext is our extracted name.
final_gdf = gdf_schools_in_bbox[['id', 'name_ext', 'confidence', 'addr_freeform', 'addr_locality', 'addr_postcode']].copy()
final_gdf.columns = ['id', 'name', 'confidence', 'address_freeform', 'address_locality', 'address_postcode']
# Output to geojson
final_gdf.to_file('tokyo_schools.geojson', driver='GeoJSON')
print(f"Saved {len(final_gdf)} schools to tokyo_schools.geojson")
print(final_gdf.head())
Wrote /work/solve.py (2973 chars).
1:11
Bash
python solve.py
python solve.py
Traceback (most recent call last):
File "/work/solve.py", line 43, in <module>
gdf_places['is_school'] = gdf_places['categories'].apply(is_school)
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/pandas/core/series.py", line 5084, in apply
).apply()
~~~~~^^
File "/usr/local/lib/python3.14/site-packages/pandas/core/apply.py", line 1520, in apply
return self.apply_standard()
~~~~~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/pandas/core/apply.py", line 1578, in apply_standard
mapped = obj._map_values(mapper=curried)
File "/usr/local/lib/python3.14/site-packages/pandas/core/base.py", line 1022, in _map_values
return algorithms.map_array(arr, mapper, na_action=na_action)
~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/pandas/core/algorithms.py", line 1715, in map_array
return lib.map_infer(values, mapper)
~~~~~~~~~~~~~^^^^^^^^^^^^^^^^
File "pandas/_libs/lib.pyx", line 3071, in pandas._libs.lib.map_infer
File "/work/solve.py", line 39, in is_school
if 'school' in alternate:
^^^^^^^^^^^^^^^^^^^^^
TypeError: argument of type 'NoneType' is not a container or iterable
1:16
Edit
/work/solve.py
Edited /work/solve.py.
1:17
Bash
python solve.py
python solve.py
Traceback (most recent call last):
File "/work/solve.py", line 84, in <module>
final_gdf.to_file('tokyo_schools.geojson', driver='GeoJSON')
^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/pandas/core/generic.py", line 6206, in __getattr__
return object.__getattribute__(self, name)
~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^
AttributeError: 'DataFrame' object has no attribute 'to_file'. Did you mean: 'to_pickle'?
1:21
Edit
/work/solve.py
old_string and new_string are identical; no edit to make
1:25
Edit
/work/solve.py
Edited /work/solve.py.
1:55
Write
/work/solve.py
content · 87 lines · py
import pandas as pd
import geopandas as gpd
from shapely import wkb
import numpy as np
# Load the places
places = pd.read_parquet('tokyo_places.parquet')
# The geometry is stored as bytes (WKB)
def load_wkb(wkb_bytes):
try:
return wkb.loads(wkb_bytes)
except:
return None
places['geometry'] = places['geometry'].apply(load_wkb)
places = places.dropna(subset=['geometry'])
# Convert to GeoDataFrame
gdf_places = gpd.GeoDataFrame(places, geometry='geometry', crs="EPSG:4326")
# Load the bbox
bbox = gpd.read_file('tokyo_23wards_bbox.geojson')
bbox_geom = bbox.geometry.iloc[0]
# Filter for schools
def is_school(cat_dict):
if not isinstance(cat_dict, dict):
return False
primary = cat_dict.get('primary', '')
if primary == 'school':
return True
alternate = cat_dict.get('alternate', [])
if isinstance(alternate, (list, np.ndarray)) and 'school' in alternate:
return True
return False
gdf_places['is_school'] = gdf_places['categories'].apply(is_school)
gdf_schools = gdf_places[gdf_places['is_school']].copy()
# Filter by bbox
gdf_schools_in_bbox = gdf_schools[gdf_schools.within(bbox_geom)].copy()
# Function to extract properties
def extract_properties(row):
name_dict = row['names']
addr_list = row['addresses']
name = None
if isinstance(name_dict, dict):
name = name_dict.get('primary') or name_dict.get('common')
addr_freeform = None
addr_locality = None
addr_postcode = None
if isinstance(addr_list, list) and len(addr_list) > 0:
addr = addr_list[0]
if isinstance(addr, dict):
addr_freeform = addr.get('freeform')
addr_locality = addr.get('locality')
addr_postcode = addr.get('postcode')
return pd.Series([name, addr_freeform, addr_locality, addr_postcode])
# Applying extraction
extracted = gdf_schools_in_bbox.apply(extract_properties, axis=1)
gdf_schools_in_bbox[['name_ext', 'addr_freeform', 'addr_locality', 'addr_postcode']] = extracted
# Final selection and renaming
final_gdf = gdf_schools_in_bbox[['id', 'name_ext', 'confidence', 'addr_freeform', 'addr_locality', 'addr_postcode']].copy()
final_gdf.columns = ['id', 'name', 'confidence', 'address_freeform', 'address_locality', 'address_postcode']
# We need to preserve the geometry for the GeoJSON output
# Since final_gdf is a slice of gdf_schools_in_bbox, it should have geometry if we are careful.
# Let's re-attach geometry to be sure.
final_gdf['geometry'] = gdf_schools_in_bbox['geometry']
# Convert to GeoDataFrame to use to_file
final_gdf_gdf = gpd.GeoDataFrame(final_gdf, geometry='geometry', crs="EPSG:4326")
# Output to geojson
final_gdf_gdf.to_file('tokyo_schools.geojson', driver='GeoJSON')
print(f"Saved {len(final_gdf_gdf)} schools to tokyo_schools.geojson")
print(final_gdf_gdf.head())
Wrote /work/solve.py (2872 chars).
1:57
Bash
python solve.py
python solve.py
Saved 1542 schools to tokyo_schools.geojson
id ... geometry
0 0036772c-0c6a-4a84-a168-95f8f2097d0b ... POINT (139.73725 35.6388)
8 014f234d-3bd8-49dc-b5c5-ddaa294accfa ... POINT (139.76743 35.67563)
19 024f6273-2a30-4f8e-85fc-1c0bb5338894 ... POINT (139.73486 35.56782)
28 02eb2153-e773-4f8e-a837-8eed7c04e12d ... POINT (139.78017 35.69899)
31 0310d475-4ce2-4ce5-993e-6a8f7a36b1ca ... POINT (139.77075 35.69557)
[5 rows x 7 columns]
1:59
Bash
head -n 20 tokyo_schools.geojson
head -n 20 tokyo_schools.geojson
{
"type": "FeatureCollection",
"name": "tokyo_schools",
"crs": { "type": "name", "properties": { "name": "urn:ogc:def:crs:OGC:1.3:CRS84" } },
"features": [
{ "type": "Feature", "properties": { "id": "0036772c-0c6a-4a84-a168-95f8f2097d0b", "name": "ちいさなピカソ(親と子の絵画・造形)", "confidence": 0.73368567228317261, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.73725, 35.6388 ] } },
{ "type": "Feature", "properties": { "id": "014f234d-3bd8-49dc-b5c5-ddaa294accfa", "name": "speek", "confidence": 0.61169266700744629, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.767425808810003, 35.675630313048998 ] } },
{ "type": "Feature", "properties": { "id": "024f6273-2a30-4f8e-85fc-1c0bb5338894", "name": "大田区大森町の個別指導塾 学習塾|英才個別学院 大森町校", "confidence": 0.7147371768951416, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.73486, 35.56782 ] } },
{ "type": "Feature", "properties": { "id": "02eb2153-e773-4f8e-a837-8eed7c04e12d", "name": "Anabuki College Japanese Tokushima", "confidence": 0.9157291054725647, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.7801658, 35.6989944 ] } },
{ "type": "Feature", "properties": { "id": "0310d475-4ce2-4ce5-993e-6a8f7a36b1ca", "name": "奥田 開業実践塾", "confidence": 0.53572899103164673, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.770748615260004, 35.695573779420997 ] } },
{ "type": "Feature", "properties": { "id": "0323c2d7-cae1-440e-96ab-e161d14d5045", "name": "橋口孝司・燻製料理とお酒の教室", "confidence": 0.77710115909576416, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.737121582029999, 35.649295806885 ] } },
{ "type": "Feature", "properties": { "id": "033782f2-c4ab-4365-b896-ac8d7d7cda97", "name": "港区立青山小学校", "confidence": 0.97962337732315063, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.718974839999987, 35.669981 ] } },
{ "type": "Feature", "properties": { "id": "04cf8f56-b70a-4172-ba72-0a6229eae2da", "name": "Hakubikyotokimonogakuin Ikebukuro School", "confidence": 0.15898820757865906, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.71289237, 35.73151149 ] } },
{ "type": "Feature", "properties": { "id": "04dbc83d-c0e9-4ae8-b1e5-8ac13b99ce7c", "name": "渋谷の彫金教室 studio crucible | スタジオ クルーシブル", "confidence": 0.72532671689987183, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.70508039, 35.660468 ] } },
{ "type": "Feature", "properties": { "id": "05ad0db9-8086-43f3-93c6-6d115bb67698", "name": "スクールPTI", "confidence": 0.45639809966087341, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.58135815, 35.70130062 ] } },
{ "type": "Feature", "properties": { "id": "05b1d280-23ee-45f2-92f2-6d6f0b029936", "name": "セルモ 東十条教室", "confidence": 0.48836621642112732, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.73019102, 35.76182502 ] } },
{ "type": "Feature", "properties": { "id": "077ed42c-a03f-47ed-a515-7f45a470cb17", "name": "スタジオアロイ タイ料理教室", "confidence": 0.65927082300186157, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.6953844, 35.6074778 ] } },
{ "type": "Feature", "properties": { "id": "095ccfbf-7a4b-4979-b291-c38d5cfe7c64", "name": "桜坂トーストマスターズクラブ〜話し方で四季を彩る〜", "confidence": 0.65645790100097656, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.6697436, 35.5900133 ] } },
{ "type": "Feature", "properties": { "id": "0c0197ee-d95d-4b50-8a8e-182ef02bcfc3", "name": "元気日本語文化学校東京校", "confidence": 0.6058315634727478, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.70593793, 35.69300056 ] } },
{ "type": "Feature", "properties": { "id": "0c67e981-427f-461a-8af0-a6a045c8fff5", "name": "Yoji Sansuu School Spica", "confidence": 0.95427185297012329, "address_freeform": null, "address_locality": null, "address_postcode": null }, "geometry": { "type": "Point", "coordinates": [ 139.71226324, 35.64760676 ] } },
2:28
Write
/work/solve.py
content · 85 lines · py
import pandas as pd
import geopandas as gpd
from shapely import wkb
import numpy as np
# Load the places
places = pd.read_parquet('tokyo_places.parquet')
# The geometry is stored as bytes (WKB)
def load_wkb(wkb_bytes):
try:
return wkb.loads(wkb_bytes)
except:
return None
places['geometry'] = places['geometry'].apply(load_wkb)
places = places.dropna(subset=['geometry'])
# Convert to GeoDataFrame
gdf_places = gpd.GeoDataFrame(places, geometry='geometry', crs="EPSG:4326")
# Load the bbox
bbox = gpd.read_file('tokyo_23wards_bbox.geojson')
bbox_geom = bbox.geometry.iloc[0]
# Filter for schools
def is_school(cat_dict):
if not isinstance(cat_dict, dict):
return False
primary = cat_dict.get('primary', '')
if primary == 'school':
return True
alternate = cat_dict.get('alternate', [])
if isinstance(alternate, (list, np.ndarray)) and 'school' in alternate:
return True
return False
gdf_places['is_school'] = gdf_places['categories'].apply(is_school)
gdf_schools = gdf_places[gdf_places['is_school']].copy()
# Filter by bbox
gdf_schools_in_bbox = gdf_schools[gdf_schools.within(bbox_geom)].copy()
# Function to extract properties
def extract_properties(row):
name_dict = row['names']
addr_list = row['addresses']
name = None
if isinstance(name_dict, dict):
name = name_dict.get('primary') or name_dict.get('common')
addr_freeform = None
addr_locality = None
addr_postcode = None
if isinstance(addr_list, list) and len(addr_list) > 0:
addr = addr_list[0]
if isinstance(addr, dict):
addr_freeform = addr.get('freeform')
addr_locality = addr.get('locality')
addr_postcode = addr.get('postcode')
return pd.Series([name, addr_freeform, addr_locality, addr_postcode])
# Applying extraction
extracted = gdf_schools_in_bbox.apply(extract_properties, axis=1)
gdf_schools_in_bbox[['name_ext', 'addr_freeform', 'addr_locality', 'addr_postcode']] = extracted
# Final selection and renaming
final_gdf = gdf_schools_in_bbox[['id', 'name_ext', 'confidence', 'addr_freeform', 'addr_locality', 'addr_postcode']].copy()
final_gdf.columns = ['id', 'name', 'confidence', 'address_freeform', 'address_locality', 'address_postcode']
# We need to preserve the geometry for the GeoJSON output
final_gdf['geometry'] = gdf_schools_in_bbox['geometry']
# Convert to GeoDataFrame to use to_file
final_gdf_gdf = gpd.GeoDataFrame(final_gdf, geometry='geometry', crs="EPSG:4326")
# Output to geojson
final_gdf_gdf.to_file('tokyo_schools.geojson', driver='GeoJSON')
print(f"Saved {len(final_gdf_gdf)} schools to tokyo_schools.geojson")
print(final_gdf_gdf.head())
Wrote /work/solve.py (2737 chars).
2:29
Bash
python solve.py
python solve.py
Saved 1542 schools to tokyo_schools.geojson
id ... geometry
0 0036772c-0c6a-4a84-a168-95f8f2097d0b ... POINT (139.73725 35.6388)
8 014f234d-3bd8-49dc-b5c5-ddaa294accfa ... POINT (139.76743 35.67563)
19 024f6273-2a30-4f8e-85fc-1c0bb5338894 ... POINT (139.73486 35.56782)
28 02eb2153-e773-4f8e-a837-8eed7c04e12d ... POINT (139.78017 35.69899)
31 0310d475-4ce2-4ce5-993e-6a8f7a36b1ca ... POINT (139.77075 35.69557)
[5 rows x 7 columns]
3:09