Skip to content

The first batch of connectors is loaded and queryable. Each connector is a pipeline that fetches the source release, harmonises it, and writes partitioned parquet to R2, where the sandbox reads it with DuckDB.

Live now:

  • G-NAF: quarterly release, partitioned by state, about 15.3 million addresses.
  • Cadastral parcels: parcel geometry for the states with open cadastre.
  • Geoscience Australia bushfire boundaries: national historical fire extents.
  • SILO: gridded daily climate observations from 1889, aggregated to the metrics we need.
  • DEA Coastlines: annual shoreline positions since 1988.

Notes from loading them

Format variance was a larger problem than volume. G-NAF ships as pipe-delimited text with a relational schema across a dozen tables. SILO is NetCDF. DEA Coastlines is vector. Cadastre varies by jurisdiction in format and attribute naming, and the same field has four different names across four states.

The harmonisation layer was most of the work and most of the value. A query that spans addresses, parcels and climate grids should not need to know the source format of any of them.

Next

NSW flood studies, CMIP6 downscaled projections, NGER emissions and the ACCU register. Flood will take longest, for the reasons described here.

Run it on your own portfolio

Zenancy is in private preview with Group 2 reporters and their advisers.

Request access