- add a reusable registry-source workflow so every source runs in parallel rather than as
a hand-written job; adding a registry becomes one matrix entry plus one script
- add build_registry.py to align the per-source CSVs on the union of columns and emit a
single table discriminated by the source column
- reindex each frame to the union before concatenating, so a source missing a column
yields an empty cell rather than a shifted row
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- registrant_* rather than owner_*, status rather than registration_status, so both
registries describe the same concept with the same column name
- registrant_zip_code carries the Canadian postal code: a union table needs one column per
concept, not one per country's vocabulary
- set source="TC", matching the discriminator the FAA frame already carries
- raises the column names shared with the FAA frame from 9 to 21
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- publish street, city, postal code and care-of, which the FAA asset already carries as
registrant_* for 99.7% of US registrants; dropping them here left one repository with
two different postures on the same class of data
- take the address from the single MAIL_RECIPIENT row rather than merging across parties,
since a co-owned mark lists several people in different cities
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- fall back to a single-day rebuild only on FileNotFoundError; a rate limit, schema change
or truncated download previously took the same path and republished one day as the whole
dataset, which the next run then read back as its base
- move the monotonic download_date assert out of the try so corruption cannot select the
destructive branch
- walk back through releases like the ADS-B reader, so one missing optional asset does not
strand the accumulation
- authenticate release reads and verify downloaded asset size
- read the previous CSV with keep_default_na=False so literal NA values round-trip
- write the download atomically and reject non-zip responses
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- reject Mode S fields that are not 24 binary digits; a short field zero-padded into a
plausible address belonging to a different aircraft
- check the "N rows selected." footer against the parsed row count and enforce a row floor,
so an upstream short export fails instead of publishing as a smaller register
- prefer ACTIVE_FLAG "A" parties but fall back to all: 1,932 Registered marks carry only "I"
rows, and those are the MAIL_RECIPIENT, so filtering on "A" alone drops real owners
- count distinct owner names rather than rows, fixing 154 marks labelled Co-owner in error
- match the spool footer by pattern instead of a brittle ragged-row count
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- parse the headerless latin1 CCARCS export against its declared column layout
- derive transponder_code_hex from the 24-bit Mode S binary, populated for all 34,913 rows
- expand marks to C- and vintage CF- registrations and drop owner mailing addresses
- mirror the FAA build: same concat-with-latest-release dedup and output conventions
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- part id is 0-indexed in both --help and the loader docstring, matching the matrix
- daily release cron comment said 6:00pm while the expression fires at 06:00 UTC
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- module called an undefined master_txt_to_releasable_csv and could never run
- nothing in the repo imported or referenced it
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- query the repo for its default branch instead of assuming main
- use that branch as both the fork point and the pull request base
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- write through get_schema_path() instead of a literal v1 filename
- drop the unreferenced SCHEMA_PATH backwards-compatibility shim
- widen the community PR trigger to schemas/** so a version bump still fires it
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- fail when no parquet parts exist and --concat_with_latest_csv is not set
- keep the fallback path that re-releases the latest CSV when adsb.lol is late
- correct a log line that reported "no parquet files" on the populated branch
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- replace two independent copies with FINAL_COLUMN_ORDER
- polars concatenates by position, so a forked copy corrupted releases silently
Generated-by: Claude Opus 5 <noreply@anthropic.com>
- stop retrying a nonexistent repo 10 times at 5-minute intervals
- cut the dead-end Dec-31 next-year probe from ~45 minutes to under a second
- leave 403, 429, 5xx and network errors on the existing retry path
Generated-by: Claude Opus 5 <noreply@anthropic.com>
add clickhouse_connect
use 32GB
update to no longer do df.copy()
Add planequery_adsb_read.ipynb
INCREASE: update Fargate task definition to 16 vCPU and 64 GB memory for improved performance on large datasets
update notebook
remove print(df)
Ensure empty strings are preserved in DataFrame columns
check if day has data for adsb
update notebook