The problem
Most critical events in the safety domain of aviation are first introduced through radio between air traffic controllers and pilots. That information is usually scattered and heavily rely on the experiece of controllers operators. The recent runway accident at LaGuardia, an Air Canada aircraft colliding with an airport fire truck on the movement area, is a clear case: the tower was understaffed, one controller was on duty, and that controller was tracking many things at once.
This project builds an automated pipeline that converts raw ATC communication into a fact sheet: a normalized, readable record fro any machine of the operational details, location, runway, flight identifier, transponder (squawk) code, incident category, generated within seconds. When warranted, it also drafts a Notice to Air Missions (NOTAM). The goal is to cut manual effort, make the reading of ATC comms more consistent, and reduce incidents caused by miscommunication.
Data
The corpus pulls from many sources: YouTube metadata, video, audio, subtitle files, OCR output, ASR transcripts, that differ widely in structure and quality. To manage that, the platform uses a Medallion architecture: data is refined progressively across three layers rather than transformed all at once.
- Bronze — landing zone. Raw ingested data and file references (downloaded video, audio, subtitles, source metadata). Preserves the original and keeps traceability to downstream tables.
- Silver — cleaned and enriched. Transcripts parsed, duplicates removed, media converted to formats ready for analysis. The main processing environment where most transformation logic lives.
- Gold — business ready. Curated datasets consolidating OCR, ASR, subtitle, and metadata sources, the single source of truth that feeds into extraction pipeline, NOTAM generation, and analysis.
Databricks environment
Everything was built and run on Databricks, the capstone partner’s platform. The pipeline runs on serverless compute, which restricts external network access, so Python packages couldn’t be installed via pip. Dependencies were packaged as wheel files and uploaded to a Databricks Volume, with configuration managed through .yml files. Data lives in Unity Catalog under the medallion layers, each mapping to a schema of Delta tables. PySpark and SQL handle transformation. Videos were downloaded locally and uploaded to a Volume alongside intermediate assets such as extracted frames and .srt files. The final application deploys in the same environment.
Scraping and ingestion
The primary source is the YouTube channel VASAviation; roughly 2,500 videos across 8 playlists covering medical emergencies, engine failures, bird strikes, smoke and fumes, near collisions, landing-gear failures, and real ATC communications. Everything was scraped through a Python script using yt-dlp batch downloads audio and autogenerated English subtitles per playlist, saving MP3 audio and SRT subtitles named by YouTube video ID. Because of the workspace network restriction and YouTube’s cookie uth requirement and the serverless environment, downloads ran locally and were uploaded to Databricks.
Source of truth
To obtain a dataset to give reliable reference for validating extracted fields, a source of truth was extracted from Youtube metadata like titles, descriptions and some summaries that were previously processed by the channel’s author. To keep it reproducible and ‘true’, extraction only used deterministic methods like regex matching and dctionary lookups against aviation mini databases.
Outcome: In practice the dataset built from this source is not reliable enough to serve as dependable ground truth. Its results are treated as exploratory, not as a benchmark. The ideal scenario would be a bigger database with gold data from the aviation company.
Cleaning data
ETLs
A series of ETL steps moves data through the medallion layers, each doing one job:
- Raw multimedia and metadata ingested into Bronze
- Source records standardized and linked to their assets
- Videos converted into image frames
- Audio transcribed through ASR
- Subtitle files parsed and cleaned
- OCR applied to extracted video frames
- OCR, ASR, subtitle, and metadata results consolidated
- Curated datasets loaded into Gold
- Extraction pipeline consumes Gold to generate incident records

ASR - speech to text
Audio tracks are extracted from the downloaded videos, converted to a standard format, and linked to their video ID. A speech recognition model transcribes them at scale; transcripts are then cleaned, formatting issues removed, aviation terminology normalized, and stored in Silver. ASR mainly covers videos with missing or incomplete subtitles, and acts as a complementary source that can validate or supplement the other extractions.
OCR - video to image to text
A three stage OCR pipeline extracts the screen subtitle text VASAviation overlays on its videos. Raw Bronze videos are split into PNG frames with OpenCV, sampling one frame per second, each named with its index and millisecond timestamp to keep temporal context. In Silver, frames go through cleaning before OCR:
- Average pixel brightness is checked to drop pure black frames.
- The bottom 17% of each frame is analyzed via HSV color thresholding to detect green (pilot) or yellow (tower) subtitle text; frames with neither color are discarded, the rest cropped to that region.
- Identical consecutive frames are removed with difference hashing (
dhash) restricted to the center 70% width, ignoring the edge channel logo. Two frames within a hash distance of 5 count as duplicates. EasyOCRreads the text off each cleaned frame. A few regex filters clean up what comes back: they strip out watermarks, URLs, and social media handles, and drop any entry that’s just telemetry from the overlay. There’s also a check for whether an entry actually looks like words, if fewer than 55% of the tokens pass a test for vowels and consonant clusters, it gets tossed. To keep that filter from throwing away real aviation vocabulary, an allowlist protects terms like ICAO codes,squawk, andwilco.
Remaining text is written to per-video .srt files, then parsed into the aero_corpus Delta table, filling the ocr_text column and metadata.
Subtitles autogenerated by Youtube
A script parses raw SRT into clean plain text: removing sequence numbers and timestamp lines via regex, stripping tags like [Music] and [Applause] that do not describe text, and collapsing the rest into one continuous string per video. This produces one text file per video for downstream NER and LLM extraction. Cleaned transcripts and audio were uploaded to Unity Catalog Volumes and registered as Bronze Delta tables.
Context Engineering
Instead of one large prompt that extracts everything at once, extraction is a chain of four focused LLM calls, each with its own small purpose prompt, where one stage feeds the next. All calls hit a Databricks serving endpoint (databricks-gpt-5-4-mini). Splitting the work this way reduces hallucination and confusion in long contexts.
- Spotter — lists every mention of an aircraft, runway, airport, or airline. Told not to judge importance or deduplicate: collect everything.
- Classifier — normalizes each mention into a clean entity, assigns a role (focal, operator, diversion), and extracts details only when stated. Injects matched glossary terms ( around 1,000 FAA entries) via a regex matcher. Returns null rather than guessing, to avoid hallucination.
- Linker — deduplicates aircraft variants, builds relationships, and selects the single focal aircraft.
- Summarizer — a separate text prompt writes one or two sentences describing the incident readable to humans.
Extraction Engine
| File | Role |
|---|---|
00_driver.py | Driver. Controls execution flow, loads transcripts, runs each layer in sequence, applies NOTAM logic, and writes final results. |
01_layer1_regex.py | Deterministic pattern matching for clearly structured fields; flight IDs, runways, squawk codes, emergency declarations. Fast and reliable. |
02_layer2_dictionary.py | Enriches results against curated airport, airline, and flight databases to recover entities pattern matching alone misses. |
03_layer3_llm_fallback.py | Optional LLM fallback. Runs only on fields still unresolved after layers 1 and 2 to improve coverage while limiting cost. |
04_notam_logic_v2.py | NOTAM engine with confidence scoring, validation checks, reasoning output, and structured ICAO NOTAM generation. |
_shared_schema.py | Standard schema, extraction fields, confidence thresholds, and data-structure requirements shared across components. |
_shared_helpers.py | Reusable utilities: spoken-number normalization, NATO phonetic processing, callsign reconstruction. |
_shared_glossary.py | Loads aviation terminology, scans transcripts for known ATC terms, and adds definitions to the output record. |
_prompt_templates.py | Centralized LLM prompt templates, so prompts can change without touching extraction logic. |
build_aero_corpus_summary.py | Consolidates subtitle, ASR, and OCR outputs, applying resolution rules to pick the most reliable value per field. |

Application
Databricks app with two modes: a Live ATC tool for controllers and a YouTube corpus explorer for browsing incidents previously processed.
Settings
On first use, the app opens to a settings screen with three configurations, which then live in the upper-right corner and can be changed anytime:
- Mode — toggle between the Live ATC tool (parses comms in real time) and the YouTube corpus (browse processed incidents). The two settings below apply only to live operation.
- Facility airport — the controller’s own airport as a four letter ICAO code (defaults to Vancouver International, CYVR).
- Emergency contacts — a roster of local services (fire, medical, hospital, police, hazmat, security) with contact methods, added / edited / removed and stored locally.
Live tool
Comms feed on the left, extracted results on the right, matching the controller’s workflow: capture the exchange, run the pipeline, review the output.
- Input panel — a fully editable text area; the controller pastes or types an exchange. A single Parse action runs the pipeline while a processing indicator masks the right panel.
- Flight information factsheet — the focal aircraft and its extracted fields: flight details (callsign, aircraft type with “heavy” indicator, tail number, origin, destination), incident details (category, runway, squawk, decoded and validated so a misheard callsign isn’t read as a transponder code), and operational details (souls on board, fuel) when stated. Every field is editable; the runway is fact checked against the facility’s real runways with mismatches flagged. A model-generated summary accompanies the fields.
- NOTAM — where warranted, the tool flags that a NOTAM is required, assigns its type (e.g. runway closure), and autodrafts it in ICAO shorthand, marked unissued pending review. The controller can override, edit, regenerate, or save. The system proposes; the human confirms.
- Emergency response — recommends which local services to notify, surfacing relevant contacts for the incident (ambulance and hospital for a medical case, hazmat for a fuel spill). The “notify” button is currently a disconnected port that could later drive a phone call or email.
Design principle: Every field on the factsheet stays editable to “always have humans in the loop.” Fields the summary doesn’t cover appear empty rather than guessed, and an invalid squawk shows blank rather than misrepresented.
Findings
The project shows that you can take the back and forth of aviation radio and turn it into structured records that a system can actually work with, using a mix of data engineering and an LLM, while still leaving the safety calls to a human. Starting from public incident videos, the workflow pulls in the raw media and metadata, cleans up the transcripts coming from subtitles, ASR, and OCR, brings them together in a Medallion architecture, and then leans on the Gold layer to do the extraction, the review, and the decision support in the app.
The big takeaway from working with the videos is that everything downstream rises or falls with the quality of the transcript. YouTube subtitles cover a lot of ground when they’re there, and ASR and OCR fill in the videos where captions are missing or patchy. But each has its quirks: OCR gets thrown off by where the subtitle sits on screen, by graphics on top of the video, and by which frames you grab, and ASR can mishear aviation terms it isn’t used to. So pulling from several transcript sources at once holds up better than trusting any one of them. The Source of Truth dataset didn’t turn out reliable enough to lean on, but it still helps as a way to sanity check flight IDs, airports, and incident details.
Limitations: The dataset comes from public incident videos rather than a live feed, so the transcripts carry noise, gaps, and moments where it’s not clear who’s talking. The NOTAM drafts are a starting point, not a finished product, and some still have placeholders a person needs to fill in before anyone uses them. From here, the next steps are cleaning up transcript quality, testing against a bigger set of examples that experts have labeled, tightening the confidence scoring, and putting the tool in front of people who actually work in aviation.
Error analysis
With no reliable validation layer, the team implemented a lightweight internal consistency check as a substitute. Rather than comparing output against external ground truth, it checks the pipeline’s own output fields against each other within the final aero_corpus_summary.csv.
Extraction rate
| Field | Filled / total | Rate |
|---|---|---|
flight_id | 10 / 14 | 71.4% |
tail_number | 9 / 14 | 64.3% |
aircraft_type | 11 / 14 | 78.6% |
runway | 12 / 14 | 85.7% |
squawk_code | 2 / 14 | 14.3% |
incident_category | 14 / 14 | 100.0% |
origin | 7 / 14 | 50.0% |
destination | 7 / 14 | 50.0% |
| Total | 72 / 112 | 64.3% |
Acknowledgments
This project was a shared effort by Selene, Infinity, Rayan, and Vincent.
We’re grateful to our capstone partner, Jeppesen ForeFlight, for the problem, the guidance, and the Databricks environment that made the work possible, and to the MDS program and its instructors for the support and feedback along the way.
Any strengths in this project belong to the whole team.