live
./docs / data-pipeline

Data Pipeline

Jobs

All jobs live in data-engine/src/jobs/. They are invoked via src/main.py --job <name>.

Main Race Jobs

JobPurpose
sync_schedulePopulates the races table for a season from FastF1 (includes sprint dates, event_format)
sync_seasonPopulates teams and drivers for a season from FastF1
ingest_qualifyingIngests Q1/Q2/Q3 qualifying session data and sector times — 2018+
ingest_qualifying_legacyIngests qualifying from Ergast — pre-2018
ingest_fp2Ingests practice long-run stints into fp2_long_run_times — FP2 primary, FP1 fallback on sprint weekends (recorded in session_type)
ingest_raceIngests race results, lap times, and race conditions (weather, SC/VSC, temps) — 2018+
ingest_race_legacyIngests race results from Ergast (no lap data) — pre-2018
compute_season_statsAggregates driver_season_stats and team_season_stats for a year (includes sprint aggregates)
compute_featuresComputes the 12 feature scores per driver for a grand prix
compute_predictionsRuns softmax on feature scores, writes win probabilities and predicted positions

Sprint Race Jobs

JobPurpose
ingest_sprint_qualifyingIngests SQ session — stores SQ1/SQ2/SQ3 times + sector times + speed in sprint_results; uses messages=True; has date guard
ingest_sprintIngests sprint race results, sprint lap times, and sprint conditions (weather, SC/VSC, temps)
compute_sprint_featuresComputes 8 sprint-specific feature scores per driver
compute_sprint_predictionsRuns softmax on sprint feature scores → sprint win probabilities
data_quality_auditMeasures per-table completeness/coverage across a season; writes data_quality_runs + data_quality_issues
data_quality_repairRe-ingests data for open, fixable issues reported by the audit, then recomputes the affected features/predictions/season-stats

Race Events Jobs (OpenF1)

Enrich a completed race weekend with data FastF1 doesn’t provide, sourced from the free OpenF1 API instead. OpenF1 has no coverage before 2023, and its free tier only serves data once a session is >30min past its end, so these jobs self-skip until the weekend is completed and race_date_utc is safely past that window — the Saturday sprint session is always well outside its own live window by then too, so one gate covers both. Each accepts --session Race (default, the Sunday GP) or --session Sprint (the Saturday sprint race on a sprint weekend); a Sprint request self-skips as a no-op on a non-sprint weekend, since races.sprint_date is null there.

Wired into auto_runner.py’s MAIN_RACE sequence (see below) — every completed race weekend now gets all three, for both sessions where applicable, automatically. Each runs via a non-fatal wrapper (_run_openf1_step) so a transient OpenF1 hiccup logs and moves on rather than blocking or reverting the core race-completion pipeline it rides on.

JobPurpose
ingest_race_controlIngests OpenF1 flags/safety-car/VSC/incident messages into race_control_messages — 2023+ only
ingest_overtakesIngests OpenF1 on-track overtakes into race_overtakes, resolving each driver number to a drivers.id via build_driver_number_map — 2023+ only
ingest_team_radioIngests OpenF1 team radio clips into team_radio_clips, same driver-number resolution — 2023+ only; F1 doesn’t release radio for every session, so a partial or empty result is normal

Backfill across a year range with python scripts/backfill_race_control.py <start> <end>, python scripts/backfill_overtakes.py <start> <end>, or python scripts/backfill_team_radio.py <start> <end> (data-engine/, paced to stay under OpenF1’s free-tier rate limit). These backfill scripts currently only cover the Race session; pass --session Sprint to the underlying src.main job directly to backfill a specific sprint weekend’s race events by hand.


Job Chain

Conventional Weekend

sync_schedule        races table must exist before ingest can find rounds
      ↓
sync_season          teams and drivers must exist before results can reference them
      ↓
ingest_qualifying    qualifying_results must exist before compute_features
      ↓
ingest_race          race_results and lap_times ingested; race.status → completed
      ↓
compute_season_stats driver_season_stats and team_season_stats must be fresh
                     before compute_features reads them
      ↓
ingest_race_control, ingest_overtakes, ingest_team_radio (Race, then Sprint if applicable)
                     Race Events (OpenF1) — non-fatal, 2023+ only
      ↓
compute_features     produces driver_prediction_features (raw weighted scores)
      ↓
compute_predictions  reads raw scores, runs softmax, writes probabilities

Sprint Weekend

sync_schedule + sync_season
      ↓
ingest_sprint_qualifying   → sprint_results (sq1/sq2/sq3 times + grid)
                             race.status → sprint_qualifying_done
      ↓
compute_sprint_features    → driver_sprint_features
compute_sprint_predictions → sprint_predictions
      ↓
ingest_sprint              → sprint_results (finish positions, points)
                             → sprint_lap_times, races (sprint conditions)
                             race.status → sprint_done
      ↓
compute_season_stats       → updates sprint_wins, sprint_total_points, etc.
      ↓
ingest_qualifying          → qualifying_results
                             race.status → qualifying_done
      ↓
compute_features + compute_predictions  → main race prediction
      ↓
ingest_race                → race_results + lap_times
                             race.status → completed
      ↓
compute_season_stats       → final season stats update
      ↓
ingest_race_control, ingest_overtakes, ingest_team_radio (Race, then Sprint)
                            Race Events (OpenF1) — non-fatal, 2023+ only; the Sprint pass
                            covers Saturday's sprint session under the same race_id

compute_season_stats should be re-run after each race so that rolling stats (win rate, DNF rate, avg position gain, sprint aggregates) are up to date before the next prediction.

Automated Polling (Render)

Instead of relying on fixed days and times (which is fragile due to global timezones and API delays), the pipeline uses an automated polling architecture hosted on a free Render Web Service.

A persistent web service runs continuously:

Service TypeCommandPurpose
Web Servicepython -m src.serverHosts a live HTML dashboard (and health check) to monitor the engine’s status, and runs auto_runner.py every hour in the background.

How auto_runner.py works:

  1. Queries the database for the most recent active race (where status != 'completed').
  2. Fetches the official F1 schedule via FastF1 to get exact UTC session times.
  3. Checks if a session ended recently:
    • Sprint Qualifying + 1.5 hours
    • Sprint Race + 1.5 hours
    • Main Qualifying + 2 hours
    • Main Race + 3 hours
  4. If the time has passed, it attempts to download the data.
    • If F1 data is delayed, FastF1 throws a DataNotLoadedError. The script catches this, exits cleanly, and tries again next hour.
  5. Once ingestion succeeds, it automatically chains the downstream jobs (features, predictions, stats). If any job in the sequence fails, it sets the race status to whatever the sequence actually last committed — not blindly back to the pre-sequence value — so the next hour’s retry resumes from where it left off instead of redoing already-completed steps.
  6. After MAIN_RACE’s core steps (ingest_race, compute_season_stats) commit completed, it also runs the OpenF1 Race Events jobs (race control, overtakes, team radio) for both the Race and Sprint sessions. These run through a non-fatal wrapper — a failure here is logged but never reverts status or fails the cycle, since Race Events are supplementary and safe to pick up on a later manual run if OpenF1 has a transient issue.
  7. Because that failure is swallowed rather than retried, and the race is already completed by the time it runs (so it drops out of step 1’s query on the next poll), run_cycle also re-checks the current race-weekend window’s race on every cycle it’s still inside that window (RaceWeekendWindow’s 24h tail — see _retry_missing_race_events). If that race is completed but has zero race_control_messages rows (race control is virtually never legitimately empty for a session that happened, unlike team radio), it retries all three OpenF1 jobs again. This is the safety net for the case where FastF1’s and OpenF1’s timing genuinely diverge — e.g. a red-flag-delayed race pushes the real session end past OPENF1_LIVE_WINDOW_BUFFER, so OpenF1 wasn’t actually ready the first time.

Running Jobs Locally

cd data-engine
source venv/bin/activate

# Sync
python src/main.py --job sync_schedule     --year 2026
python src/main.py --job sync_season       --year 2026 --round 1

# Main race pipeline
python src/main.py --job ingest_qualifying --year 2026 --round 6
python src/main.py --job ingest_race       --year 2026 --round 6
python src/main.py --job compute_season_stats --year 2026
python src/main.py --job compute_features  --race_id 42
python src/main.py --job compute_predictions --race_id 42

# Sprint pipeline
python src/main.py --job ingest_sprint_qualifying --year 2026 --round 9
python src/main.py --job compute_sprint_features  --race_id 55
python src/main.py --job compute_sprint_predictions --race_id 55
python src/main.py --job ingest_sprint            --year 2026 --round 9

# Race Events (OpenF1) — --session defaults to Race; pass Sprint for the sprint session
python src/main.py --job ingest_race_control --year 2026 --round 9 --session Sprint
python src/main.py --job ingest_overtakes    --year 2026 --round 9 --session Sprint
python src/main.py --job ingest_team_radio   --year 2026 --round 9 --session Sprint

# Data-quality audit
python src/main.py --job data_quality_audit --year 2026     # latest season
python src/main.py --job data_quality_audit --all            # every season in the DB

# Repair fixable gaps found by the audit (act on the latest run for the year)
python src/main.py --job data_quality_repair --year 2026
python src/main.py --job data_quality_repair --year 2026 --resolve_run 5

Repair maps each fixable issue to the ingest+recompute jobs that own that table’s data (per quality_utils.resolve_issue_actions). It does NOT attempt data the source API doesn’t provide — e.g. FP2 long-run coverage is flagged as informational (the model falls back to historical circuit pace) rather than re-ingested, because a driver that ran no long-run stint can’t be reconstructed.

For pre-2018:

python src/main.py --job ingest_qualifying_legacy --year 2015 --round 5
python src/main.py --job ingest_race_legacy       --year 2015 --round 5

Historical Backfill

Two backfill scripts in data-engine/scripts/:

Full backfill (scripts/backfill_full.py)

Runs the complete pipeline (sync → ingest → sprint pipeline → season stats → features → predictions) for every round in a year range. Sprint weekends are automatically detected and handled.

cd data-engine
source venv/bin/activate

python scripts/backfill_full.py                        # 2000–2026 (legacy years skip gracefully)
python scripts/backfill_full.py --start 2018           # 2018–2026 (recommended — full FastF1 coverage)
python scripts/backfill_full.py --start 2025 --end 2025

Sprint-only backfill (scripts/backfill_sprint.py)

Re-runs just the sprint pipeline for specific years (useful after sprint schema changes).

python scripts/backfill_sprint.py --years 2021 2022 2024 2026
python scripts/backfill_sprint.py --years 2026          # single year

Data coverage by era:

YearsQualifyingLap timesSprint
2018–presentFull Q1/Q2/Q3 + sector timesFull per-lapFull SQ + sprint lap times
2006–2017Q1/Q2/Q3 timesNoneNone (sprint format started 2021)
2000–2005Single best lap onlyNoneNone
1990–1999None (fell back to race starting grid)NoneNone

Sprint format years: 2021 (3 rounds), 2022 (3 rounds), 2023 (6 rounds), 2024 (6 rounds), 2025 (6 rounds), 2026+.

Both scripts use _run() wrappers — a single failing round does not abort the whole year. Failures print as [SKIP] lines.


FastF1 Cache

FastF1 caches API responses locally. Enable it in development to avoid re-fetching:

import fastf1
fastf1.Cache.enable_cache('./cache')

The cache is already enabled via src/config.py. The cache/ directory is gitignored.


Idempotency

Every job uses INSERT ... ON CONFLICT DO UPDATE. Running any job twice produces identical results — no duplicate rows. This makes retries safe.

The sprint qualifying ingest uses exclude_update to prevent SQ data from overwriting already-ingested sprint race finish positions (finish_position, points, status, total_sprint_time_ms, fastest_lap are never overwritten by SQ ingest).

ingest_qualifying and ingest_sprint_qualifying both have a date guard: they check qualifying_date/sprint_qualifying_date <= today before loading FastF1, and raise an error if no results came back. This prevents future rounds from being incorrectly set to qualifying_done/sprint_qualifying_done by a backfill run.


Error Handling

  • Jobs exit with code 1 on failure so Render marks the job failed for manual retrigger.
  • Never use sleep() inside jobs — Render has a job timeout.
  • Use structured logging: {"job": "ingest_race", "round": 14, "status": "failed", "error": "..."}.