Solution · Real-Time Dashboard

Real-time operations dashboard for Apache Airflow

Refreshing the Airflow UI is not monitoring. BatchFoundry deploys a live operational dashboard fed by Prometheus and the Airflow REST API — what was a Control-M ViewPoint screen, now built on open standards.

Why real-time matters

Critical batch windows — overnight settlements, market open, payroll cutoff — need NOC-style screens that update every few seconds. Airflow's UI auto-refresh queries the metadata database and becomes expensive at scale. A dedicated observability stack (Prometheus + Grafana) handles this with sub-second latency at any workload size.

The reference architecture

Airflow statsd metrics → statsd_exporter → Prometheus

Custom DAG-level metrics via callbacks → Prometheus

Prometheus → Grafana dashboards

Airflow REST API → per-task drill-down panels in Grafana

Optional: thin React dashboard for executive viewing on lobby displays.

Standard panels we ship

  • Live job count: queued, running, succeeded, failed in the last 60 minutes
  • SLA breach ticker: DAGs missing SLA right now
  • Worker pool saturation by queue
  • Success-rate heatmap by hour-of-day × day-of-week
  • Top-N longest-running tasks with trend vs. 7-day baseline
  • Cluster-wide latency from queue → start
  • Per-region split for multi-region Airflow deployments

What BatchFoundry delivers

  • statsd_exporter + Prometheus + Grafana stack as code (Helm charts or Terraform)
  • Pre-built dashboard panels — importable JSON, version-controlled
  • Custom DAG callback library emitting business-metric events alongside infrastructure metrics
  • Optional NOC kiosk view for operations floor display

Pitfalls we plan around

  • Prometheus retention vs. cardinality — don't tag by dag_id if you have 10k+ DAGs
  • XCom-based metric exfiltration — use callbacks, not XCom reads
  • Authn / authz on the Grafana dashboard — lock to corporate SSO, not open by default

What we won't tell you

A simple Slack #batch-alerts channel is enough for some teams. We'll tell you when a 24/7 NOC dashboard is solving a problem you don't have.

Talk to an observability engineer

We'll scope the dashboard stack against your workload size and critical batch windows.

Talk to an observability engineer