Real-time operations dashboard for Apache Airflow
Refreshing the Airflow UI is not monitoring. BatchFoundry deploys a live operational dashboard fed by Prometheus and the Airflow REST API — what was a Control-M ViewPoint screen, now built on open standards.
Why real-time matters
Critical batch windows — overnight settlements, market open, payroll cutoff — need NOC-style screens that update every few seconds. Airflow's UI auto-refresh queries the metadata database and becomes expensive at scale. A dedicated observability stack (Prometheus + Grafana) handles this with sub-second latency at any workload size.
The reference architecture
Airflow statsd metrics → statsd_exporter → Prometheus
Custom DAG-level metrics via callbacks → Prometheus
Prometheus → Grafana dashboards
Airflow REST API → per-task drill-down panels in Grafana
Optional: thin React dashboard for executive viewing on lobby displays.
Standard panels we ship
- Live job count: queued, running, succeeded, failed in the last 60 minutes
- SLA breach ticker: DAGs missing SLA right now
- Worker pool saturation by queue
- Success-rate heatmap by hour-of-day × day-of-week
- Top-N longest-running tasks with trend vs. 7-day baseline
- Cluster-wide latency from queue → start
- Per-region split for multi-region Airflow deployments
What BatchFoundry delivers
- statsd_exporter + Prometheus + Grafana stack as code (Helm charts or Terraform)
- Pre-built dashboard panels — importable JSON, version-controlled
- Custom DAG callback library emitting business-metric events alongside infrastructure metrics
- Optional NOC kiosk view for operations floor display
Pitfalls we plan around
- Prometheus retention vs. cardinality — don't tag by dag_id if you have 10k+ DAGs
- XCom-based metric exfiltration — use callbacks, not XCom reads
- Authn / authz on the Grafana dashboard — lock to corporate SSO, not open by default
What we won't tell you
A simple Slack #batch-alerts channel is enough for some teams. We'll tell you when a 24/7 NOC dashboard is solving a problem you don't have.
Talk to an observability engineer
We'll scope the dashboard stack against your workload size and critical batch windows.
Talk to an observability engineer