---
title: "Databricks’ native monitoring resources"
description: "Learn how to collect Databricks and Spark telemetry with system tables, the Databricks and Spark UIs, job notifications, and the Databricks REST API."
author: "Aaron Kaplan, Ryan Warrier"
date: 2026-10-02
tags: ["data observability", "data jobs monitoring", "databricks"]
blog_type_id: the-monitor
locale: en
---

In the [first part](https://www.datadoghq.com/blog/key-metrics-for-databricks-monitoring.md) of this series, we cataloged key metrics for Databricks data engineering, analytics, and Model Serving workloads. In this post, we'll discuss how to collect those metrics and other telemetry data from Databricks and Apache Spark, which powers Databricks under the hood. We'll cover [collecting and querying telemetry data via system tables](#collect-and-query-telemetry-data-via-system-tables), as well as the other primary sources of visibility into:

- [Compute performance](#compute-performance)
- [Job- and pipeline-level performance](#job-and-pipeline-level-performance)
- [Query-level performance](#query-level-performance)
- [Data quality and lineage](#data-quality-and-lineage)
- [Model Serving performance](#model-serving-performance)

## Collect and query telemetry data via system tables

[System tables](https://docs.databricks.com/aws/en/admin/system-tables/) are the primary resource for Databricks performance, cost, and usage analysis. They provide a wide range of telemetry data, from audit logs and cost records to performance metrics and SQL warehouse events.

**The data contained in system tables is not real-time; its latency varies by schema**. So while system tables provide far-reaching visibility into Databricks environments, don't rely on them for up-to-the-minute alerting. For real-time monitoring of Databricks, you can use [job notifications](https://docs.databricks.com/aws/en/jobs/notifications) and [Jobs API](https://docs.databricks.com/api/workspace/jobs) queries, as well as host-based telemetry data collected via monitoring agents or cloud providers.

By [enabling system tables](https://docs.databricks.com/aws/en/admin/system-tables/#enable-system-tables) and [delegating access via Unity Catalog](https://docs.databricks.com/aws/en/admin/system-tables/#grant-access-to-system-tables), you can [query](https://community.databricks.com/t5/technical-blog/top-10-queries-to-use-with-system-tables/ba-p/82331) a wide range of Databricks observability signals and analyze Databricks performance, cost, and usage through the Databricks UI as well as through external monitoring tools. From the Databricks UI, you can also create custom SQL-powered [dashboards](https://docs.databricks.com/aws/en/dashboards/) and [alerts](https://docs.databricks.com/aws/en/sql/user/alerts/) using any data source.

![A Databricks dashboard alongside the visualization configuration panel.](https://web-assets.dd-static.net/42588/1790884701-databricks-sql-powered-dashboards-and-alerts.png)

External tools can pull data from system tables via the Databricks SQL API for reporting and analysis. While system tables include operational data for all workspaces deployed within the same cloud region for your account, they can only be accessed through [Unity Catalog](https://docs.databricks.com/aws/en/data-governance/unity-catalog/)-enabled workspaces. The billing schema is enabled by default, providing access to the [`system.billing.usage`](https://docs.databricks.com/aws/en/admin/system-tables/billing) table, which includes cost data for all billable usage. All other system table schemas must be [enabled manually](https://docs.databricks.com/api/workspace/systemschemas/enable).

Throughout this post, we'll cover the system tables that are particularly useful for monitoring Databricks performance, including [`system.compute.*`](https://docs.databricks.com/aws/en/admin/system-tables/compute), [`system.query.history`](https://docs.databricks.com/aws/en/admin/system-tables/query-history), and [`system.lakeflow.*`](https://docs.databricks.com/aws/en/admin/system-tables/jobs#job-tasks), as well as complementary methods of monitoring performance through the Databricks UI.

## Compute performance

System tables provide visibility into both [classic](https://docs.databricks.com/aws/en/compute/use-compute) and [serverless](https://docs.databricks.com/aws/en/compute/serverless/) Databricks compute. The compute schema exposes infrastructure-level performance data for all-purpose compute and classic compute jobs and pipelines ([`system.compute.clusters`](https://docs.databricks.com/aws/en/admin/system-tables/compute#cluster-table-schema), [`system.compute.node_timeline`](https://docs.databricks.com/aws/en/admin/system-tables/compute#node-timeline-table-schema), [`system.compute.node_types`](https://docs.databricks.com/aws/en/admin/system-tables/compute#node-types-table-schema), [`system.compute.instance_events`](https://docs.databricks.com/aws/en/admin/system-tables/compute)) as well as SQL warehouses ([`system.compute.warehouse_events`](https://docs.databricks.com/aws/en/admin/system-tables/warehouse-events), [`system.compute.warehouses`](https://docs.databricks.com/aws/en/admin/system-tables/warehouses)). For serverless compute and SQL warehouse monitoring, the query schema exposes query-level metrics from the [`system.query.history`](https://docs.databricks.com/aws/en/admin/system-tables/query-history) table, and [the Databricks UI offers query insights](https://docs.databricks.com/aws/en/compute/serverless/notebooks#view-query-insights) for serverless query performance profiling.

Classic compute clusters emit metrics and logs that can be collected via monitoring agents, forwarded to external storage, and accessed in Databricks via the Compute page.

![Metrics tab of a Databricks cluster detail page showing server load distribution, CPU utilization and active nodes, container memory usage, JVM heap usage, and GC pause time.](https://web-assets.dd-static.net/42588/1790884706-databricks-compute-page-cluster-metrics.png)

See the Databricks [documentation](https://docs.databricks.com/aws/en/compute/configure#cluster-log-delivery) to learn more about configuring compute log delivery destinations and retention.

The [Spark UI](https://docs.databricks.com/aws/en/optimizations/spark-ui-guide) offers another layer of visibility into classic compute, enabling you to closely track and analyze Databricks job, pipeline, and query performance by monitoring the Spark jobs powering your Databricks workloads under the hood. You can use it to troubleshoot and optimize Spark job-level execution via the [jobs timeline](https://docs.databricks.com/aws/en/optimizations/spark-ui-guide/jobs-timeline); [DAG visualizations](https://docs.databricks.com/aws/en/compute/troubleshooting/debugging-spark-ui#job-details-page); metrics for [skew](https://docs.databricks.com/aws/en/optimizations/spark-ui-guide/long-spark-stage-page#skew), [spill](https://www.databricks.com/discover/pages/optimize-data-workloads-guide#data-spilling), [shuffle](https://www.databricks.com/discover/pages/optimize-data-workloads-guide#data-shuffling), and [streaming jobs](https://docs.databricks.com/aws/en/compute/troubleshooting/debugging-spark-ui#streaming-tab); [JVM thread dumps](https://docs.databricks.com/aws/en/compute/troubleshooting/debugging-spark-ui#thread-dump); and [driver and executor logs](https://docs.databricks.com/aws/en/compute/troubleshooting/debugging-spark-ui#driver-logs). It is particularly useful for [troubleshooting slow performance](https://docs.databricks.com/aws/en/optimizations/spark-ui-guide/long-spark-stage).

![Spark UI Spark Jobs page with an event timeline.](https://web-assets.dd-static.net/42588/1790884711-spark-ui-spark-jobs-event-timeline.png)

### SQL warehouse performance

Alongside the [`system.compute.warehouse_events`](https://docs.databricks.com/aws/en/admin/system-tables/warehouse-events) and [`system.compute.warehouses`](https://docs.databricks.com/aws/en/admin/system-tables/warehouses) tables, the Databricks UI provides a dedicated dashboard for each of your [SQL warehouses](https://docs.databricks.com/aws/en/compute/sql-warehouse/). [SQL warehouse dashboards](https://docs.databricks.com/aws/en/compute/sql-warehouse/monitor/#activity-details) visualize a range of metrics alongside logs of active and historical query and cluster activity.

![Monitoring tab of a Databricks SQL warehouse showing peak query count, completed query count, running clusters, and a query history table.](https://web-assets.dd-static.net/42588/1790884717-databricks-sql-warehouse-monitoring-dashboard.png)

## Job- and pipeline-level performance

Tracking Databricks job and pipeline performance is essential to ensuring reliability and enabling optimization and troubleshooting of tasks and scheduling. The [`system.lakeflow.*`](https://docs.databricks.com/aws/en/admin/system-tables/jobs#job-tasks) tables provide job metadata such as job and task definitions and schedules ([`system.lakeflow.jobs`](https://docs.databricks.com/aws/en/admin/system-tables/jobs#jobs-table-schema), [`system.lakeflow.job_tasks`](https://docs.databricks.com/aws/en/admin/system-tables/jobs#job-tasks)), as well as job performance metrics ([`system.lakeflow.job_run_timeline`](https://docs.databricks.com/aws/en/admin/system-tables/jobs#runs), [`system.lakeflow.job_task_run_timeline`](https://docs.databricks.com/aws/en/admin/system-tables/jobs#job-task-run-timeline-table-schema)). Because `system.lakeflow.*` data is not real-time, these tables should not be relied on for live alerting on job and task failures. To receive real-time notifications when jobs succeed, fail, start, or exceed expected durations, Databricks provides built-in [job notifications](https://docs.databricks.com/aws/en/jobs/notifications) that can be routed to email, Slack, Microsoft Teams, PagerDuty, or [any HTTP webhook](https://docs.databricks.com/aws/en/jobs/notifications#http-webhook-payloads).

Databricks also provides detailed visibility into [jobs](https://docs.databricks.com/aws/en/jobs/#monitoring-and-observability) and [pipelines](https://docs.databricks.com/aws/en/ldp/observability) via the jobs and pipelines page, the [pipeline event log](https://docs.databricks.com/aws/en/ldp/monitor-event-logs), and the [Spark Streaming Query Listener](https://docs.databricks.com/aws/en/structured-streaming/stream-monitoring). The pipeline event log is [SQL-queryable](https://docs.databricks.com/aws/en/ldp/monitor-event-logs#query-the-event-log) via the Pipelines API, and with the Spark Streaming Query Listener, you can [send streaming metrics to external services](https://docs.databricks.com/aws/en/structured-streaming/stream-monitoring).

![Databricks job Runs page.](https://web-assets.dd-static.net/42588/1790884723-databricks-job-runs.png)

## Query-level performance

Along with being useful for monitoring serverless compute and SQL warehouses, visibility into query performance is important for root cause analysis and fine-tuned optimization of your Databricks jobs and pipelines. In addition to the [`system.query.history`](https://docs.databricks.com/aws/en/admin/system-tables/query-history) table, which you can use to track and dissect query durations, the [query history](https://docs.databricks.com/aws/en/sql/user/queries/query-history) page in the Databricks UI enables in-depth [query performance profiling](https://docs.databricks.com/aws/en/sql/user/queries/query-profile), and query insights surface Spark statement metrics inline for queries run on serverless compute.

![Databricks notebook cell querying system.access.audit, with an inline performance panel.](https://web-assets.dd-static.net/42588/1790884727-databricks-query-insights.png)

You can also [access query history via the Databricks REST API](https://docs.databricks.com/api/workspace/queryhistory).

## Data quality and lineage

Monitoring data lineage can help you find upstream root causes of issues affecting your data and, more broadly, [follow the movement and development of data throughout your pipelines](https://www.datadoghq.com/blog/data-lineage.md). Databricks provides [data profiling](https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-quality-monitoring/data-profiling/) via Unity Catalog and a [Data Quality Monitoring UI](https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-quality-monitoring/anomaly-detection/#-view-data-quality-monitoring-results-in-the-ui) to help you track data quality and integrity. You can [define profiles through the Databricks UI](https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-quality-monitoring/data-profiling/create-monitor-ui) or [the API](https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-quality-monitoring/data-profiling/create-monitor-api) in order to monitor critical metrics such as data completeness and freshness and use them to set SQL alerts. Databricks also provides [anomaly detection](https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-quality-monitoring/anomaly-detection/) that can help you track these metrics throughout schemas using the Data Quality Monitoring UI as well as health indicators in the [Catalog Explorer](https://docs.databricks.com/aws/en/catalog-explorer/).

![Databricks Data Quality Monitoring page.](https://web-assets.dd-static.net/42588/1790884729-databricks-data-quality-monitoring.png)

The [`system.access.table_lineage`](https://docs.databricks.com/aws/en/admin/system-tables/lineage) and `system.access.column_lineage` tables record lineage events, and their `entity_metadata` fields can be used to trace data events to specific jobs and notebooks. Databricks [audit logs](https://docs.databricks.com/aws/en/admin/account-settings/audit-logs), which are queryable via the [`system.access.audit`](https://docs.databricks.com/aws/en/admin/system-tables/audit-logs) table, can be used to track actions taken across your account, such as configuration, schema, and permission changes. Together, these tables enable root cause analysis of issues affecting data quality. You can also [enable audit log delivery](https://docs.databricks.com/aws/en/admin/account-settings/audit-log-delivery) to external storage. The Databricks UI visualizes lineage with [lineage graphs in the Catalog Explorer](https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-lineage), and [Unity Catalog automatically captures lineage](https://docs.databricks.com/aws/en/data-governance/unity-catalog/data-lineage#data-lineage-overview) down to the column level for queries run on Databricks. The lineage system tables retain data over a rolling one-year window, while the Catalog Explorer retains lineage data indefinitely. Unity Catalog also allows you to [add lineage metadata from external sources](https://docs.databricks.com/aws/en/data-governance/unity-catalog/external-lineage).

![Unity Catalog lineage graph tracing upstream views and tables into a compact_car_sales table and on to downstream features, an external location, and a model version.](https://web-assets.dd-static.net/42588/1790884734-unity-catalog-lineage-graph.png)

Finally, you can collect Spark-level lineage metadata from Databricks by [installing the OpenLineage Spark integration](https://openlineage.io/docs/integrations/spark/quickstart/quickstart_databricks/) on classic compute clusters. From there, you can store and visualize this data using tools like [Marquez](https://marquezproject.ai/) (a popular open source solution) or Datadog (as we'll discuss in [Part 3](https://www.datadoghq.com/blog/how-to-monitor-databricks-with-datadog.md) of this series).

## Model Serving performance

[Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/) deploys AI/ML models and agents behind HTTPS endpoints on Databricks-managed serverless compute. Databricks provides a [range of resources](https://docs.databricks.com/aws/en/machine-learning/model-serving/monitor-diagnose-endpoints) for tracking Model Serving performance in order to optimize and troubleshoot AI/ML workloads. It [exposes endpoint health metrics](https://docs.databricks.com/aws/en/machine-learning/model-serving/metrics-export-serving-endpoint) in the [OpenMetrics](https://openmetrics.io/) format at `https://[DATABRICKS_HOST]/api/2.0/serving-endpoints/[ENDPOINT]/metrics`, so you can track live request and error counts, latency, and compute usage with Prometheus, Datadog, or any OpenMetrics-compatible agent.

From the Databricks Serving UI, you can access endpoint health metrics from the previous 14 days. You can also access [service](https://docs.databricks.com/api/workspace/servingendpoints/logs) and [build logs](https://docs.databricks.com/api/workspace/servingendpoints/buildlogs) from the UI or through the API; service logs are ephemeral, while build logs are retained for up to 30 days.

![Databricks Model Serving endpoint page for an A/B test, showing two model versions split 60/40 with latency and request rate charts below.](https://web-assets.dd-static.net/42588/1790884739-databricks-model-serving-endpoint-metrics.png)

Databricks also logs Model Serving endpoint inputs and responses in [inference tables](https://docs.databricks.com/aws/en/machine-learning/model-serving/inference-tables), which capture requests and responses for AI/ML performance analysis. Inference tables must be [enabled for individual endpoints](https://docs.databricks.com/aws/en/machine-learning/model-serving/inference-tables#enable-and-disable-inference-tables). You can track model performance by defining data profiles and SQL alerts on your inference tables, and Databricks provides customizable [LLM judges and scorers](https://docs.databricks.com/aws/en/mlflow3/genai/eval-monitor/concepts/scorers) for tracking model and agent performance.

For Model Serving endpoints that sit behind the Databricks [AI Gateway](https://docs.databricks.com/aws/en/ai-gateway/), you can also access metrics via the [`system.ai_gateway.usage`](https://docs.databricks.com/aws/en/ai-gateway/usage-tracking-beta#usage-table-schema) table for performance monitoring and cost tracking.

## Comprehensively monitor Databricks

In this post, we've surveyed the primary monitoring solutions provided by Databricks, including system tables, the Databricks UI and SQL API, the Spark UI, and built-in job notifications. In [Part 3](https://www.datadoghq.com/blog/how-to-monitor-databricks-with-datadog.md) of this series, we'll show you how to use Datadog's Databricks integration and Data Observability: Jobs Monitoring for comprehensive monitoring of your Databricks data engineering, analytics, and Model Serving workloads.