Databricks platform lessons
These are recommendations, not product limits.
Where these lessons come from#
These lessons reflect 536Tech client work on production Azure Databricks platforms. For questions or guidance, feel free to reach out at hello@536tech.com.
Plan before anyone builds#
- Decide the workspace and catalog layout before teams start. A workspace per environment gives you a place to test storage, compute policies, and init scripts before production.
- One pattern has worked well for me: combine the environment with the medallion layers. Use bronze, silver, and gold catalogs in each workspace. Use schemas for sources or data products. Use catalogs as access and isolation boundaries.
- Do not let notebooks come first. Create the catalogs, groups, policies, and service principals before teams build workloads.
Pick one owner for each resource#
- Terraform is a good fit for platform resources such as networking, storage, workspaces, Unity Catalog, identity, and compute policies.
- Declarative Automation Bundles are a good fit for jobs, pipelines, notebooks, and other resources tied to one workload. The workload author owns the bundle.
- Both tools can manage some of the same resources. That is where teams get into trouble. Two repositories are fine. Two owners for the same remote resource are not. See Terraform and bundles.
Control Unity Catalog grant ownership#
databricks_grantsis authoritative. It removes any grant that it does not declare. A bundle changes catalog or schema grants only when its configuration includes them. If both tools need to change grants on one object, usedatabricks_grantfor granular Terraform ownership or pick one grant owner. See Unity Catalog grants.- Grant access to groups, not users. Map Entra ID or Okta groups to Databricks account groups, then grant access to those groups.
Set sensible compute defaults#
- There is no useful universal auto-termination value. Pick a timeout that fits the workload and enforce it with a compute policy.
- Serverless SQL warehouses typically start in two to six seconds. Pro and classic warehouses typically take about four minutes. Size each warehouse for its workload and set a short automatic stop period.
- Use compute policies to limit what users can create.
Machines run production jobs#
- Run production jobs as service principals. Use a separate principal for each environment when the environments need separate permissions.
- Keep access groups in Entra ID. Use Automatic Identity Management when the account supports it. Keep SCIM where AIM does not fit.
- Use workload identity federation for CI/CD. Do not keep long-lived PATs or service-principal secrets in the pipeline.
Separate secrets by use#
- Put secrets that a job reads in a Databricks secret scope.
- Keep CI/CD secrets outside Databricks. Use Key Vault, GitHub Secrets, or the external secret store that the organization already manages. GitHub Secrets is enough for many GitHub-based pipelines.
Plan network capacity before deployment#
- Azure does not let you resize an assigned subnet CIDR in place. You can now replace the workspace
subnets or move the workspace to another VNet without replacing the workspace. Stop all compute
first and use API version
2026-01-01. Terraform does not support this update yet. - Plan enough capacity before deployment. A later subnet change is possible, but it is disruptive.
- Classic compute does not need public IP addresses. It still needs outbound access through a NAT gateway, firewall, or another supported route.
- The Azure workspace pricing SKU lists Standard and Premium. Private Link requires Premium.
- New Standard workspaces have not been available since
2026-04-01. Azure upgrades any Standard workspace left on2026-10-01to Premium. - Classic compute downloads its images and startup artifacts from regional Blob storage. Use an Azure Storage service endpoint and service endpoint policy to keep that traffic off the NAT gateway or firewall. Classic Private Link handles cluster-to-control-plane traffic instead.
Documentation checked 2026-08-27: Manage your subscription, End of life for Standard tier workspaces, Private Link concepts, Update workspace networking, Artifact storage networking.
Cost and observability belong to the platform team#
- Billing, compute, and jobs system schemas are on by default in Unity Catalog workspaces. If a query fails, check the grants first. Enable another system schema only when it is available but disabled. See the system tables documentation.
- Do not guess cost from the SKU name. Check the unit price and the consumed quantity.
Cost spike triage (first 15 minutes)#
Start with the system tables.
- Calculate spend by SKU. Join
system.billing.usagetosystem.billing.list_pricesfor the correct price period. Group the result bysku_name, then inspect the resources behind the largest changes. - Check classic compute. Use the latest record in
system.compute.clustersto find a timeout of0orNULL. That is a warning, not proof that the cluster is running or idle. Confirm waste with compute activity data. - Trace jobs that use all-purpose compute. Start with the high-cost cluster IDs, then find the jobs that use them. Shared all-purpose compute cannot attribute its cost precisely to one job. Job compute or serverless gives you better attribution and usually less idle cost.
- Check the tags. Serverless usage policies put cost attribution tags in the billing records. Check the selected policy on each asset with missing tags. Use the procedure below.
If someone cannot query the billing or compute tables, fix the Unity Catalog grants first.
Use Databricks' cost query examples to investigate daily spend, warehouse trends, and costs by tag.
Attribute serverless costs to workloads and teams#
Use billing records to assign serverless costs to workloads and teams. Job ownership alone leaves notebook and shared-platform costs unresolved. Agree which field takes precedence when tags and resource ownership identify different teams.
- Select a billing period and workspace scope in
system.billing.usage. - Include correction records, such as retractions and restatements, in the usage total.
- Separate usage by SKU and unit before you compare quantities.
- Join applicable price periods from
system.billing.list_pricesto estimate list-price cost. - Apply that attribution rule to policy tags, resource owners, and workload identities.
- Report shared and unresolved costs separately when you cannot identify a responsible team.
List-price cost is an estimate. It does not include the customer's negotiated discounts. See the billing schema and serverless billing examples.
| Workload | Attribution fields and limits |
|---|---|
| Serverless job | Use usage_metadata.job_id and job_run_id to identify the job and run. |
| Serverless notebook | Use usage_metadata.notebook_id. On shared notebooks, identity_metadata.run_as identifies the session creator. |
| Pipeline | Use usage_metadata.dlt_pipeline_id and the policy selected on the pipeline. |
| Shared platform use | Retain a shared-cost category with a named operational owner. |
| Unresolved usage | Report the unattributed amount instead of assigning it to an arbitrary team. |
Serverless usage policies are in Public Preview. They add tags for attribution; they are not a spending cap. Existing notebooks, jobs, and pipelines need explicit policy assignment after their owners receive policy access. A pipeline triggered by a job does not inherit that job's policy. Policy changes affect usage initiated afterward, rather than an active run.
Check a new billing record for the expected tags after each policy change. Report how much cost has an assigned team and how much remains unresolved. Name an owner to investigate the unresolved costs.
Community example: cost attribution across jobs, notebooks, and shared platform use.
Bring a brownfield platform into Terraform#
- Importing a resource does not change it. The first plan can. Import in phases: networking, storage, workspace, grants, then cleanup. Read each plan for a replace or destroy action before you apply it.
- People still make emergency production changes. Add each valid change to Terraform quickly, then return to PR-based change control. Otherwise, drift becomes normal.
Keep the change process boring#
- Put Terraform changes through a pull request. Post the plan. Use CODEOWNERS for each environment. Lock apply concurrency.
- Pin provider versions. Pin GitHub Actions to full commit SHAs. Keep plan artifacts for the period that the audit policy requires. GitHub uses 90 days by default.
Add a service only when it fills a gap#
- Lakeflow Jobs handles dependencies between Databricks tasks. Use ADF when the workflow must also coordinate other Azure services.
- Unity Catalog governs Databricks data and AI assets. Use Purview when the organization needs a catalog across platforms.
User interface#
- Databricks has a dark theme in the user settings.