AI and ML platform guide
Use this page to design and operate machine learning, generative AI, and natural-language analytics workloads on Azure Databricks. Product names and preview status can change. Check the current Azure documentation and Azure release notes before you select a feature.
Platform map#
| Need | Start with | Focus |
|---|---|---|
| Data and AI governance | Unity Catalog | Access and lineage |
| Experiment lifecycle | MLflow | Reproducibility and promotion |
| Reusable ML features | Feature Store | Reuse and training-serving consistency |
| Business questions | Genie Agents and metric views | Meaning and answer accuracy |
| Model and MCP access | Unity AI Gateway | Policy, traffic, and cost |
| Online inference | Model Serving | Latency, scale, and health |
| AI retrieval | AI Search | Access, freshness, and relevance |
| Agent or RAG system | MLflow 3 and serving | Evaluation and rollback |
| Agent-assisted delivery | Databricks Agent Skills | Routing and access |
Do not start with the newest feature. Start with the workload's data boundary, quality target, latency, traffic, and failure mode. These requirements determine whether a batch job, SQL function, model endpoint, Genie Agent, or application is the correct delivery surface.
Production baseline#
- Use one governance layer. Put data, features, models, functions, model services, and evaluation datasets under explicit Unity Catalog ownership. Use groups and least privilege.
- Separate each environment. Keep development, staging, and production data and resources separate. Promote tested code and versioned artifacts through an automated release process.
- Standardize the lifecycle. Use MLflow to record experiments, models, prompts, traces, evaluations, and production feedback. Keep the code and data version with the result.
- Evaluate the actual task. Use representative business examples and measure the errors that matter. Generic scores do not prove that a model, Genie Agent, or agent is ready.
- Make authorization explicit. Confirm which identity reads each source and invokes each tool. Record any intended use of owner or service principal privileges.
- Trace and monitor. Capture latency, failures, versions, retrieval context, and quality signals. Apply a data policy before you store request text, responses, or traces.
- Limit cost and failure impact. Set owners, budgets, rate limits, timeouts, fallbacks, and a rollback target before you open a workload to broad use.
- Treat preview features as replaceable. Isolate a preview behind a small interface. Record its status and region support during each design review.
Machine learning lifecycle#
- Use MLflow tracking or autologging for each training run. Record parameters, metrics, dependencies, data lineage, and artifacts needed to reproduce the result.
- Use Feature Store when teams must reuse features, apply point-in-time joins, or prevent training-serving skew. Keep feature tables and specifications in Unity Catalog.
- Register production models in Unity Catalog. Use a deployment job and a model alias to identify an approved model version.
- Use a Lakeflow Job for batch inference. Use Model Serving when the application needs a managed, low-latency REST endpoint.
- Monitor model quality and input changes after release. Also monitor endpoint latency, errors, resource use, and cost.
Generative AI and agents#
- Trace the complete application with MLflow 3. Include model calls, retrieval, tool calls, and important application steps.
- Build an evaluation set from real tasks, known failures, and expert feedback. Use business scorers and safety checks in development and in the release gate.
- Reuse the same scorers on a sample of production traces. A human must review high-risk outputs and judge failures that an automated scorer cannot resolve.
- Use AI Search for retrieval from approved sources. Test authorization, source freshness, retrieval quality, answer grounding, and citations as separate concerns.
- Store production traces and endpoint telemetry in Unity Catalog when the data policy permits it. Set access, sampling, retention, and redaction before you enable payload logging.
- Version the prompt, model, tool set, retrieval configuration, evaluation set, and release result. Keep a tested fallback for a model retirement or provider failure.
Genie#
Use Genie One as the user experience and Genie Agents as curated, domain-specific interfaces to business data. Treat Genie Code as an authoring aid. Do not use one large agent as a replacement for data modeling or a catalog.
- Start with one business domain and a small set of well-documented Unity Catalog assets.
- Put reusable measures, dimensions, joins, and filters in Unity Catalog metric views. Use clear business names and resolve ambiguous terms.
- Add SQL expressions, example queries, instructions, and trusted assets for known business rules. Keep instructions short and remove conflicts.
- Use a serverless SQL warehouse when it is available. Users still need access to the Genie
Agent and
SELECTon its data, although the author's credentials provide warehouse access. - Test realistic questions and review the generated SQL. Add two to four phrasings for each important benchmark and compare results with approved SQL or a trusted function.
- Review usage, slow queries, failed answers, and user feedback. Feedback does not change the agent automatically. A named curator must approve each context change.
Unity AI Gateway#
Use Unity AI Gateway as the control plane for model and MCP access. Apply the same pattern to Databricks-hosted models and external providers.
- Create approved model services as Unity Catalog objects. Grant
EXECUTEto groups and review the default access to services insystem.ai. - Register external providers so callers can use them without receiving provider credentials.
- Review the service owner's privileges. A model service uses the owner's privileges to invoke its primary and fallback destinations.
- Set service and user rate limits. Add budgets and request tags for cost ownership. Clients must handle HTTP 429 responses with exponential backoff.
- Configure routing and ordered fallbacks when the workload needs rollout control or provider
resilience. Verify routing decisions and usage in
system.ai_gateway.usage. - Use inference tables only when policy permits payload storage. They support quality review and diagnosis, but they do not provide a complete audit record.
- Apply service policies for sensitive data, prompt injection, unsafe content, and custom rules. Service policies are Beta, so keep application controls and evaluation gates in place.
Review questions#
- What data may the workload read, and can the answer reveal data the caller could not query?
- What test set blocks a bad model, prompt, retrieval change, or agent tool from promotion?
- Which identity invokes the workload, and which identity reaches downstream data or services?
- Where are traces retained, who can read them, and how are sensitive values removed?
- What happens when the model is slow, unavailable, over budget, or confidently wrong?
- Which model services can each group use, and did you review the default
system.aigrants? - Who curates each Genie Agent, and which benchmark proves its important answers?
- Can the team reproduce and roll back the exact production version?
Official sources#
- Azure Databricks machine learning overview — https://learn.microsoft.com/azure/databricks/machine-learning/
- MLflow on Azure Databricks — https://learn.microsoft.com/azure/databricks/mlflow/
- Feature Store — https://learn.microsoft.com/azure/databricks/machine-learning/feature-store/
- Model Serving — https://learn.microsoft.com/azure/databricks/machine-learning/model-serving/
- Model Serving monitoring — https://learn.microsoft.com/azure/databricks/machine-learning/model-serving/monitor-diagnose-endpoints
- AI Search — https://learn.microsoft.com/azure/databricks/generative-ai/vector-search
- MLflow 3 evaluation and monitoring — https://learn.microsoft.com/azure/databricks/mlflow3/genai/eval-monitor/
- Curate a Genie Agent — https://learn.microsoft.com/azure/databricks/genie-agents/best-practices
- Test and monitor a Genie Agent — https://learn.microsoft.com/azure/databricks/genie-agents/monitor
- Unity Catalog metric views — https://learn.microsoft.com/azure/databricks/business-semantics/metric-views/
- Unity AI Gateway — https://learn.microsoft.com/azure/databricks/ai-gateway/
- Govern model services — https://learn.microsoft.com/azure/databricks/ai-gateway/govern-model-services
- Track model usage — https://learn.microsoft.com/azure/databricks/ai-gateway/usage-tracking
- Azure Databricks release notes — https://learn.microsoft.com/azure/databricks/release-notes/