Bring data quality together. In Databricks and beyond.

Data teams at
Data contracts
Bring every check for a dataset together. Manage technical checks and business requirements in shared contracts, across Databricks and your other sources.
- row_count:never empty - freshness:updated within 6 hours - duplicate:no repeated orders - missing:never empty - invalid:paid, shipped or refunded data_type: decimalalways a number
vs DatabricksData observability
Find the issue behind the metric. Use monitoring to spot changes and business checks to find records that break your rules.
vs DatabricksAI-assisted coverage
Cover more data without writing every check. Let Autopilot draft contracts. Add your team’s business requirements with Copilot, the UI or code.
Contract AutopilotReady to draft
vs DatabricksMigration testing
Verify what reaches Databricks. Find missing records and changed values before the switch. Keep checking quality after the move.
vs DatabricksInfrastructure
Keep failed records in your own warehouse. Investigate issues in your diagnostics warehouse, with shared check results across Databricks and your other sources.
Cloud
vs DatabricksHow 2K, HelloFresh and Make use Soda.
“Data contracts brought transparency and also bigger cooperation between different teams.”


We went from detection to prevention. Most data quality tooling will tell you something broke after it broke. Our goal is to catch it before that reaches the dashboards and datasets in the first place.

Investing in data quality is key for cross-functional teams to make accurate, complete decisions…

A business analyst or data analyst can write and provision checks themselves through self-service.

Soda has integrated seamlessly into our technology stack.

We're catching issues we never would've noticed before.

It's about how it reshaped our approach to data quality and product mindset.

Frequently asked questions
Choosing a data quality platform.
Who writes the checks. How you get coverage. Where your data stays.
Explore the documentationAdd Soda when your data quality requirements span Databricks, other data sources and the teams that use them. Soda brings business checks, data contracts, monitoring and diagnostics into a shared platform. Engineers can work in code while business users contribute requirements through the UI. You can keep Databricks as your data platform and use Soda to manage quality across your data sources.
See Soda working with DatabricksLakeflow expectations validate records as pipelines run. Delta NOT NULL and CHECK constraints enforce conditions when data is written; primary and foreign keys are informational. Unity Catalog quality monitoring detects changes in table health and profiles data. DQX is a separate Databricks Labs framework for Spark validation, with rules, quarantine, a browser UI and other quality features. These tools have different execution models, availability and support terms.
Soda lets teams author shared contracts through code, a no-code editor or AI. DQX Studio also offers browser-based rules, AI suggestions, approvals and ODCS contract import. Its current documentation labels Studio Experimental. Both provide ways to capture business requirements; the choice is how you want to manage those requirements across your teams and data sources.
Yes. DQX can train models on representative data to identify unusual records and explain which fields contributed to each result. This differs from native table monitoring for freshness and completeness. DQX documents batch training and streaming scoring; its models do not currently retrain automatically. Teams need to plan for reviewing results and retraining as normal data patterns change.
DQX is a Databricks Labs project provided without a formal Databricks support SLA. Its maintainers direct issues to GitHub, where they are reviewed as time permits. This applies to DQX, not to Databricks’ supported native products. Teams evaluating DQX should account for operating and upgrading the framework. Soda offers commercial product support, with premium support on its Enterprise plan.
Yes. Existing controls can continue running while you introduce Soda on selected datasets or requirements. Keeping native controls does not itself import their definitions or results into Soda; migration and integration requirements need a separate review. Existing business logic can be adapted into Soda contract checks, including custom SQL, with validation before older checks are retired.
Explore Soda’s contract checksYes. Soda supports Databricks SQL alongside sources such as Snowflake, BigQuery, PostgreSQL and SQL Server. Teams can manage quality requirements across those systems in Soda rather than limiting their quality program to one platform. Connection methods and available features vary by source, so review the supported integration for each system you want to cover.
See Soda’s supported data sourcesYes. Soda supports metric and record-level reconciliation between source and target datasets. Keyed row comparisons can identify missing records and mismatched values before a migration is complete. Teams can then keep running quality checks against the target. Run reconciliation through Soda Runner or the reconciliation extension for Python workflows; confirm plan and source support when planning the migration.
Explore source-to-target reconciliationSoda offers hosted and self-hosted Runners. With a self-hosted Runner, checks execute in your environment; a Diagnostics Warehouse keeps failed-row data in your warehouse for investigation. Soda Cloud receives metadata, check results and monitoring metrics. The exact data flow depends on your deployment, so choose the Runner and diagnostics configuration that meets your access and privacy requirements.
Review Soda’s data flows
Introduction meeting 30 minutes Zoom
Put your checks where put theirs.
Thirty minutes with the Soda team to learn how your checks work today and how Soda can help you and your team.
- Two short steps, then pick a time
- Soda runs inside your own network
- 1You
- 2Your company
- 3Your time
vs