Databricks hit a $6.9bn annualised revenue run rate in June 2026, growing around 80% year on year, and is valued at $134bn with reports of a new round in the $165–175bn range. Its net dollar retention has stayed above 140% — meaning the average existing customer spends 40% more this year than last.
That last number is the one worth pausing on, and it cuts both ways. Retention above 140% means customers who adopt it expand aggressively. It also means your spend will grow every year whether or not you planned for it. Both facts should be in your business case.
This article explains what the product actually does, what it costs, and how to tell whether you need it — without the architecture diagrams.
The problem it exists to solve
For thirty years, large organisations have run two separate data estates.
The warehouse holds clean, structured, organised data. Sales by region, claims by month, transactions by account. It answers questions you already knew you'd ask. It's what your BI dashboards read from. It's expensive per terabyte, and getting data into it requires a project.
The lake holds everything else. Log files, documents, images, sensor readings, contracts, call recordings — raw, cheap to store, and mostly unusable without an engineer.
The split existed because the technology forced it. Warehouses were good at structured queries and bad at scale and unstructured content. Lakes were the reverse.
The consequence, which any executive who's tried to fund a data project will recognise: your reporting runs on one estate and your machine learning runs on the other, they disagree, and reconciling them is a permanent tax. Two copies of the truth, two teams, two tools, and a recurring argument about which number is correct.
Databricks' proposition — the "lakehouse" — is to collapse the two into one. One place to store everything, one place to govern access, one place from which both the quarterly report and the fraud model read.
That's the whole idea. Everything else is implementation.
Why this matters more now than it did five years ago
Two things changed.
Machine learning became business-critical rather than experimental. When ML was a research function, it was tolerable that data scientists worked from exports. Now that models sit in claims processing and credit decisions, they need live, governed, auditable access to the same data your finance team uses. Exports don't survive an audit.
Generative AI made unstructured data valuable. Your contracts, emails, call transcripts and PDFs were previously storage costs. They're now the raw material for anything useful an LLM can do inside your organisation. Which means the "lake" side of the estate — the messy side — suddenly needs the governance the warehouse side always had.
Both changes point the same way: the split between the two estates went from an annoyance to a blocker.
What it actually costs
Vendors publish per-unit pricing that tells you almost nothing about your bill. Here's the honest structure.
Consumption, not licences. You pay for compute used, measured in the vendor's own units, plus your cloud provider's charges for the underlying machines and storage. There is no seat cost and no floor. This is genuinely attractive — and genuinely dangerous, because nobody has to approve anything for the bill to go up.
Storage is nearly free. Compute is not. Storing a petabyte costs very little. Repeatedly scanning it because a query was written carelessly costs a great deal. In practice, most of the variance between a well-run and badly-run deployment is query discipline, not platform choice.
The people cost exceeds the platform cost, usually by a lot. Budget for platform engineers who understand the tool. On most programmes I've seen, the fully-loaded cost of the team is two to four times the platform bill in year one.
Migration is the invisible line item. Moving from an existing warehouse is a year of work for a large estate, and it must run in parallel with the old system for much of that time. You pay for both. This is the number most business cases understate.
A realistic total cost of ownership for a large enterprise deployment is: platform consumption + cloud infrastructure + a platform team + parallel running during migration + a governance function that doesn't currently exist. If your business case has only the first line, it isn't a business case.
The single most important cost control: put a named owner on the consumption budget from day one, with a monthly review. Consumption pricing without an owner is how organisations discover a seven-figure overrun in month nine.
Where the return actually comes from
Vendor case studies quote dramatic percentages. Treat them as marketing. Here are the return mechanisms that hold up under scrutiny, roughly in order of reliability.
Retiring duplicate infrastructure. The most defensible saving, because it's a line you can point at. If consolidation lets you switch off a warehouse, a set of ETL tools and a reporting stack, that's a hard cost removed. But it only materialises if you actually decommission — and decommissioning is the step most programmes skip. A consolidation programme that doesn't turn anything off has increased your costs, not reduced them.
Time-to-answer for new questions. Before: a new analytical question needs a pipeline, which needs a project, which needs a quarter. After: someone queries existing governed data in a day. This is real and large, but hard to put in a spreadsheet because nobody tracks the cost of questions that were never asked.
Model deployment cycle time. If your data scientists spend 60–70% of their time acquiring and cleaning data — the commonly reported range, and consistent with what I see — then halving that is equivalent to hiring a third more data scientists at no cost.
Audit and regulatory response. Underrated. In a regulated industry, being able to answer "who accessed this data, when, and on what basis" from one system rather than five is worth substantial money the first time a regulator asks.
AI use cases that were previously impossible. The largest potential value and the least predictable. Don't put it in the business case as a number. Put it in as an option you're buying.
A defensible business case leans on the first and third mechanisms, mentions the fourth, and treats the fifth as upside.
How to tell whether you need it
Six questions. Answer them honestly.
- Do your reporting and your models read from the same data? If no, you're paying the reconciliation tax whether you've measured it or not.
- How long does a new analytical question take, end to end? If the answer is measured in months, the constraint is architectural.
- What proportion of your data scientists' time goes to data acquisition? If it's above half, you're funding expensive people to do plumbing.
- Can you answer an access-audit question from one system? If it takes a week and three teams, that's a compliance risk with a price tag.
- Do you have unstructured data with obvious business value that nobody can use? Contracts, claims correspondence, call recordings.
- Can you name what you would switch off? If you can't, stop. You're adding a platform, not consolidating one.
Question 6 is the one that matters. It's the difference between a consolidation programme and an expansion programme, and organisations routinely believe they're doing the first while doing the second.
When it's the wrong answer
Consultants are poorly incentivised to tell you this, so:
- Your data is small. If everything fits comfortably in a conventional database, you're buying capability for a scale you don't have.
- You're purely BI, with no ML ambition. A warehouse plus a BI tool will be cheaper and simpler.
- You're deep in one cloud vendor's ecosystem and happy there. The integrated option from your existing provider may cost less in total once you count the people.
- You haven't fixed ownership. This is the important one. A platform does not create data governance. If nobody can currently say which system is authoritative for customer address, migrating to a new platform relocates the confusion at considerable expense.
That last point generalises beyond this product: tooling doesn't fix an ownership problem, it makes it more expensive.
What I'd put in front of a board
Not "should we buy Databricks." That's the wrong question and it invites a vendor-led answer.
The right question: what is our current cost of maintaining two data estates, and what would we switch off if we had one?
Cost the reconciliation. Cost the duplicate infrastructure. Cost the engineering time spent moving data between systems. Then look at the consolidation options — this one, its competitors, or your existing cloud provider's offering — and compare against that baseline.
If you can't build the baseline, you're not ready to choose a platform. And building the baseline is a four-week exercise that will improve your decision more than any vendor evaluation.
Platform licensing is only part of the bill. For the consulting and data engineering side of the number, see What AI and Data Consulting Actually Costs.
Sources: Databricks company newsroom and press releases, 2026; Sacra and press reporting on valuation and revenue run rate, June 2026. Figures accessed 2026-07-26.
Fares runs QartMina Labs, an independent backend and AI engineering practice. No vendor relationships.