Try this. Ask five people in your organisation what "digitally transformed" would look like. Not the vision statement — the finish line. What would be true that isn't true today?
You'll get five answers. Usually: a technology answer, a customer-experience answer, a cost answer, an organisational answer, and one person who says it's a journey and never finishes.
That's the actual state of the discipline. Bridge engineering has standards. Accounting has standards. Even project management has a body of knowledge people argue about. Digital transformation — a category that absorbs an enormous share of large-company capital expenditure — has no agreed sequence, no shared definition of done, and no milestones that mean the same thing in two different companies.
This isn't a complaint. It's a diagnosis, and it explains a set of symptoms that are usually blamed on people.
The symptoms that get misdiagnosed
Your steering committee disagrees with itself, meeting after meeting. This gets called misalignment, and someone suggests a workshop. But the CFO is measuring cost per transaction, the CTO is measuring deployment frequency, and the COO is measuring customer wait time. All three are correct. There's no shared scale on which their positions can be compared, so the disagreement is structural, not interpersonal. A workshop won't fix an ontology problem.
Nobody can say whether you're ahead or behind. Board asks "where are we versus our competitors?" and the honest answer is that nobody knows, because there's no unit. So you get proxies: a maturity assessment from a vendor whose scale conveniently peaks at their product, or a benchmark based on self-reported surveys. Both are unfalsifiable.
The programme never ends, it just gets renamed. Digital transformation becomes cloud transformation becomes data transformation becomes AI transformation. Same budget line, same people, new noun. This happens because there was never a completion criterion — and a programme without a completion criterion cannot complete, it can only be rebranded.
Direction from the top is genuinely unclear, and it's not because the board is weak. Boards are being asked to set direction on a technology whose capabilities changed materially in the last quarter and will change again in the next. Any specific instruction they give risks being obsolete before it's implemented. So they give directional instructions — "we must be AI-first" — which are unfalsifiable and therefore un-implementable. The teams below then interpret them in incompatible ways, in perfectly good faith.
That last one deserves saying plainly, because it's usually treated as a failure of leadership. It isn't. When capability changes faster than an organisation's decision cycle, vagueness at the top is a rational response to genuine uncertainty. The failure is not in the vagueness — it's in not building a mechanism to convert it into something testable further down.
Why the frameworks don't fill the gap
There is no shortage of frameworks. Every large consultancy has one, most have several, and they share a family resemblance: a maturity ladder from one to five, a set of dimensions, a radar chart.
They fail for three reasons that are worth naming precisely, because the failure is instructive.
They measure intent, not evidence. Almost every maturity assessment is built on interviews and questionnaires. You are asking an organisation to grade itself, usually through the people most invested in a favourable answer. Ask a company whether it has "a data governance capability" and it will say yes, because a policy document exists and someone owns it. Ask whether an engineer can find the system of record for customer address in under five minutes, and you get a different answer.
They have no time axis. A maturity assessment is a photograph. It tells you where you are, not which direction you're moving or how fast. For anyone deciding where to invest, direction and speed matter more than position — a company at 45 and climbing five points a quarter is in a completely different situation from one at 60 and flat, and the photograph shows the second one as healthier.
They're not comparable across companies. Every framework is applied by a different assessor with different judgement. Two companies scored "level 3" by two different firms have no established relationship to each other. Which makes the number decorative.
The result is that organisations spend serious money to receive a radar chart that confirms what everyone already believed, contains no falsifiable claim, and is quietly forgotten within a quarter.
An alternative: score what you can observe from outside
Here's the approach I use. I'm describing it in full, including its weaknesses, because the point isn't that this particular method is correct — it's that any method with explicit rules and observable inputs beats a method built on self-assessment.
The constraint I set was deliberately harsh: score only on evidence that can be observed from outside the company. No interviews, no questionnaires, no access. If a claim can't be sourced to a public artefact with a date, it doesn't count.
That constraint sounds limiting. It's the whole value. It removes the assessor's judgement about what people told them, and it makes every score reproducible by someone else.
Six dimensions, each scored 0–100, weighted:
| Dimension | Weight | Observable evidence |
|---|---|---|
| Cloud & infrastructure | 20% | Migration announcements, infrastructure roles in job postings, disclosed vendor contracts |
| Data platform | 20% | Named platforms, data engineering postings, architecture talks |
| AI in production | 25% | Named systems that are live, with a date. Not announcements |
| Engineering culture | 15% | Active technical blog, open source activity, conference talks, named technical leadership |
| Investment | 10% | Disclosed budgets, acquisitions, partnerships |
| Technical talent | 10% | Engineering headcount in absolute terms and as a share of total |
Two rules do most of the work.
A press release is not evidence of production. It counts toward investment, never toward AI-in-production. This single rule separates companies that ship from companies that announce, and it's astonishing how much it changes a ranking.
A job posting is a better source than a press release. A press release describes what a company wants you to believe; it's written by communications, cleared by legal, and the verb is almost always in the future tense. A job posting describes what the company is currently paying for. It's written by the manager who has the seat to fill. It names the real stack, the real project, the real team — and it only exists because a budget was approved.
That second rule is the most useful idea in this article, and you can apply it today without adopting anything else. Read your competitors' job postings. Read your own. Somebody is reading yours right now and inferring your actual strategy from them, not from your annual report.
What the method surfaces that a survey wouldn't
A worked example, from a panel of large European companies I score weekly.
One software company scores 87 out of 100 — the highest in the panel. It's shipping agentic AI in its own products, training its own models, and hiring principal-level agentic AI engineers. Real, verifiable execution.
The same company had 182 open positions in its home country on the day I measured, for a workforce well over 100,000. Its board had, three weeks earlier, restricted new hiring to core AI roles and created a spending council to tighten external costs.
A survey-based assessment would have scored this company highly on every dimension and stopped there. The observable method produces something more useful: a company executing extremely well and contracting hard. Those two facts together mean something specific — for its employees, for its suppliers, and for anyone reading it as a signal about where the software market is going.
That's what a measurement gives you that a maturity rating doesn't. Not a grade. A situation.
The three milestones that actually mean something
If you want completion criteria — real ones, that can be true or false — I'd propose these. They're deliberately few, and deliberately hard.
Milestone 1: You can name the system of record for your top twenty data entities, and an engineer can query each one live. Not "we have a data catalogue." A named system per entity, reachable at runtime. Most organisations that believe they've completed a data transformation fail this.
Milestone 2: A new production use case can reach the data it needs without a new integration project. This is the real test of whether a platform exists. If every use case still requires a bespoke pipeline, you have a collection of integrations, not a platform — regardless of what the architecture diagram says.
Milestone 3: You can turn something off. The most reliable indicator of transformation is decommissioning. Anyone can add a system. Organisations that have genuinely transformed have switched the old one off — and that's the step almost everyone skips, which is why the estate keeps growing and the run costs keep rising.
Three criteria. Each is binary. Each can be verified by someone who doesn't work for you. None of them mentions AI, which is the point: they're the preconditions that determine whether anything you build on top will survive contact with production.
What I'd actually tell a board
There is no standard process, and waiting for one to emerge is not a strategy — the technology is moving faster than any standards body could ratify anything.
What you can do is refuse to accept unfalsifiable statements. Every claim about progress should be checkable by someone outside the team making it. "We have improved our data maturity" is not checkable. "An engineer can retrieve a customer's current address from the system of record in under five minutes, and here's the recording" is.
That's a lower bar than a framework and a higher bar than what most transformation programmes currently clear.
Sources: company disclosures, official career portals and public technical communications, collected 2026-07-26 as part of an ongoing weekly review of 54 large companies across France, Germany and the Middle East. Scoring methodology described above is applied consistently across the panel and revised quarterly.
Fares runs QartMina Labs, an independent backend and AI engineering practice.