The Data Debt Comes Due
Vishal Sachar
Co-Founder & CEO of CLRT
Every organisation runs on data it has quietly agreed not to trust. The revenue figure that lags a week. The CRM stage that means whatever the last salesperson wanted it to mean. The spreadsheet everyone calls temporary that has been the real ledger for six years. None of this appears in the architecture diagram, and none of it has ever caused a visible failure, because the people who use the data every day route around the bad parts without thinking. Then an agent joins. An agent is the first user in the history of your company that reads everything, believes everything, and acts on everything at machine speed. It treats your data with a literalism no human survivor of your organisation would ever apply. That is the moment two decades of quiet data debt stops accruing and comes due.
Data debt has stayed invisible for a structural reason, not an accidental one. Every organisation's data is worse than its dashboards admit, and nothing breaks, because the people who use that data carry a private map of its lies. The operations manager knows the status field has meant nothing since the migration. The analyst knows to cross-check billing before quoting a revenue number. The salesperson knows which pipeline stages are real and which are theatre. None of this is written down anywhere a system could read it, because it never needed to be. It is folklore, transmitted at handovers and over lunch, and it works. A field can be wrong for years and cost nothing, provided every reader knows it is lying. The interest on the debt was being paid all along, in small silent acts of human judgment that no ledger ever recorded.
The agent breaks this equilibrium twice over. First, it is a literalist. It has no tenure, no folklore, and no lunch. Where a veteran reads an active status and silently discounts it, the agent reads an active status and acts on it, because the record is the only reality it has. Second, it operates at scale. A human misreading a stale field produces one error, usually caught by the next human in the chain. An agent reading the same field inside an automated workflow turns the error into a policy, applied uniformly to every record it touches, at a speed no review cycle matches. And when the output arrives confidently wrong, the post-mortem will blame the model, because the model is the new and visible component. Look closer and the model reasoned faithfully. The ground it reasoned over was false.
The numbers say this is the main event, not an edge case. In February 2025, Gartner predicted that through 2026, organisations will abandon 60 percent of AI projects unsupported by AI-ready data, and the scope clause carries the meaning: the projects that die are the ones whose ground was never made ready. The same survey of 1,203 data management leaders found that 63 percent either lack the right data management practices for AI or are unsure they have them. That is a separate finding from Gartner's earlier call, in July 2024, that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025; the two predictions have different scopes and should not be merged. What they share is the direction of blame. The industry keeps auditing the engine. The failures keep rooting in the ground.
The debt also carries a second ledger, exposure, because the first user that reads everything will read things nobody intended to publish. IBM's 2025 Cost of a Data Breach study, vendor research and best read directionally, found that 97 percent of breached organisations that had suffered an AI-related security incident lacked proper AI access controls, and 63 percent had no AI governance policy at all. Heavy use of unsanctioned shadow AI added roughly 670,000 dollars to the average breach, against a global average cost of 4.44 million dollars. Set that beside a stranger fact from Salesforce's own survey of some 6,000 professionals: more than half say they do not trust the data used to train today's AI. Organisations, in other words, distrust the ground and are wiring agents into it anyway, with the reading list ungoverned and the doors unlocked.
The instinctive response is a cleaning programme, an enterprise project to make the whole estate trustworthy, and it is the wrong response. Most of the estate is never read by any workflow that matters; cleaning it is cost without return, and it rots again the moment the project ends. AI-ready is not a property of a warehouse. It is a property of a workflow. The work that pays is triage: choose the workflow, list every field it consumes, trace each one to where it is born and where it rots, then decide, field by field, whether to fix the source, fence the agent away from it, or verify it at the point of use. That decision needs the folklore, the map that lives in your veterans' heads, and extracting it is a read on the business, not a data engineering ticket.
A field can be wrong for years and cost nothing, provided every reader knows it is lying. The agent is the first reader that does not.
A deeper dive
The second-order traps are where in-house efforts founder, and they are worth naming precisely. The first is aiming at the pile instead of the workflow. Data quality programmes scoped to the whole estate fail on their own economics, which is one reason the debt survived every previous modernisation: the cost of cleaning everything always exceeded the visible cost of cleaning nothing. Workflow-first triage inverts the economics, because a revenue-bearing workflow rarely depends on more than a few dozen fields, and making those few trustworthy is tractable. The second trap is treating contradiction as a technical defect. When the CRM and the billing system disagree about a customer, choosing which one wins is not an integration task. It is a governance decision that encodes who in the company owns which truth, and no pipeline can make it for you. The third is verifying in the wrong place. The check that matters runs at the point of consumption, on the specific fields the agent is about to act on, against a source the model does not control. Asking the model to assess its own inputs is not verification. It is the folklore problem restated, with the least experienced reader in the building as the arbiter.
There is also a discovery effect that surprises every organisation the first time. Connecting an agent is an audit you did not commission. The agent finds the forgotten export full of customer records, the share nobody remembered was open, the permissions that accreted over a decade of reorganisations, and it finds them at machine speed, which is why IBM's findings on access controls and breach costs belong in this story rather than in a separate security one: data debt and access debt compound. Underneath all of it sits the reason the demo misleads. The demo ran on clean data, often data the vendor prepared, and the pilot ran on the one system the project team understood. Production is where the folklore lives, and the folklore cannot simply be requested from the people who hold it, because its defining property is that nobody ever needed to write it down. Surfacing that map, deciding which of its lies matter for this workflow, and engineering the checks that make the ground hold is the unglamorous majority of every agent build that works.
Work with CLRT
This triage is where CLRT starts, because it is where the outcome is decided. We do not begin with a model choice. We begin with the workflow, the fields it stands on, and the folklore your veterans are silently applying, then we build the verification and the fences that let an agent read your business without being lied to. If you want to see where your own debt sits before an agent finds it for you, our diagnostic at ascent.clrtstudio.com maps exactly this, workflow by workflow. Bring us the process you want automated, and we will show you the ground it has to stand on first.

Vishal Sachar
Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.


