Solving the Garbage In, Garbage Out Data Crisis | Molecule One
Back to Insights
Data Foundations

Solving the garbage in, garbage out data crisis

In enterprise spend analytics the technology was never the bottleneck: the fragmented data underneath was. Here is why one supplier ends up wearing four different names, and how AI-native entity resolution finally gives procurement and finance numbers they can trust.

DC
Deepak Chander
Co-Founder, MoleculeOne.ai
July 2026 5 min read
Entity Resolution Spend Data Vendor Master Data Procurement AI Data Quality
Four different names, one supplier: Molecule One entity resolution unifying ABC AG, ABC Corp, ABC INDUSTRY INC and ABC (UK) Ltd into a single resolved entity across ERPs, currencies and regions

We spent years inside large enterprises before we built MoleculeOne, and here's the thing nobody wants to admit out loud: the technology was never the bottleneck. It was the data underneath.

Imagine, you walk into a Fortune 500 procurement function with a shiny new analytics tool, an expensive spend cube, a dashboard that looked great in the demo. And then you try to answer a simple question like "how much do we spend with this vendor globally" and the whole thing would fall apart.

Why the spend file is quietly lying to you

Here's why. The same supplier shows up as "ABC AG," "ABC Corp," "ABC INDUSTRY INC," and "ABC (UK) Ltd" across four ERPs, six currencies, and three languages. One entry has a typo. Another was created after an acquisition and nobody ever merged the vendor master records. A regional team in Brazil registered the same supplier under a local tax ID that looks nothing like the one used in Germany. Multiply that by thousands of vendors and dozens of subsidiaries, and you get a spend file that looks precise down to the decimal point and is quietly lying to you.

This isn't a new problem. Companies have thrown everything at it. Armies of analysts doing manual cleanup in spreadsheets, quarter after quarter, only to have new dirty data pour in the moment they finish. Rule-based matching tools that work fine until a vendor name has one extra space or a country abbreviation nobody coded for. Outsourced data teams that clean the file once and hand back a report that's stale by the time anyone opens it. None of it holds. The problem isn't effort. It's that fragmented spend data doesn't follow rules, it follows the messy reality of how large organisations actually operate, with M&A, regional autonomy, and decades of legacy systems all bolted together.

What entity resolution actually does

This is where entity resolution actually earns its keep, and it's worth explaining plainly because the term gets thrown around loosely. Entity resolution is the process of figuring out when different records, written differently, actually refer to the same real-world thing. Not by matching strings character for character, but by understanding context: addresses, tax IDs, transaction patterns, language variants, industry codes, the works. A human expert could eventually work this out for any single case. The problem was always scale. No team of analysts can hold that kind of judgment across millions of line items.

Native AI changes the math here, not because it's smarter in some abstract sense, but because it can apply that same contextual judgment consistently, at volume, without getting tired or inconsistent on record 40,000. Older rule based tools needed someone to anticipate every variation in advance. AI models built for this actually learn the patterns of how entities fragment and reassemble across a real dataset.

340 vendors that were really 190

Picture a mid size industrial company with operations in twelve countries. Before cleanup, their spend file shows 340 distinct "vendors" in a category that, once resolved, turns out to be 190 actual suppliers. One supplier alone was split across 14 different entries because of currency formatting differences and a 2019 acquisition that never got reconciled. The company had no idea it was buying the same raw material from the same supplier at three different price points across three regions, because the data made it look like three different vendors entirely.

340190
"vendors" resolved to real suppliers
141
fragmented entries, one supplier
3
price points for one raw material

Once that gets cleaned up, the conversation changes completely. Procurement teams can actually see total spend with a given supplier and walk into a negotiation with real numbers instead of a guess. Sourcing teams can spot consolidation opportunities that were invisible before. Finance can trust the numbers going into a board deck instead of caveating every slide with "directionally correct." None of this requires better strategy or smarter people. It requires the data to actually reflect reality first.

We built MoleculeOne because we got tired of watching good teams get blamed for bad decisions that were really just downstream of bad data nobody had the tools to fix properly. This isn't about replacing procurement or finance judgment. It's about finally giving that judgment something solid to stand on.

Curious how many "different" suppliers in your own systems are actually the same one wearing four different names.

Frequently asked questions

Entity resolution is the process of figuring out when different records, written differently, actually refer to the same real-world thing. In procurement it means recognising that "ABC AG", "ABC Corp", "ABC INDUSTRY INC" and "ABC (UK) Ltd" are the same supplier, not by matching strings character for character, but by understanding context: addresses, tax IDs, transaction patterns, language variants and industry codes. It is what lets you see true total spend with a supplier across every ERP, currency and region.
Because large organisations run on messy reality, not clean rules. The same supplier ends up spread across multiple ERPs, currencies and languages; entries carry typos; vendor master records never get merged after an acquisition; and regional teams register the same supplier under a local tax ID that looks nothing like the one used elsewhere. Multiply that by thousands of vendors and dozens of subsidiaries and a spend file that looks precise is quietly wrong.
The problem is not effort. Manual cleanup in spreadsheets is stale the moment new dirty data pours in. Rule-based matching works until a vendor name has one extra space or a country abbreviation nobody coded for. Outsourced teams clean the file once and hand back a report that ages instantly. Fragmented spend data does not follow rules, it follows how organisations actually operate, with M&A, regional autonomy and decades of legacy systems bolted together.
AI applies the same contextual judgment a human expert would, but consistently and at volume, without getting tired or inconsistent on record 40,000. Older tools needed someone to anticipate every variation in advance; models built for this learn how entities fragment and reassemble across a real dataset. Once the data is resolved, procurement can negotiate with real total-spend numbers, sourcing can spot consolidation opportunities, and finance can trust the figures going into a board deck.

Deepak Chander is Co-Founder of MoleculeOne.ai, an AI-native procurement consultancy that trains and builds alongside procurement and finance teams turning messy spend data into decisions they can trust.

Molecule One

How many of your suppliers are secretly the same one?

We resolve fragmented vendor and spend data across your ERPs, currencies and regions, so procurement and finance finally work from numbers that reflect reality. Let's find the duplicates hiding in your systems.