We spent years inside large enterprises before we built MoleculeOne, and here's the thing nobody wants to admit out loud: the technology was never the bottleneck. It was the data underneath.
Imagine, you walk into a Fortune 500 procurement function with a shiny new analytics tool, an expensive spend cube, a dashboard that looked great in the demo. And then you try to answer a simple question like "how much do we spend with this vendor globally" and the whole thing would fall apart.
Why the spend file is quietly lying to you
Here's why. The same supplier shows up as "ABC AG," "ABC Corp," "ABC INDUSTRY INC," and "ABC (UK) Ltd" across four ERPs, six currencies, and three languages. One entry has a typo. Another was created after an acquisition and nobody ever merged the vendor master records. A regional team in Brazil registered the same supplier under a local tax ID that looks nothing like the one used in Germany. Multiply that by thousands of vendors and dozens of subsidiaries, and you get a spend file that looks precise down to the decimal point and is quietly lying to you.
This isn't a new problem. Companies have thrown everything at it. Armies of analysts doing manual cleanup in spreadsheets, quarter after quarter, only to have new dirty data pour in the moment they finish. Rule-based matching tools that work fine until a vendor name has one extra space or a country abbreviation nobody coded for. Outsourced data teams that clean the file once and hand back a report that's stale by the time anyone opens it. None of it holds. The problem isn't effort. It's that fragmented spend data doesn't follow rules, it follows the messy reality of how large organisations actually operate, with M&A, regional autonomy, and decades of legacy systems all bolted together.
What entity resolution actually does
This is where entity resolution actually earns its keep, and it's worth explaining plainly because the term gets thrown around loosely. Entity resolution is the process of figuring out when different records, written differently, actually refer to the same real-world thing. Not by matching strings character for character, but by understanding context: addresses, tax IDs, transaction patterns, language variants, industry codes, the works. A human expert could eventually work this out for any single case. The problem was always scale. No team of analysts can hold that kind of judgment across millions of line items.
Native AI changes the math here, not because it's smarter in some abstract sense, but because it can apply that same contextual judgment consistently, at volume, without getting tired or inconsistent on record 40,000. Older rule based tools needed someone to anticipate every variation in advance. AI models built for this actually learn the patterns of how entities fragment and reassemble across a real dataset.
340 vendors that were really 190
Picture a mid size industrial company with operations in twelve countries. Before cleanup, their spend file shows 340 distinct "vendors" in a category that, once resolved, turns out to be 190 actual suppliers. One supplier alone was split across 14 different entries because of currency formatting differences and a 2019 acquisition that never got reconciled. The company had no idea it was buying the same raw material from the same supplier at three different price points across three regions, because the data made it look like three different vendors entirely.
Once that gets cleaned up, the conversation changes completely. Procurement teams can actually see total spend with a given supplier and walk into a negotiation with real numbers instead of a guess. Sourcing teams can spot consolidation opportunities that were invisible before. Finance can trust the numbers going into a board deck instead of caveating every slide with "directionally correct." None of this requires better strategy or smarter people. It requires the data to actually reflect reality first.
We built MoleculeOne because we got tired of watching good teams get blamed for bad decisions that were really just downstream of bad data nobody had the tools to fix properly. This isn't about replacing procurement or finance judgment. It's about finally giving that judgment something solid to stand on.
Curious how many "different" suppliers in your own systems are actually the same one wearing four different names.
Frequently asked questions
Deepak Chander is Co-Founder of MoleculeOne.ai, an AI-native procurement consultancy that trains and builds alongside procurement and finance teams turning messy spend data into decisions they can trust.