Skip to main content
Insights

So Salesforce Bought Informatica: Your Salesforce Org is Still Full of Duplicates

By September 30, 2026No Comments

Salesforce completed its $8B acquisition of Informatica in November 2025. Here’s what it changes for enterprise data management, what it doesn’t change inside your Salesforce org, and where a purpose-built Salesforce data quality tool still fits.

salesforce informatica

Key Takeaways

  1. Salesforce didn’t buy Informatica to clean up your Salesforce org. It bought a governance and context layer so Agentforce can be trusted across the enterprise by having core data management functions: catalog, lineage, metadata, multidomain MDM. Its other main strategic objective is to allow Salesforce to maintain its “metadata moat”.  Otherwise, Salesforce risks becoming a secondary spoke to Snowflake’s or Databrick’s hub. If that happens, Salesforce loses control over its customer base, which today is locked in due to huge substitution costs.

    The strategy, the positioning, the buyer, and the roadmap all point outward and upmarket, not downward into the core Salesforce data model.
  2.  A golden record somewhere else doesn’t delete the duplicate account your rep opened this morning. Enterprise MDM assembles a trusted view downstream. Operational Salesforce data quality keeps the actual records — the ones reps, service agents, and Agentforce read and write — clean at the moment of entry and on an ongoing basis. Different layer, different clock, different owner.
  3. Do capabilities overlap?   Somewhat, but that doesn’t help your Salesforce team. Informatica is a Gartner Leader in data quality. It’s also a program with a data-governance owner, a 6–18 month runway, a six-digit price tag, and a domain queue that puts ERP and supplier data ahead of your leads and cases. DataGroomr installs from AgentExchange and starts matching the same day, in real-time. The two provide a seamless data quality strategy. They are not alternatives.

Introduction

Informatica is a company whose storied history as a data-focused integration layer goes back to the founding days of the Internet.  Salesforce purchased it in November 2025 for a total price of ~$8B.  It was a strategically interesting acquisition for Salesforce driven, as we will discuss, by weaknesses that were impacting enterprises’ uptake of Agentforce.  The question we are being asked as a result of this acquisition is how DataGroomr and Informatica play together in the Salesforce “sandbox”.  Are they competitors or are they complementary?  This article will clarify what each product does and show that they are actually separate pieces of an end-to-end data quality strategy that every enterprise needs to ensure that its people, and need we say, its agents, are using only the highest quality, most accurate data to power their decisions.

What Informatica does

Informatica was a small company in Redwood City, CA when it started selling a product called PowerCenter in 1993.  Powercenter provided a visually-driven easy-to-implement ETL backbone to database-driven systems.  At the time, this was a radical departure from then current practices. Before PowerCenter, ETL layers were custom-written between data sources in COBOL, PL/SQL, shell scripts or C++.  But the advent of e-commerce drove a huge increase in cross-platform database integration, to the point that doing custom ETL layers between data sets or databases became unsupportable. This increasing pain drove the demand for any product that would allow companies to easily transform and load data from one database or data format to another. PowerCenter’s core thesis was that a visually-driven user interface could simplify the process and make it easily scalable by allowing the database engineer to visually map data structure relationships rather than writing code.  

That thesis proved immensely powerful. I was very much in this space at that time, building and integrating ecommerce databases. By the mid-90s and early 2000s, PowerCenter had become the go-to ETL layer for most of the Internet. If you worked in a Fortune 500 data-focused team  from the mid-90s until even recently, you were likely to have been using PowerCenter in at least one data warehouse.

The problem challenge for Informatica was that PowerCenter was designed before the concept of the Internet, much less the later concepts of the Cloud or SaaS, existed.  PowerCenter was completely legacy on-prem.  The growth of the Internet in the early 2000s, especially the notion of “Software as a Service” espoused by Marc Benioff,  caused that legacy to become a critical weakness.  As software-as-a-service (SaaS) systems like Salesforce grew, companies needed a way to move data to and from the Internet without exposing their secure internal databases. Informatica engineered a hybrid model that launched in 2006 as Informatica Cloud Secure Agent. The design interface and task scheduler moved to a public website accessible via standard web browsers (eliminating local client software installations). An internal piece of software called the Secure Agent sat behind the customer’s firewall to run the actual data processing.  At the same time, Informatica began to show its support for Salesforce by introducing Informatica On-Demand.  This was Informatica’s earliest internet-hosted service, primarily used to synchronize data between local databases and early Salesforce environments.

Since then, Informatica has been evolving with database technology to cloud-first platforms like AWS, Google Cloud, and Azure.  That evolution has been concurrent with the increasing sophistication with which enterprises digest and use data assets in their business.  On top of that, companies have seen an exponential growth in the volume of data, and the number of data sources they must manage.  The average enterprise went from having 5 SaaS systems they depended on in 2015 to 115 SaaS systems in 2024. Add to that unstructured data, like web scraping, application files like .csv , .xlsx, and structured data formats like JSON.  The increase in data volume, the growth in the number of data sources, and the growing variety of data types created demand for a whole new class of databases that would provide a single repository for all enterprise data: data lakes like Delta Lake and data lakehouses like Apache Iceberg, as well as SaaS-hosted platforms like Snowflake and Databricks that provide the data lakehouse  along with value-added products built on top of it.

With that came the need for tools to help manage all this data.  I tend to classify these tools under the category of “dataOps”, in the same way we now have MLOps for managing data science model development and deployment (and increasingly AI and agent deployment).  These tools include customer data platforms (CDPs), master data management systems (MDMs), data integration tools, and data governance tools, among others.  

Informatica’s response has been a growing suite of products combined under the name Intelligent Data Management Cloud (IDMC). This cloud-native rewrite, branded as a unified platform includes:

  • Cloud Data Integration (CDI) — the modern equivalent of PowerCenter, running as a managed cloud service
  • Cloud Data Quality — profiling, standardization, and rule-based cleansing
  • MDM (Master Data Management) — the canonical “single source of truth for customers/products/etc.” product, where Informatica is genuinely a category leader
  • Cloud Data Governance and Catalog — metadata management, data lineage, business glossary
  • CLAIRE — Informatica’s metadata-driven AI engine, used for automated mapping suggestions, data quality recommendations, and now LLM-style natural language interfaces
  • API & Application Integration — the iPaaS pieces, competing with MuleSoft and Boomi

The IDMC pitch is completeness: integration, quality, governance, MDM, cataloging, and APIs in one platform. Its weakness is clock speed and price.  A typical deployment timeline for IDMC is 6-18 months.  Independent estimates put the Year 1 cost to the enterprise in low- to mid-six figures.  

What Informatica brings to Salesforce today

I hypothesize that Salesforce bought Informatica for three main reasons:

Agentforce had a credibility gap on data, and Data Cloud alone couldn’t close it. Marc Benioff made an interesting comment at the closing that gives a clear rationale for the acquisition: “You have to get your data right to get your AI right… without clean, connected, trusted data there is no intelligence – only hallucination.” Data Cloud (now Data 360) is good at ingesting and harmonizing data for activation. It is not a catalog, a lineage engine, a governance framework, or a multidomain MDM hub. Those are precisely the four things a CIO or CDO demands before they will let an autonomous agent take action against enterprise data. Salesforce could not sell enterprise Agentforce without them, and building them would have taken years.

Metadata is Salesforce’s stated moat, and its metadata stopped at the CRM boundary.  Salesforce’s moat has always been a very limited set of data fields (company name, contact name, leads, actions) tied to the extensive customization most companies needed to map it to their unique processes.  Salesforce extended that moat by adding AppExchange.  As a result, not only does a company’s CRM system run off this limited data set, many other functions driven by custom workflows do as well.  Removing Salesforce thus has huge substitution costs that most enterprises are unwilling to consider.

The advent of multiple data sources, as well as their aggregation into data lakes and data lakehouses, now threatens Salesforce’s moat.  If, as a corporation, I can combine all my data in one place from which it is consumed by all applications across the enterprise in a single, standard way, then it now makes sense to move my Salesforce data into the data lake and make Salesforce a feeder into the larger system. This is the value proposition that Databricks and Snowflake are selling today. The savings I incur by moving to that new approach far outweigh the cost of restructuring Salesforce to support the new architecture.  Salesforce doesn’t go away, but it loses its premier position as the central hub from which a large number of enterprise applications, especially marketing applications, run. Salesforce needed some way to stay relevant across the broader enterprise in the world of data lakes. Informatica gives it that.

It fills the gap MuleSoft never covered and extends its application-layer lock-in. MuleSoft is an application and API integration platform. Informatica is bulk data integration, ETL/ELT, and pipeline management. With both combined, Salesforce can claim it is an end-to-end solution from the presentation layer (browser) to the logic layer (code), to the data layer for enterprise application development. As a Salesforce platform developer, you can now develop your core shared functions, then deploy them via an API layer and Mulesoft API gateway. This abstracts the functionality to make Salesforce-based software functions easily available for other developers in the organization to use. Then developers across the entire company, via Data 360 and Informatica, can leverage the comprehensive enterprise data set to develop a wide range of applications outside the CRM. Salesforce stresses the importance of this as a best practice:

“APIs remain a cornerstone of digital innovation, with 99% of organizations using them to streamline and automate business processes. And for the most part, business leaders are taking notice: 40% of respondents attribute a significant portion of their revenue to APIs, up from 33% last year. “

This expansion outside the CRM, which Salesforce has been perfecting since the early days of AppExchange, allows Salesforce to maintain its preeminent position as the “hub” of computing functionality for its customers in the era of data lakes/lakehouses. It is an extension of the same lock-in strategy Salesforce has used for years.    

Moreover, every indication is that Informatica is not going to be used to push these capabilities down into the CRM platform. Everything is upwards and outward, trying to shore up its metadata moat.  What makes me say that:

  • The focus on an MCP interface means the interface is for agents and engineers, not Salesforce admins in setup.
  • The headless approach, combined with cross-cloud neutrality, means Salesforce customers will spend their engineering effort making Informatica work well on Databricks and Snowflake, not on Salesforce object semantics.
  • The clear strategic fit of extending its lock-in that began with Mulesoft.
  • The CIO coverage of the Agentforce 360 integration noted openly that Salesforce gave no licensing or pricing details for the combined platform and no explanation of how existing Informatica customers get accommodated. That is not a product on its way to a self-serve AgentExchange listing.

What Informatica does not offer Salesforce customers is an admin-owned, same-day, in-org tool that merges the actual duplicate account a rep is looking at, prevents the next one at the point of creation, and does it across leads, cases, email messages, and custom objects. Nothing in Informatica’s roadmap is aimed there, and the strategy actively points away from it.

What Informatica doesn’t handle: the last mile

So let’s go back to the original question: is Informatica a replacement for DataGroomr or is it complementary, giving the enterprise a set of capabilities Informatica doesn’t have?  The answer is they are complementary, and here’s why:

Informatica governs the system of reference. DataGroomr governs the system of record at the moment of entry.

There are seven concrete differences that lie beneath this statement that make clear that these are complementary technologies.

1. Direction of flow. MDM and Data 360 pull data out of Salesforce, resolve identity, and produce a unified profile somewhere else. That profile does not delete the duplicate account in the org. The rep can still see two or more duplicate customer records after an upload of the thousand latest leads from the conference you just sponsored. Routing, territory assignment, entitlements, campaign membership, forecasting, and approval flows all run on the data in your Salesforce instance — not on the golden record across the enterprise. Unless something merges, updates, and writes back into the actual Salesforce records, the operational problem of duplicates that you need solved now is untouched.

2. Clock speed. Enterprise data quality runs as pipelines, jobs, and stewardship queues. Salesforce duplicates are created continuously — by reps, web-to-lead, list imports, integrations, and enrichment pushes. A reconciliation that doesn’t happen in real-time means the organization is tolerating dirty data all day.  Users are working from bad records in the interim. Real-time prevention at create and update is an essential function, not a nice-to-have.

3. Different owners and target customers. Informatica is a program owned by the CIO or CDO across an entire enterprise.  It requires extensive data modeling, source onboarding, match-rule design, stewardship workflow, a governance team, months of elapsed time. DataGroomr is an AgentExchange install that a Salesforce admin does in an afternoon. Critically — even at companies that already own Informatica, the Salesforce team usually cannot get MDM cycles for “our accounts have duplicates.” The domain queue is ERP, supplier, product, and finance first. 

4. Object coverage and Salesforce semantics. MDM masters party, customer, product, and supplier domains. Salesforce data quality problems live in leads, contacts, accounts, opportunities, cases, email messages, and custom objects. They are also grounded in  Salesforce-specific mechanics: lead conversion, record ownership, merge survivorship across related lists, activity and task reparenting, campaign member history, record types, sharing rules. An enterprise MDM hub has no concept of lead conversion.

5. The Agentforce sequencing argument. Salesforce’s own thesis is that agents need trusted context. Informatica supplies enterprise context. But an Agentforce service agent runs on the Salesforce organization’s CRM records: it opens a case, walks to the contact, walks to the account. Three contacts for the same person and it reads the wrong history, with total confidence. The potential for a Salesforce-localized, automated disaster is significant should that happen. Enterprise governance does not fix that. Real-time, record-level hygiene inside the organization’s CRM system does.

6. Enrichment is a duplicate factory. ZoomInfo and Apollo pushes create the mess they were bought to fix. That is a write-path problem inside Salesforce — orchestrate enrichment and dedupe at “the point of input”. It is not an MDM concern and never will be.

7. The overlap is real but it is the wrong shape. Informatica does profile, standardize, cleanse, match, and manage duplicates, and is a Gartner Leader for it. But for a Salesforce team, that overlap sits inside the most expensive, slowest-to-deploy, least Salesforce-aware part of the stack. The claim is not “no overlap.” The claim is: the capabilities overlap; the scope, the owner, the timeline, the cost, and the write path do not.

The result is pretty clear. Informatica’s value for Salesforce customers is to allow them to manage their entire, complex data stack across multiple platforms like Snowflake, Databricks and others –  from within Salesforce itself. Their alternative is to go the other direction – make Snowflake, Databricks, or Apache Iceberg their main system of record and data management platform with Salesforce as a spoke from that hub. It is this latter scenario that Salesforce is trying to counter. Informatica is big, it takes a huge amount of time to implement, it is expensive, and handling duplicates within the Salesforce platform is not its focus or strength.

Datagroomr fills the gap – “the last mile” – in data quality and accuracy that Informatica does not.  It ensures, in real-time, that the records your salespeople have in front of them in Salesforce – accounts, leads, cases, email messages, and custom objects – are unified in a single “best” record and are as accurate as they can be.  As a result, the two together provide an integrated, end-to-end solution for data quality for large enterprises.

Arthur Coleman

Arthur Coleman is a fractional Chief Product Officer, data scientist and builder of AI products. He is both deeply technical and a seasoned business leader who believes building a great product that uniquely fills an essential customer need requires attention to the tiniest details. With a special appreciation for elegant user interfaces, Arthur is a DataGroomr superfan.