Skip to main content

Insights / Articles

Meet Veridion: How ML Models Classify Companies When All You Have Is a Name

Struggling with predicting industry codes? Veridion's ML models offer a revolutionary way to classify companies using just their names.

SG
Stefan Gergely
Stefan Gergely
in 17 hours9 min read
AI & Automation in DataConcept Explainer
Key takeaways
  • 1 in 4 CRM admins trust less than half their data.
  • Classification starts with identifying the correct company.
  • ML combines multiple public signals to build continuously updated company profiles.

Every business decision starts with one basic question: Who is this company?

The answer sounds simple until you only have a company name to work with. 

Thousands of businesses share similar names, operate across multiple countries, change ownership, or expand into entirely new industries over time. 

That makes identifying and classifying companies far more difficult than matching a name in a database.

In this article, we'll look at why accurate company classification is important business data and the role of machine learning in modern company classification. 

Why Accurate Company Classification Matters

Think about the last time you searched for a company online.

You probably typed its name into Google or a business database and expected the right company to appear.

Most of the time, it does.

But what happens when all you have is the company name? 

No industry, website, or description.

For instance, just imagine how many businesses share common words like 'Prime,' 'Pioneer,' 'Summit,' or 'United.' 

They all sound familiar. 

But they don't tell you whether the business is a manufacturer, a software company, a bank, or a food supplier.

That's the problem.

A company name is just a label.

Before you can enrich a company profile, find it as a supplier, or add it to your CRM, you first need to answer one simple question.

What kind of company is this?

That's exactly what company classification solves.

It identifies what a business actually does instead of relying on how it describes itself or what someone manually entered years ago.

Now, company classification may sound simple, but in reality, it's one of the hardest problems in business data.

Companies can shut overnight, move industries, or even get acquired by competitors. 

Such events can change a company’s official name, work address, email IDs, and phone numbers.

Now, just imagine tracking this over millions of companies!

When these changes occur, your datasets decay rapidly.

According to The State of CRM Data Management in 2024, nearly one in four CRM administrators say less than half of their CRM data is accurate and complete.

24% of CRM administrators report incomplete and inaccurate data statistic

Illustration: Veridion / Data: Validity

If your CRM tags a manufacturer as a distributor or a software reseller as a software developer, every team in your company using that data starts from the wrong place.

Marketing targets the wrong audience; procurement wastes time reviewing companies that never matched the requirement in the first place

Take WeWork (flexible workspace provider) as an example.

For years, WeWork was presented as a technology company.

Its filings repeatedly (110) used the word 'technology,' and many investors viewed it alongside fast-growing software businesses.

The problem was that its core business worked very differently.

WeWork signed long-term leases on office buildings, redesigned those spaces, and rented them out to businesses on short-term memberships. 

Technology supported the experience, but it wasn't the product. 

That distinction mattered.

In a 2019 analysis, Vijay Govindarajan and Anup Srivastava argued that WeWork lacked the characteristics that define modern technology companies. 

Unlike software platforms, it couldn't scale with minimal costs, required significant capital investment, lacked strong network effects, and generated revenue through a real estate model rather than a digital platform.

Classifying a business wrongly changes how investors understand its risks, growth potential, and value.

The same idea applies far beyond financial markets.

Let's look at something much closer to everyday business.

In 2019, an investigation by The Wall Street Journal found millions of fake business listings on Google Maps.

Fake business listings and addresses on Google Maps

Some businesses appeared under names they didn't own.

Others used misleading identities or fake locations to attract customers.

People searching for legitimate businesses often ended up contacting entirely different companies. 

Trust suffered because users couldn't always tell who they were really dealing with.

Ethan Russell, Director of Product at Google Maps, puts it plainly.

Russell quote

Illustration: Veridion / Quote: Phone Arena

When company-related inaccuracies take over, search becomes less useful, and discovery becomes harder. 

Not to mention, a direct hit on revenue too. 

The same State of CRM Data Management in 2024 mentioned above highlighted that 31% of CRM administrators said low-quality CRM data costs their company at least 20% of annual revenue! 

Validity statistic

Illustration: Veridion / Data: Validity

Even more worrying, 41% said poor data forced their organization to delay or completely stop valuable business initiatives during the previous year. 

Validity statistic

Illustration: Veridion / Data: Validity

These numbers aren't only about company names, duplicate records, or missing phone numbers.

They reflect a much bigger issue.

When businesses can't trust their data, they struggle to trust the decisions built on top of it.

And classification plays a big role in that.

After all, how can you identify the right prospects if you don't know what industry they belong to?

In other words, it wouldn't be wrong to say that company classification is the basis of all other classification. 

If done correctly, it can help you find the right suppliers and prospects, analyze markets with confidence, and make smarter business decisions.

The challenge is that doing this accurately for millions of companies, using nothing more than a company name, isn't something traditional databases were built for. 

That's where machine learning changes the whole equation. 

How Veridion Classifies Companies With Just a Name

What can a machine really learn from just a company name?

Suppose all you know is 'Prime Technologies' in Canada. 

Search that name online, and you'll probably find dozens of businesses with similar names. 

Some may be software companies, others manufacturers, and some may not even be operating anymore.

That is where company classification actually begins. 

Before a model can determine what a business does, it must first answer a much harder question: Which company are we talking about?

Veridion approaches this as a data resolution problem before it becomes a classification problem.

Veridion dashboard

Source: Veridion

Here’s how:

Veridion First Identifies the Right Company Behind the Name

A company name is rarely unique enough to identify a business. 

Commercial names are reused across countries, and sometimes acquisitions leave behind historical names that continue appearing across websites and registries for years.

Instead of assuming the name is enough, Veridion's Match & Enrich API starts by treating it as only one signal among many.

Veridion dashboard

Source: Veridion

The minimum input is simply a company name and country, but the API can also accept an address, website, registry identifier, or phone number whenever those details are available. 

Each input is first standardized into a common format before matching begins:

  • Country names are normalized
  • Websites are reduced to their root domains
  • Registry numbers are converted into jurisdiction-specific formats, and 
  • Phone numbers are standardized regardless of how they were originally written 

Now, not every signal carries the same weight. 

A verified company website is usually a stronger identifier than a business name. 

Likewise, an official registry number provides much stronger evidence than a social media profile. 

The matching engine continuously weighs these signals against one another rather than relying on a single attribute.

That flexibility becomes particularly useful when the input itself contains mistakes. 

A company name may be misspelled, an address incomplete, or a website accidentally associated with the wrong business. 

Instead of rejecting the request outright, the system evaluates every available signal together.

Sometimes there simply isn't enough evidence to make one definitive decision. 

Rather than forcing a potentially incorrect match, Veridion returns multiple ranked candidates together with an explainability layer showing how closely each company matches the original input. 

Every resolved entity also carries a confidence score so downstream systems know exactly how much trust to place in the result before enrichment even begins.

ML Models Analyze Business Signals Beyond the Name

After finding the right company, the next challenge is understanding what that company actually does.

If you relied only on official registries, the answer would often be incomplete because companies can shift their business models long before those changes appear in government records.

Veridion therefore builds company profiles by combining signals from across the public web instead of depending on any single source.

Veridion dashboard

Source: Veridion

Its crawling infrastructure continuously collects information from company websites, registries, filings, press releases, news, maps, and social platforms. 

Together, these sources process well over a billion web pages every month. 

Language models extract only business-relevant facts such as locations, offerings, technologies, contact details, ownership information, and operational activities.

Veridion further compares evidence across multiple sources before deciding which attribute becomes part of the final company profile.

Every attribute also keeps its own provenance. 

Alongside the value itself, the platform records where that information came from, when it was last verified, and how confident the system is in its accuracy. 

Consider the case of a Brazilian credit bureau that wanted to identify businesses operating outside official registries. 

Using digital signals from websites, maps, and social platforms, Veridion discovered 15,000 active businesses in São Paulo that had no registry record but could still be verified through their public digital presence. 

Every business was matched to a physical location, enriched with standardized business information, and confirmed to have no legal registry entry before being added to the bureau's addressable universe.

Veridion Maps Business Activities Into Structured Categories

Traditional company classification usually ends with an industry code like NAICS or SIC. 

While useful for reporting, those codes often say very little about how a company actually operates.

Veridion therefore treats industry codes as only one layer of a much richer company profile. 

It goes further to add six proprietary layers on top: 

  1. Business Model (how the company generates revenue, like SaaS or marketplace)
  2. Target Markets (which industries it actually serves)
  3. Core Offerings (what it delivers, described in plain operational terms rather than a category label)
  4. Technology Focus
  5. Certification Focus, and 
  6. Supply Chain Focus

This creates a profile that is considerably more descriptive than a single industry label. 

A commercial insurance underwriter, for example, can identify businesses based on their actual operations and risk profile instead of relying on broad industry classifications.

Equally important, every attribute follows a consistent schema. 

Instead of producing different free-text descriptions each time a model runs, the classification engine returns structured outputs that downstream applications can reliably consume through APIs, warehouses, batch files, or streaming pipelines without additional interpretation.

Classification Improves Through Large-Scale Data and Continuous Updates

Veridion's graph is built to catch changes rather than wait for a scheduled refresh. 

Core profiles update continuously, fast-moving signals refresh daily, and technographic data runs on a rolling 90-day window.

When a source updates, for instance a registry filing showing a company moved its headquarters, the system doesn't just patch one field. It re-derives the record from every available source to keep the full profile consistent. 

The scale behind this is what makes it hold up across a database of more than 80 million companies in 240-plus countries

  • 507 million legal entities analyzed
  • 186 million validated digital entities, and
  • Around 472,000 company records updated in just the last 30 days

Veridion’s objective is not to place businesses into predefined boxes. 

It is to continuously build, verify, and refine a living company profile that evolves as the business itself evolves. 

That is ultimately what allows classification to remain useful long after the first match has been made. 

Conclusion

Company classification is much more than assigning an industry label. 

That is why modern classification models focus on evidence instead of assumptions. 

By combining entity resolution, large-scale data collection, structured machine learning, and continuous validation, they can build company profiles that stay relevant as businesses evolve. 

As company data continues to grow in both volume and complexity, accurate classification will become a foundational capability for every organization that depends on business intelligence.

Articles

Discuss how these trends affect your organization.

Our analysts are available for a short call. Bring a specific question and we will ground it in the data.