Engineer says firmographic data needs transparency to build user trust

Rohit Muthyala says confidence scores source tracking and repeatable pipelines can keep business records useful

A company’s industry, employee count and estimated revenue may look like simple database fields. Behind each number is a difficult question: Which source should be trusted?

This information is known as firmographic data. Sales teams use it to find potential customers, divide territories and decide which accounts deserve attention. It also feeds pricing systems, lead-scoring models and automated outreach.

When a company record is wrong, the error can spread through every system that uses it.

“A mislabeled industry or a shaky headcount looks small, but it spreads into targeting, pricing and models before anyone notices,” Rohit Muthyala said.

Muthyala is a principal software engineer at ZoomInfo who works on large data and machine learning platforms. His peer-reviewed research on automated industry classification has been presented at IEEE conferences, including the International Conference on Semantic Computing.

His work focuses on making company information more accurate while showing users where estimates came from and how confident the system is in them.

Why company labels become outdated

Industry classification is harder than it appears.

A company may operate in several industries. It may also change its products or business model without updating every public profile. Websites often use marketing language that does not clearly describe what the company does.

Different data providers can assign different industry codes to the same business. Older information may remain in a database long after the company has changed direction.

Muthyala worked on machine learning pipelines that classified companies using NAICS, SIC and LinkedIn industry categories. The systems examined text from company websites and other sources to predict the most likely classification.

According to Muthyala, the pipelines covered more than 18 million companies.

The model used text patterns found on company websites, but it did not treat every prediction as equally reliable. Confidence thresholds allowed the system to hold back uncertain results instead of forcing a label onto every record.

“If the model is guessing, I want the system to admit it,” Muthyala said. “Confidence is part of the product.”

That approach is especially important for industries with fewer examples. A model may perform well across the full database while still struggling with smaller categories.

Accuracy and coverage pull in different directions

A data provider wants to fill as many fields as possible. Customers do not want company records with missing employee or revenue estimates.

Filling every field, however, can create false confidence.

Employee counts may differ depending on whether a source includes contractors, international workers or subsidiaries. Revenue estimates can also vary widely for privately held companies that do not publish financial reports.

Muthyala’s approach was to combine several sources instead of trusting one feed. The system normalized the information, removed duplicate records and gave different weight to each source.

It could also reject extreme estimates that did not fit the other available information. For example, a revenue estimate could be checked against the company’s reported workforce and typical revenue per employee in its industry.

The goal was not to hide uncertainty. It was to provide useful coverage while placing limits on estimates that lacked strong evidence.

Changes should be explainable

A company’s firmographic record does not remain still. Employee counts rise and fall. Revenue changes. Businesses merge, move or enter new industries.

Large unexplained changes can cause problems for the teams using that information.

A company that appears to jump from 50 employees to 5,000 could suddenly move into a different sales territory or pricing tier. That change may be correct, but users need to know whether it came from a reliable filing, a new data provider or a model prediction.

Muthyala said systems should preserve the source, date and confidence level behind each important field. They should also limit major changes unless stronger evidence supports them.

That makes it possible to review why a value changed instead of simply replacing the old number.

Data reliability becomes an engineering problem

In a Forbes Technology Council article, Muthyala argues that important data products should be managed with practices similar to site reliability engineering.

Teams can monitor whether data is current, complete, consistent and within expected ranges. They can also set limits for how much missing or unreliable information a system can tolerate.

When those limits are exceeded, the response should go beyond sending an alert. A pipeline may need to quarantine suspicious records, return to an earlier stable version or pause a risky update.

“If you cannot replay the decision, you cannot defend it,” Muthyala said. “Reproducibility is how you earn trust at scale.”

A repeatable pipeline allows engineers to run the same process again, compare results and identify what changed. That becomes important when a vendor modifies its data or a machine learning model begins producing different results.

Resolving duplicate companies

Another challenge is entity resolution, which is the process of determining whether several records refer to the same real company.

A business may appear under a legal name, a shortened brand name and several website addresses. A subsidiary may share contact information with its parent company. Different offices may also be mistaken for separate businesses.

Muthyala discussed this problem in an IEEE Computer Society webinar about moving from noisy records to trustworthy entities.

The process begins by standardizing names, websites, addresses and phone numbers. The system can then compare records using a mix of exact and approximate matches.

Human review remains useful for difficult cases, especially when merging records could erase meaningful differences between a parent company, subsidiary and individual location.

Better models still need better records

New AI systems can classify and enrich company information faster, but larger models do not remove the need for reliable sources.

The harder problem is showing why a result should be believed.

That requires confidence scores, source tracking, repeatable processing and clear rules for handling conflicts. It also means leaving a field uncertain when the evidence does not support a firm answer.

“If you want people to trust firmographics, you have to show your work,” Muthyala said. “The label is the last step, not the first.”