Skip to main content

Insights / Articles

Cracking the Black Box: Why Underwriters Need Defensible and Traceable ESG Screening

Tired of opaque ESG ratings? Build a defensible ESG screening model that provides clear, traceable insights for underwriters.

SG
Stefan Gergely
Stefan Gergely
in 3 days10 min read
ESG & SustainabilityConcept Explainer
Key takeaways
  • Review of Finance found correlations of only 0.38 to 0.71 between six major ESG raters, a pattern known as ESG rating disagreement.
  • ESG rating methodologies stay hidden, so one input can dominate a score without anyone outside the agency knowing it.
  • Traceable screening ties every score component to a citable source, which gives underwriters a reason they can defend when the decision is reviewed.

An environmental, social, and governance (ESG) score arrives as a single number, and that number carries real weight in commercial underwriting decisions. 

What it rarely carries is an explanation. 

As a result, underwriters who price risk on that basis can find themselves unable to say what the score was built from. 

This article looks at why aggregate ratings resist scrutiny, and what evidence-backed screening does differently.

Why Traditional ESG Ratings Function as a Black Box

A single ESG score compresses a large amount of analytical work into one figure, and almost all of that work stays in a “black box,” or unknown. 

What the underwriter receives is a number they can use but cannot examine.

Three characteristics of aggregate ratings explain why this is the case.

Scores Rarely Match Across Providers

Two agencies looking at the same company often reach different conclusions, and neither explains why its number differs from the other's. 

Researchers call this ESG rating disagreement, and it has been measured directly.

In fact, a study published in the Review of Finance compared the ratings of six major providers assigned to the same companies: KLD, Sustainalytics, Moody's ESG, S&P Global, Refinitiv, and MSCI. 

Correlations between them ranged from 0.38 to 0.71, signaling a substantial divergence. 

The researchers traced where the disagreements originate, with the results shown below.

Breakdown of ESG rating disagreement: 56% measurement, 38% scope, and 6% weighting

Illustration: Veridion / Source: Review of Finance

Measurement accounts for the largest share, meaning agencies assessing the same category, such as labor practices, still arrive at different conclusions about it. 

Scope is the second driver. Two agencies may cover a different set of categories, so their ESG ratings are not describing the same thing.

As an example, researchers writing in the International Journal of Finance & Economics scored the same 67 firms using data from Bloomberg and S&P Global across 2018 to 2021.

The ESG score differences are shown below.

Comparison of Bloomberg and S&P Global ESG scoring methodologies and average scores

As the figures show, Bloomberg put the average at roughly double the one from S&P on identical companies over the same period.

The study attributes the split to differing criteria and focus, though neither Bloomberg nor S&P publishes enough data for an outsider to confirm this.

For an underwriter, this creates a concrete difficulty. 

Pulling a second opinion is standard practice in risk assessment work, and yet a second ESG rating cannot serve as one. 

That’s because there is no way to tell whether a lower score reflects worse ESG performance or simply a different rating method. 

Confidence in the figure therefore rests entirely on trusting the agency that produced it.

Underlying Inputs Stay Hidden

That difficulty connects directly to a second issue, which concerns what actually goes into the scores. 

If methodologies were published in full, the differences between raters could be inspected, and an underwriter could judge which approach better suits the risk being priced. 

Unfortunately, competitive pressure works against disclosure, so weighting frameworks and input hierarchies mostly stay behind closed doors.

The scale of what stays hidden becomes clearer once you consider the range of sources that can feed a single ESG figure.

Company disclosures, filings, controversy screening, and estimates combined into one ESG score

Source: Veridion

Multiply that by every ESG score component an agency tracks before combining them in a single score, and the differences between scores start to become substantial. 

What reaches the underwriter, then, is the product of many layered judgments, none of them visible.

On top of that, weighting adds another layer. 

The weight assigned to an input determines how much it impacts the final score, and those weights differ from one rater to the next. 

Researchers behind the International Journal of Finance & Economics report reverse-engineered Bloomberg’s and S&P’s scores, and found a single input had an especially large weighting.

That input was net income, weighted at 99%.

International Journal of Finance & Economics statistic

In other words, a company’s profitability was shaping the ESG figures that were meant to describe environmental and social performance. 

For a different example, consider this Bloomberg report that examined how their competitor MSCI assigned its ratings.

Bloomberg’s team found that one ESG rating upgrade for McDonald’s illustrated an issue particularly well.

Simpson, Rathi and Kishan quote

Illustration: Veridion / Source: Bloomberg

This upgrade followed a decision to drop carbon emissions from McDonald's calculation, on the reasoning that climate change posed no material financial risk to the company. 

For an investor weighing financial exposure, this may be a defensible methodological choice. 

But an underwriter assessing environmental liability needs this exact variable, as the ESG assessment depends on it.

And without published inputs or their weighting, teams are not aware what they’re basing their decisions on.

A Single Number Can't Be Defended Later

Divergent scores and hidden inputs combine into one practical weakness, which is that the rating cannot be defended once somebody examines it. 

Underwriting decisions get reviewed regularly, whether by an internal auditor, a reinsurer assessing a treaty, or a regulator running a thematic review. 

None of those reviewers will treat an ESG rating agency's authority as the end of the discussion. 

Instead, they work backwards from the decision toward the evidence, asking some of the following questions.

Key questions for evaluating the traceability and reliability of an ESG score

Source: Veridion

Every one of those questions asks for something an aggregate score does not contain. 

This means an underwriter holding a number and nothing else has no straightforward answer, which puts the decision that relied on it in question as well.

The concern deepens once you consider that ESG scores often draw on what companies say about themselves rather than what they have been shown to do. 

Regulators, for their part, have already penalised firms over that gap between claim and substance, as the example below shows.

SEC press release on Invesco Advisers’ misleading statements about ESG investment considerations

Source: SEC

According to the U.S. Securities and Exchange Commission (SEC), Invesco had no written policy defining what their ESG integration initiative actually meant, so the promising figures they published could not be substantiated under examination. 

An underwriting file sits in the same position when the decision behind it rests on a claim nobody verified, which is why insurance data increasingly has to carry its own evidence.

Luckily, European regulators have begun addressing the problem at its source with EU Regulation 2024/3005 setting stricter requirements from ESG rating providers.

Key requirements of EU Regulation 2024/3005 on ESG rating activities

Illustration: Veridion / Data: EUR-Lex

Published methodologies and separate pillar ratings, as set out in the EUR-Lex text of the regulation, would remove much of what makes today's aggregate scores unexaminable. 

However, the rules are still recent, and they govern rating providers rather than the insurers relying on them. 

Underwriters working today therefore need their own basis for a decision, one that holds without waiting for the ratings market to reform.

What Defensible, Traceable ESG Screening Looks Like

Screening built around evidence behaves differently from screening built around an aggregate score. 

The reasoning becomes inspectable, which in turn changes what an underwriter can say about a decision months after making it. 

Three shifts describe what that looks like in practice.

Every Score Traces to a Specific Source

A defensible screening model connects the overall score, and every component inside it, back to a specific piece of evidence somebody can read directly. 

Picture a manufacturer carrying an elevated environmental risk score. 

That score may well originate in a report of a chemical spill at one of its sites, and yet knowing the report exists is not the same as knowing how it impacted the number. 

An exact traceable chain like the one below closes that distance.

Five-step ESG score traceability process from source document to score contribution

Source: Veridion

The steps above are simplified, of course, and the principle holds regardless of how a particular model is built. 

The source document is retained rather than summarised away, together with the claim drawn from it and its publication date. 

A confidence value then records how reliable that source is, which matters because a regulatory filing and an unverified news report should not carry equal weight. 

Only afterwards does the evidence contribute to the score, with the size and direction of that contribution also recorded. 

As you might expect, providers that document their methodology openly make this chain auditable from the outside.

What that means in practice is that an underwriter revisiting the file can move from the number back to the document without relying on anyone's recollection. 

Verena Ross, Chair of the European Securities and Markets Authority, the EU's securities markets supervisor, has argued along similar lines. 

Speaking at the European Financial Congress, she set out why this kind of visibility matters.

Ross quote

Illustration: Veridion / Quote: ESMA

Ross extends the point past the ratings themselves to the sustainability claims underneath them, where the risk is that investors and consumers are misled by statements nobody has tested. 

Put simply, ESG screening that records its own reasoning can be checked. 

Screening that delivers only a headline score leaves nothing to verify, which means taking the rating agency's word for it.

Company-Level Justifications Replace Guesswork

Traceability also changes what an ESG profile is able to record about a company. 

Instead of one blended figure covering everything at once, an ESG profile can hold individual findings that each stand on their own, with a source attached. 

That structure, in turn, supports a distinction between a company’s public ESG claims and real-world actions.

Comparison of public ESG claims with verified actions and regulatory evidence

Source: Veridion

Public statements like press releases, website copy, or sustainability pledges are statements of intent. 

By contrast, a verified action holds evidence of things that actually happened. 

Both belong in an ESG assessment, but treating them as equivalents is where guesswork creeps in.

The mismatch between making sustainability claims and the reality of them being fulfilled is called greenwashing

Left unchecked, it can lift a company's ESG rating well above what its behaviour actually supports.

TotalEnergies, a French multinational oil and gas producer, offers a documented case. 

If we look at the image below, the company states its ambition plainly on its own site.

TotalEnergies webpage outlining its sustainable development and carbon neutrality ambitions

Claims of this kind feed naturally into the environmental component of an ESG rating. 

But if such a claim is scored without a corresponding record of conduct, it raises the total score based on the strength of an intention alone.

A French court took a different view of similar wording, as France24 reported.

Headline reporting a French court ruling against TotalEnergies over misleading greenwashing claims

Source: France24

In October 2025, the Paris Judicial Court found three passages on TotalEnergies' French consumer website to be misleading commercial practices, ordering their removal within a month and publication of the court ruling on the site. 

For underwriters, this case is a reminder that a pledge and a verified outcome are not the same evidence, and a rating built on the former alone can misstate the risk it's meant to capture.

Traceable, dated, and sourced findings let an ESG profile hold public claims and verified conduct as separate, comparable records.

Underwriters Can Defend the Decision

Evidence-backed screening ultimately gives underwriters something they can stand behind under review. 

A decision stops being a reaction to a score and becomes a record of a specific event alongside the action taken as a result. 

Captured properly, the result is an audit-ready record with the components shown below.

Audit-ready ESG record checklist for underwriting review and documentation

Source: Veridion

Each field answers one of the questions a reviewer would otherwise raise, and it does so without anybody reconstructing the reasoning from memory. 

Most importantly, teams end up with a documented reason tied to named evidence. 

Records like these depend on ESG data that stays current, which is where Veridion fits. 

Veridion is a business data service that builds company ESG profiles from real-world signals rather than self-reported disclosures alone. 

It tracks every active company globally and uses a range of primary data sources including company websites, news sources, public filings.

Data from these sources are then classified against a detailed ESG taxonomy, shown below.

ESG taxonomy covering environmental, social, and governance risk categories

Source: Veridion

Machine learning models refresh this data weekly, which keeps each profile aligned with what a company is doing currently. 

Plus, Veridion’s coverage reaches past disclosure-based datasets into private and small and medium-sized businesses (SMBs), which is where most underwriting portfolios actually sit.

Northbridge, a Canadian commercial insurer, saw the effect on its SMB book, with the results illustrated next.

Northbridge case study showing SMB match rate rising from 15% to 60% with Veridion

Source: Veridion

In practical terms, quadrupling the match rate moved most of that SMB portfolio from guesswork into assessable data, and a fifth of policies were repriced based on the data. 

Ultimately, this level of data completeness and accuracy is what makes confident, defensible underwriting decisions possible.

Conclusion

That covers why aggregate ESG ratings tend to struggle under scrutiny, along with what changes once screening is built on traceable evidence instead.

Hopefully, you now have a clearer sense of the difference between a score you can use and one you can actually stand behind.

From here, a practical step is to take a recent underwriting decision and see how far back you can trace the ESG score that informed it.

Articles

Discuss how these trends affect your organization.

Our analysts are available for a short call. Bring a specific question and we will ground it in the data.