Weight of Evidence, Information Value and Logistic Regression for Credit Scorecards

By Hafizh Yuwan Fauzan · 2026-10-01

Part 3 of 7 in a series on building credit risk scorecards. Part 1 showed how points turn into odds; Part 2 settled what "bad" means and who belongs in the sample.

Take the applicants aged 55 and over in our development sample. They make up 13.86% of all good accounts but only 4.83% of the bad ones, so they are 2.87 times more common among goods than among bads. Take the natural log of that ratio and you get 1.053, this age group's weight of evidence. A few steps later, that number becomes the 124 points a 55-year-old applicant receives on the finished scorecard, which Part 5 builds in full.

This part is about those few steps: how raw characteristics such as age are cleaned, grouped, measured and combined into a model. It is also the part of a scorecard build that takes the longest, and I will explain why I still do most of it by hand.

All numbers here come from the same simulated portfolio of 40,000 applicants I built for my refresher on N. Siddiqi's Credit Risk Scorecards (Wiley, 2006), the book this series follows. The model is fitted on a 70% development sample of 28,000 accounts (25,187 goods and 2,813 bads), and the Gini figures come from the 30% holdout. The fine and coarse classing chart uses a smaller 2,810-account subsample, so its weights differ from the table (55+ is +1.22 there against +1.053), and the smaller sample makes the noise in narrow bands easier to see.

Clean before you group

Before any grouping, the data has to be explored, because every problem left in it ends up baked into the scorecard.

Issue Treatment options What scorecard practice does
Missing values Exclude, impute, or keep as a separate group Keep "Missing" as its own group and let its weight speak. Here, missing time at address has a weight of −0.27: riskier than average, which is information worth keeping
Outliers Cap, trim, or verify against source systems Once confirmed as genuine, grouping neutralises extreme values on its own
Correlation Correlation matrix, variable clustering Keep the strongest, most explainable member of each correlated group
Data quirks Check with operations for system changes Look for sudden shifts in distributions over the sample window

The missing-value row is worth dwelling on. Imputing a missing time at address with the average would hide the fact that applicants who leave it blank behave differently. In a scorecard, "didn't answer" is often a characteristic in its own right.

Weight of evidence: one scale for every characteristic

Weight of evidence (WOE) measures how strongly each group of a characteristic separates goods from bads:

WOE = ln( % of goods in the group / % of bads in the group )

Here it is for age:

Age Goods Bads % of goods % of bads WOE
18–24 1,585 382 6.3% 13.6% −0.769
25–29 3,522 645 14.0% 22.9% −0.495
30–34 4,377 621 17.4% 22.1% −0.239
35–44 7,738 764 30.7% 27.2% +0.123
45–54 4,473 265 17.8% 9.4% +0.634
55+ 3,492 136 13.9% 4.8% +1.053

Reading it takes three rules:

A group with zero goods or zero bads has no finite WOE, so it has to be merged with a neighbour. The real value of WOE is that it puts every characteristic, numeric or categorical, on one comparable scale. Age, residential status and number of credit inquiries can then all enter the same model in the same units.

Fine classing, coarse classing, and why age is rarely a straight line

Grouping happens in two passes. Fine classing cuts a characteristic into many narrow bands to see its raw shape. Coarse classing merges those bands until the pattern is stable and explainable.

Two bar charts of weight of evidence by age: 14 fine bands with reversals at ages 22-24 and 70-plus, merged into 6 coarse groups that rise steadily

The fine classing shows exactly why this step exists. The 22–24 band is riskier than 18–21, and the 70+ band drops back well below 65–69. Both come from small bands, so neither is a trend I would put in a scorecard without checking it first. Merging into six groups gives a steady rise from −0.81 to +1.22 that anyone in credit can explain.

In my experience, age is very often not monotonic in raw data, and that is precisely why we use WOE binning rather than putting age into a model as a straight number. A logistic regression on raw age can only draw one straight line through the relationship; grouped WOE lets each age band carry its own weight. Siddiqi makes the same point: one advantage of grouping is that "nonlinear dependencies can be modeled with linear models", and the goal is a "logical (not necessarily linear)" relationship. He adds that some reversals reflect real behaviour, and "where valid nonlinear relationships occur, they should be used if an explanation using experience or industry trends can be made" (pp. 78 and 83–84).

Research on financial decision-making points the same way. Across ten types of credit transactions, Agarwal, Driscoll, Gabaix and Laibson found that financial mistakes, such as paying higher interest rates and fees, follow a U-shape over the life cycle: they are lowest around age 53 and higher for both younger and older borrowers (The Age of Reason, Brookings Papers on Economic Activity, Fall 2009). The study measures costly mistakes, not defaults, and its turning point need not match any given portfolio. But it is a plausible reason why an older age band can turn back down for real.

So a 70+ band that dips is not automatically noise. In the 2,810-account subsample it holds only 86 accounts, about 3%, below the 5% minimum discussed below, so it merges into 55+. With enough volume, I would test whether the dip holds before merging it away.

Why I still bin by hand

Automated binning tools will happily produce groups that maximise a statistic. What they cannot do is decide whether a reversal is a data quirk or real behaviour, or whether a group boundary will make sense to the credit committee. So I review and adjust every characteristic's grouping by hand. It is the slowest part of a build, and I think it is time well spent, because these groupings become the scorecard's points table.

Information value: ranking the characteristics

Once each characteristic is grouped, information value (IV) sums up its total separating power:

IV = Σ ( % of goods − % of bads ) × WOE     (summed over the groups)

Lollipop chart of information value by characteristic: age 0.264, delinquencies 0.147, credit inquiries 0.111, residential status 0.087, utilisation 0.064, time at address 0.040, phone type 0.002

Siddiqi's rule of thumb (p. 81) reads:

IV Predictive strength
Below 0.02 Unpredictive, not used
0.02 to 0.1 Weak
0.1 to 0.3 Medium
0.3 and above Strong

Phone type, at 0.002, is dropped: it doesn't separate goods from bads, and it is easy to misstate on an application anyway.

What about IV above 0.5?

A common shortcut says an IV above 0.5 means the variable is leaking the outcome and must go. The book is more careful than that. Siddiqi writes that characteristics with an IV above 0.5 "should be checked for over-predicting—they can either be kept out of the modeling process, or used in a controlled manner" (p. 82). His own worked example of age in the same chapter has an IV of 0.668.

That matches what I see at the bureau where I work. In generic bureau scorecards predicting probability of default, variables built on days past due routinely score above 0.5. (In this simulated application portfolio, the delinquency variable is much weaker, at 0.147.) That isn't leakage when days-past-due variables are measured before the application date; in my experience, past repayment behaviour is the strongest predictor of future repayment there is. FICO, for example, weights payment history at 35% of its score, the largest of its five factor groups (as of October 2026, per myFICO). The check that matters is timing: make sure the variable could not have been influenced by the outcome you are predicting.

When a grouping is good enough

A grouping is accepted only when it makes business sense, not merely because its IV is high:

Check What to look for Example from the simulated data
Logical trend WOE moves in one direction as the characteristic's value increases, unless a reversal can be explained Utilisation WOE falls steadily from +0.30 (below 10%) to −0.87 (90% and above)
Group size Each group holds enough accounts, typically at least 5% of the sample Age groups hold 7% to 30% of accounts
Distinct groups Neighbouring groups have clearly different WOE; otherwise merge them In the full sample, the six age groups step clearly apart: −0.769, −0.495, −0.239, +0.123, +0.634, +1.053 (smallest gap 0.26)
Missing as a group Missing values get their own group when their risk differs Missing time at address: −0.27
Business fit Legal to use, reliably captured, available in future, hard to manipulate Phone type dropped: unpredictive and easy to misstate

The 5% minimum is Siddiqi's rule of thumb too ("minimum 5% in each bucket", p. 80), and I agree with it. Smaller groups produce weights that swing from one sample to the next.

Logistic regression on WOE inputs

With every characteristic coded as WOE, a logistic regression combines them into one prediction:

ln( p / (1 − p) ) = β0 + β1·WOE1 + … + βk·WOEk

p = probability of being good, so the left side is the log of good:bad odds

Fitted on the development sample:

Characteristic Coefficient (β) Standard error Wald χ² IV
Age 1.056 0.043 606 0.264
Residential status 1.070 0.072 220 0.087
Time at address 1.094 0.102 114 0.040
Delinquencies, last 12 months 1.060 0.051 441 0.147
Credit inquiries, last 6 months 1.075 0.061 308 0.111
Revolving utilisation 1.102 0.078 197 0.064

What to read from it:

Part 5 shows how these coefficients and WOE values become points, which is where the 124 points for a 55-year-old in the opening come from.

Building the model in blocks

Rather than letting a stepwise search pick characteristics, the build here adds them in blocks that mirror a credit officer's view of an applicant:

Waterfall chart of holdout Gini as blocks are added: applicant profile 36.8 percent, plus credit bureau behaviour to 45.3 percent, plus capacity and utilisation to 47.4 percent

  1. Applicant profile: age, residential status, time at address. Gini 36.8%.
  2. + Credit bureau behaviour: delinquencies and credit inquiries. Gini 45.3%, the biggest step.
  3. + Capacity: revolving utilisation. Gini 47.4%.

Gini measures how well the score ranks good and bad borrowers on a holdout sample, from 0% (random) to 100% (perfect); Part 5 covers it in detail. These figures come from the unrounded model, so they can differ by 0.1 point from the final points-based scorecard.

The block-by-block build shows what each type of information adds, and it can keep weaker but stable characteristics that a purely automated forward, backward or stepwise search might drop. Those automated methods remain valid alternatives, but the staged build is easier to explain to the people who will use the scorecard.

Next, Part 4 tackles a problem hiding in every development sample built from approved applicants: reject inference.

Over to you

Binning is where a scorecard developer's judgement shows most, and it is also where practices differ most between teams. Do you still group characteristics by hand, rely on automated binning, or use a mix of both? And where do you draw the line on information value? I would like to hear how your team handles it: get in touch.

Hafizh Yuwan Fauzan (Hafizh Fauzan) is a credit risk data scientist in Jakarta, Indonesia, building scorecards and machine learning models on national-scale credit data.