Takeaways
The short version
1. The scale fix did nearly all the work. The same borrowed model scored 0.49 on Polish companies using the published numbers and 0.71 once each company was measured against others in its own country. Nothing about the model changed — only how the numbers were prepared.
2. The main warning signs mean the same thing in both countries. How much a company owes, whether it turns a profit, and whether it can pay this month’s bills all point the same way in Taiwan and Poland. Measures of operating speed don’t, and are weak in both places anyway.
3. Borrowing beat building for a long time. A model with no local records at all beat one built from local records until there were about 1,600 of them — roughly 60 actual bankruptcies. If you care about the top of the alert list rather than the whole ranking, that arrives sooner, near 800.
4. The sophisticated model only paid off at home. Working inside Poland, the flexible model on 17 ratios led a simple one on five 1968 ratios by a clear margin. Sent across the border, that lead almost vanished. Whatever the extra flexibility had learned belonged to the country it was built in.
5. Half-adjusting was the worst option. Nudging the borrowed model with a small amount of Polish data lost to leaving it alone in 14 of 15 attempts. A thin sample is not a weak signal — it’s a misleading one.
If you had to act on this
- Before pointing a foreign model at new data, plot the numbers side by side. A model that looks broken is often a model being fed incomparable units.
- Rank companies against the population you’re actually scoring, rather than assuming two data sources use the same scale.
- With fewer than roughly a thousand local outcomes, use the borrowed model untouched.
- Decide what you’re optimising for before choosing when to switch. Sorting the whole book and getting the top alerts right give answers that differ by about a factor of two.
- If a model has to travel, try the simple version first. The premium for sophistication mostly did not survive the move.
What this study cannot tell you
The companies differ, not just the countries. The Taiwanese firms are stock-market listed; the Polish ones are mostly private. Country, accounting rules and company type all changed at once, so I can’t attribute the drop to nationality. This is a test of whether a model survives being moved somewhere different — not a clean comparison between two countries.
Ranking throws away the absolute figure. “This company owes more than it owns” is a meaningful line to cross. Turned into a position within a queue, it becomes an ordinary high rank. Comparability was bought with real information.
Ranking needs a crowd. To rank a company you need a population to rank it against. Fine for studying a whole portfolio; a system scoring one company at a time would need a fixed reference group, and would drift as the market changed.
One pair of countries, one family of models. The 1,600-record figure belongs to this setup. The shape of the finding will likely hold elsewhere; the number itself should not be quoted as a constant.
Two of the 17 ratios are approximate matches. Their definitions differ slightly between the two datasets — one on tax treatment, one because the Taiwanese documentation doesn’t say what it divides by. Both are flagged as such wherever they’re used.
The time periods differ. Taiwan covers 1999–2009, Poland 2000–2013. Overlapping, but not the same economic weather.
How this is checked
The column matching and the ranking step are the two places where a quiet mistake would change every result that follows. Both are covered by automated checks that run on every change:
- no source column feeds two different ratios by accident
- all five of Altman’s ratios are present, and the two approximate matches stay flagged
- ranking survives rescaling — squashing numbers into a 0-to-1 range doesn’t change anyone’s position in the queue. This is the property the entire study rests on, so there is also a deliberate opposite test: flip a column’s sign and the ranking must change. Without that, the first check could pass while testing nothing.
- ranking never looks at which companies failed, so no answers leak in
- the borrowed model stays at chance on published numbers and clearly above it after the fix — locked in so a future change can’t quietly undo the main finding
Checks needing the raw datasets skip themselves when those aren’t present, so a fresh download stays honest rather than pretending they ran.
What I’d do next
Build a version closer to real deployment: fix a reference group of companies, score new ones against it one at a time, and watch how the model decays as the market changes. That is the failure this setup can’t see.
A third country would help, but only one whose ratios are defined closely enough to make the comparison fair.
Reproducing this
Both datasets are public. They aren’t stored in this repository — a script downloads them from the university archive that publishes them. From there, one command rebuilds every table and another every chart. Nothing on this site is typed in by hand; the pages read the generated files when the site is built. Full instructions are on the front page.
- Taiwan: UCI #572, Taiwan Economic Journal, 1999–2009.
- Poland: UCI #365, Zięba, Tomczak & Tomczak (2016).
Unfamiliar with a term used here? See the terms page.