The 89% Accurate Model That Was Almost Useless
What a machine learning assignment taught me about the questions every business should ask before trusting AI
After more than 40 years in business, I didn't expect a college assignment to change how I think about decision-making. But that's what happened this month.
As part of my Masters in Advanced Digital Technologies for Business, I had to analyse a real dataset from a Portuguese bank. The bank had phoned 4,521 existing customers to offer them a term deposit, and it recorded whether each person said yes. My job was to build a model that could predict, in advance, which customers were likely to say yes, so the bank could focus its calls on the right people.
It's the same question many Irish SMEs could be asking every week: which customers are likely to reorder, which quotes are worth following up, which accounts are about to go quiet. And most of them already have what they need to start answering it.
Every business is now a data business
Most owners I work with don't think of themselves as running a data business. But every day, often without noticing, they're collecting it.
Every invoice records who bought what, when, and for how much. Every quote records what was offered and whether it was won. The till or card machine knows your busiest hours to the minute. Your booking system knows who comes back and who never returns. Job sheets record how long each job really took, compared with how long it was priced for. Even your email inbox holds a record of which customers need chasing for payment.
A garage knows which customers come back for their next service and which disappear after one visit. A wholesaler can see which accounts have ordered less for three months in a row. A café's till data shows which days it's overstaffed. A trades business has years of quotes that show which kinds of job it wins, and which it only ever loses on price.
None of this requires new software or new systems. It's already sitting in your accounts package, your spreadsheets and your filing cabinets.
The real question isn't whether you have data. It's whether you're using it. For most small businesses, the honest answer is: only to send invoices and file tax returns. The same records that tell you what happened could also help you see what's likely to happen next: which customers are drifting away, where your margins are leaking, and where to focus your time.
That's what the bank in my assignment was trying to do. Its records weren't exotic: customer ages, jobs, account balances, and notes on past phone calls. It's the kind of information almost any business keeps. The question was whether those everyday records could tell the bank who to call.
The model that looked brilliant
I tested three different models. One of them stood out immediately: it was right 89% of the time. On paper, that's an excellent result. If a software vendor showed you that figure, you'd probably be impressed.
Then I looked closer.
Of the 4,521 customers, only 521 said yes. The other 4,000, about 88.5%, said no. So a "model" that did nothing clever at all, and simply predicted "no" for every customer, would be right 88.5% of the time.
My impressive model was barely better than that. When I checked what it actually did, it had correctly identified just 75 of the 521 customers who said yes. It was scoring highly mainly by saying no to almost everyone. Its accuracy was real, but it was accuracy at the wrong thing.
The model that looked worse but was far more useful
Another model had a lower headline score: 83% accuracy. By the usual standard, it came last.
But it found 213 of the 521 customers who said yes, nearly three times as many. The price was more wasted calls: it also flagged 472 people who went on to say no.
So which is better? That depends entirely on the business.
Here's a simple illustration. Suppose each call costs €5 in staff time, and each new customer is worth €500 to the business. These figures are made up for the example, but the logic holds:
- The "89% accurate" model flags 120 customers. That's €600 in calls, winning 75 customers worth €37,500.
- The "83% accurate" model flags 685 customers. That's about €3,400 in calls, winning 213 customers worth €106,500.
The lower-scoring model makes the business nearly three times as much money. If calls were expensive and customers low-value, the answer could flip. The right model depends on what a wrong answer costs you, and no accuracy figure can tell you that.
The other surprise: a column that looked too good to be true
One piece of information in the data made every model look much better: the length of the phone call. Long calls meant more yeses.
The problem is that you only know how long a call lasted after it has ended, and by then you already know the answer. A model built on that information would look superb in testing and be useless in real life. Data scientists call this "data leakage". I'd call it the business equivalent of predicting last week's sales.
The question to ask of any prediction is simple: would I actually have this information at the moment I need to make the decision?
Most of the work wasn't the AI
The last lesson was one I recognise from years of mentoring. The modelling itself took a handful of lines of code. Most of my time went on understanding the data first.
The file had no blank cells, so at first glance it looked clean. But thousands of entries said "unknown". One column used "−1" to mean "never contacted before", which could easily be mistaken for a real number. None of this was obvious until I went looking.
Every business database I've ever seen has quirks like these, including the everyday records I described earlier: duplicate customers, product names typed three different ways, dates in whatever format someone used that year. The quality of any AI tool you use will never be better than the data you feed it.
Five questions to ask before trusting any AI prediction
Whether you're evaluating new software, a vendor's pitch, or a tool your team has built, these questions will protect you:
- Accurate at what? A high accuracy figure can hide a model that misses exactly the cases you care about.
- What does a wrong answer cost us? A missed opportunity and a wasted effort rarely cost the same. The best model depends on which mistake hurts more.
- How often does the thing we're predicting actually happen? If it's rare, as most valuable things are, be especially sceptical of headline accuracy.
- Will we have this information when we need it? If the model relies on something only known after the event, it won't work in practice.
- How good is our data? Look for the "unknowns", placeholders and duplicates before you trust any result.
The real takeaway
Machine learning is now accessible to businesses of every size. The tools are free, and the techniques are well established. But the most important decisions in this project weren't technical. They were business judgements: what we were trying to achieve, what a mistake would cost, and whether the data could be trusted.
That's good news for SMEs. You're already collecting the data. You don't need to be a data scientist to use it well. You do need to ask the right questions.
If you're wondering what your own business data could tell you, whether that's sales records, customer lists or job logs, I'd be happy to talk it through. You can find me at dlconsultancy.ie.
The data used in this article is the publicly available Bank Marketing dataset (Moro, Laureano and Cortez, 2011), from the UCI Machine Learning Repository.