Popular Topics
Popular Topics
Showing posts with label data-driven vs knowledge-driven. Show all posts
Showing posts with label data-driven vs knowledge-driven. Show all posts
Thursday, November 22, 2012
A Real World Case Study: Business Rule vs Predictive Model
The following is a true story to complement earlier posts Comparison of Business Rules and Predictive Models and Predictive Modeling vs Intuitive Business Rules .
A few years ago, we built a new customer acquisition model for a cell phone service provider based on its historical application and payment data. The model calculated a risk score for each cell phone service applicant using information found in his/her credit reports. The higher the score, the higher the risk that a customer will not pay his/her bill.
A few weeks after the model was running, we received an angry email from the client company manager. In the email, the manager gave a list of applicants who had several bankruptcies. According to the manager, they should be high risk customers. However, our model gave them average risk scores. He questioned the validity of the model.
We mentioned that the model score was based on 20 or so variables, not bankruptcies alone. We also analyzed people with bankruptcies in the data that we used to build the model. We found that they paid bills on time. It might be that people with bankruptcies are more mobile and thus depend more on cell phones for communication. They may not be good candidates for mortgage. But from cell phone service providers' perspective, they are good customers.
This is the bottom line. Data-driven predictive models are more trustworthy than intuition-driven business rules.
Tuesday, September 04, 2012
Predictive Modeling vs Intuitive Business Rules
Many organizations have analysts who create business rules to identify most profitable customers, detect credit card or medical claim frauds, etc. Analysts in a company may maintain hundreds or even thousands of rules and add new ones regularly. Those rules are usually derived from human intuition and experience. The following are two imaginary credit card fraud detection rules. (They are for illustration purpose only, not actual rules used by any banks).
1. High frequency rule: If the number of transactions from a card in the past 7 days is above 75, then current transaction is fraud.
2. High dollar amount: If the amount of transactions from a card in the past 7 days is above $95,000, then current transaction is fraud.
We can plot credit card transactions on a plane as shown below. X-axis is the number of transactions in last 7 days. Y-axis is the total dollar amount of transactions in the last 7 days. Based on the two variables, every transaction is represented by a dot in the plane. In this plot, red dots represent fraud transactions and blue represent good. Fraud transactions normally have higher frequency and cumulative amount since fraudsters spend money more aggressively. Thus most of the red dots are located in the up right corner (high frequency and high cumulative amount).
1. High frequency rule: If the number of transactions from a card in the past 7 days is above 75, then current transaction is fraud.
2. High dollar amount: If the amount of transactions from a card in the past 7 days is above $95,000, then current transaction is fraud.
We can plot credit card transactions on a plane as shown below. X-axis is the number of transactions in last 7 days. Y-axis is the total dollar amount of transactions in the last 7 days. Based on the two variables, every transaction is represented by a dot in the plane. In this plot, red dots represent fraud transactions and blue represent good. Fraud transactions normally have higher frequency and cumulative amount since fraudsters spend money more aggressively. Thus most of the red dots are located in the up right corner (high frequency and high cumulative amount).
A high frequency rule detects 2 red dots (yellow rectangle) and a high dollar amount rule detection 2 red dots (blue rectangle). Totally, we detect 4 red dots. If we want to detect more red dots, we have to lower the thresholds for those rules which will lead to many normal transactions being mistakenly marked as fraud.
The most effective way to separate red and blue dots is the straight line that runs northwest-southeast direction as shown in the figure. Unfortunately, the line can not be described by intuitive rules. What a statistical predictive model does is to find such a line through learning from the data. A statistical predictive model can not be described intuitively but it is far more accurate. It is often to see that a single predictive model outperforms hundreds or thousands of intuitive business rules combined.
There is a newer post about the comparison of predictive models and intuitive rules from 13 aspects.
Subscribe to:
Posts (Atom)
