Hiding in Plain Sight
Jan 01, 0001
Jan 01, 0001
Uncovering fraud requires a highly honed set of skills that combines both the art of knowing how to find patterns and using technology to do the heavy lifting. Consider your own credit card bills.
November 2013
By Diane Burley
Fraud and falsehood only dread examination.
Truth invites it.
– Samuel Johnson
The illusionist’s slight of hand amuses and confounds us. It is right before our eyes -- but yet we can’t see it! In fact, illusionists are masters at manipulating our attention – of having us focus intently away from where the trick is being played out. Skilled perpetrators of fraud do essentially the same thing. They know instinctively, if not cognitively, that the human brain is wired so that as it focus on one set of information – it must shut out other information.
That’s why it was so easy for Cliff U., his wife Ezinne, their partners, Princewill N. and his wife Caroline, to steal more than $5 million through healthcare fraud. The two men had founded a Houston-based home health care company that said it provided skilled nursing to Medicare beneficiaries. Caroline and her staff were responsible for recruiting Medicare patients, while Ezinne was a registered nurse who, with her staff, wrote care plans for home-bound patients.
On the surface, it all looked legitimate. And indeed, there are many such legitimate operations just like these all over the country. But in fact, the healthcare group’s patients were not homebound, nor did they need (or receive) skilled nursing services. Yet they billed CMS (Centers for Medicare and Medicaid Services) and received payment for those non-services.
Finding the Cliffs and Ezinnes of the world has been near impossible for forensic accountants and investigators, because their actions, albeit criminal, look so reasonable -- which is how, according to the FBI, CMS alone is bilked of $80 billion each year. Forensic accountants are trained to look for anomalies: Is a bill excessively high? Is the same amount being processed over and over and over again? Are the number of patients for a single practitioner realistic?
Uncovering fraud requires a highly honed set of skills that combines both the art of knowing how to find patterns and using technology to do the heavy lifting. Consider your own credit card bills. Most people assume their bills are accurate -- unless they notice an excessive amount or a purchase made from out of state or even out of country. By knowing the pattern of our own purchases and behaviors we can quickly find suspect charges. And just like a homeowner might review a monthly bill looking for aberrations, so do forensic accountants. But imagine the difficulty of finding patterns when you are dealing with billions of transactions or billions of invoices.
When individuals set out to commit fraud they know the best way to not get caught is to act like everything is normal. Highly trained individuals – even highly trained machines – can’t spot ersatz charges when transactions are perceived to be like all the other claims. And spotting the illusion – and stopping it -- has been a big problem for CMS and virtually every organization in every industry. Until now.
As the home credit card bill example shows finding fraud is predicated on understanding past behaviors. For organizations and companies this means “knowing their customers” – what is the legal entity, where is it located and how does it typically behave. Discovering fraud then is in figuring out the correlations and to do so, a data scientist play out a hunch and then sees if there is a statistical correlation between that hunch and fraud.
But testing a hunch has been an expensive and time-consuming proposition, as it requires letting data scientists search and constrain on variables. When we use relational databases, the de facto technology for the last 30-40 years, database administrators (DBAs) have to determine the columns and rows – and make sure all the data fits into those columns and rows. That’s data modeling. This one task can take several weeks if not months to accomplish. Now since fraud is finding patterns of aberrant behavior across many, many different datasets, you need to have a ton of computing horsepower to search across all these tables. So by the time you finish remodeling your database and adding all the necessary hardware and software one of two things happen: Either there was no correlation after all, so you wasted a ton of time and money – or you did find a correlation – and millions were stolen in the meantime.
Enterprise NoSQL is a new type of database that I believe is going to become a crucial weapon to combat Fraud. While there are different types of NoSQL database, the one most commonly used for being schema agnostic is a document store – allowing data to be ingested “as is.” There is no need to model data into one specific schema – and there are no tables to search across. The result is sub-second search with a fraction of the computing horsepower. The beautiful thing about all of this is, when you want to test correlations, you ingest that new information -- no matter if it is text or relational, and immediately start to test your theory.
This is precisely what we did for CMS. We took one state’s claim data of more than 600 million medical claims -- with many of those claims having more than 100 line items. We then started adding in a variety of external data sources, such as Dunn & Bradstreet Reports, National Provider Identifiers database, Facebook, multiple claims, extracts claims, DEA schedules, list of excluded individuals and entities (LEIE), diagnostic codes, indictment documents, SAS fraud algorithm reports and state policy documents. All of this data incorporated allowed us to present a picture of many aspects of a providers behaviors and history – and allowed data scientists to test different theories. We could see that Mr. U. owned several medical companies – that all resided in exactly the same address (right down to the suite number.) We then were able to cross-tabulate beneficiaries’ lists of all those companies -- and sure enough they were virtually the same.
CMS was astounded at the ease in which each new external dataset could be added. By eliminating the need for modeling data CMS was suddenly able to increase the odds in its favor to interrupt and catch fraud much earlier, while decreasing costs. The Enterprise NoSQL database allowed the focus and examination to be put on the right information – not the illusion criminals had created.
Diane Burley is Chief Content Strategist for MarkLogic.