Friday, September 30, 2016

Algorithms, Big Data And Accountability

                                       Credit: www.ug.ru
Everywhere you look, you’ll find them. In trouble with the law, and facing possible jail time? She is an ever-present fixture in the judge’s chambers, and you can bet we’ll hear plenty from her. Want to obtain a car loan? You can see him sitting on the loan officer’s desk, waiting to offer his thoughts. What about when you apply for a job? Quite often, she is the ultimate gatekeeper, parsing through every single resume, and deciding who will or won’t get a call back.

We live in an era where large sets of data (often referred to as Big Data), and the formulas used to analyze this information coherently (algorithms), are used to make highly impactful decisions, affecting almost every facet of our lives. The role of these tools in our lives will continue to grow.

However, sometimes these algorithms are of highly questionable accuracy, and use flawed or incomplete data, to reach decisions which affect millions of lives, at times causing considerable harm. What’s more, despite their outsized roles in our lives, many of these algorithms remain secret, with their inner workings unknown to everyone except for their creators.

Clearly, this situation cries out for greater transparency, and more effective regulation. This can be accomplished in the form of a federal agency, similar to the FCC, which would police the implementation of those specific algorithms which can have a significant negative impact on a large swath of the American public.

ProPublica recently published a detailed piece on how algorithms known as risk assessments, are used to determine a criminal defendant’s risk of reoffending. The results of these calculations are then factored into decisions concerning bail, sentencing, parole and more. ProPublica's study of more than 7,000 individuals arrested in Broward County, Florida, uncovered some rather troubling flaws in the risk assessment model developed by Northpointe,a private software company whose secret algorithms are utilized in jurisdictions throughout the nation.

Northpointe’s algorithms are based on questionnaires answered by defendants, which probes a range of data, including life prior to arrest, by asking questions about everything ranging from the arrest records and criminal history of a defendant’s family and friends, academic track record, personality traits, and drug and alcohol usage.

These models have proven quite unreliable in predicting whether someone who was arrested, will actually commit a violent crime in the future. In a two year period following an arrest (the same benchmark used by Northpointe’s creators in designing their software), just 20% of those who were thought likely to commit a violent crime, actually did so (the recidivism rate for all crimes, at a 61% accuracy rate, was “somewhat more accurate than a coin flip”).

Under Northpointe’s model, African-Americans were almost twice as likely as whites to be wrongly labeled as likely to re-offend (that is, Northpointe’s model incorrectly predicted future criminal behavior for African-American defendants, at twice the rate of whites). Additionally, whites were considerably more likely than African-Americans to be labeled as at low risk of recidivism, and yet end up in prison again (i.e. Northpointe’s algorithm majorly underestimated recidivism risk amongst whites) .

ProPublica’s researchers wondered if this racial gap might be due to other factors, including prior criminal history, age and gender. However, the disparities stubbornly persist, even when controlling for those variables. While Northpointe disagrees with ProPublica’s findings; since the actual algorithm remains secret, there’s little way to publicly debate and assess it’s functionality.

In some jurisdictions, Northpointe’s algorithms allow judges to decide whether prisoners should be granted pretrial release, or directed towards a rehabilitation program. In other places, like La Crosse County, Wisconsin, judges have utilized risk assessment scores to identify certain individuals as being at high risk for reoffending, and handed down longer sentences as a result (even Tim Brennan, the statistician who cofounded Northpointe, has expressed opposition to using this software for sentencing ).

The use of secretive, “black box” algorithms in the criminal justice system also poses major constitutional concerns. Under the Due Process clause of the Constitution, before depriving a person of his or her rights, the government must apply processes that are, at their core, fair and non-arbitrary. How does one challenge the (often inaccurate) statistical recommendations of a computerized formula, when we can’t even know how it’s findings were in fact derived? Secretive algorithms pose a unique problem to protection of this core right.

Credit scores and data, specifically those generated by FICO, and the three major credit bureaus (Experian, Equifax, and TransUnion), raise substantial concerns of fairness and transparency. FICO scores, which are the most widely used measure of creditworthiness by lenders, employers, and others, are made up of a mix of one’s payment history, debt load, the type of credit used, and other data found in credit reports.

While we have a basic idea of which ingredients go into a FICO score, as the Fair Isaac Corporation (the parent company of FICO) acknowledges: “The importance of any one factor in your credit score calculation depends on the overall information in your credit report.....therefore, it’s impossible to measure the exact impact of a single factor in how your credit score is calculated, without looking at your entire report.” Thus, FICO (as well as the VantageScore product, created by the three credit bureaus), offers little real clarity, in terms of allowing an individual to figure out why a credit score falls exactly where it does. Changes in various data points on a credit report, can affect FICO scores in a somewhat unpredictable manner.

When one considers the very substantial impact of a credit score, this opacity and secrecy is especially troubling. FICO numbers typically play a major role in determining the interest rate one pays on a mortgage or car loan, as well as whether one is able to rent an apartment, and, in some states, be hired for a new job.

While the Fair Credit Reporting Act provides methods for customers to challenge and remove inaccurate information from credit reports (which is crucial, considering an FTC study found 25% of all credit reports contain at least one material error), there’s very little that any individual customer can do to challenge the secretive verdict handed down by FICO. If there is some sort of fundamental flaw in how FICO assesses credit, such that FICO scores are a partially inaccurate gauge of creditworthiness, there is no way to know, because just like Northpointe’s software, this pivotal algorithm remains secret, hidden from outside scrutiny.

Securing gainful employment is one of the most critical (and sometimes challenging) aspects of our lives. Here too, secret algorithms play an increasingly prominent (and rather troubling) role. With around 72% of all resumes never initially reviewed by human eyes, but rather through computer programs, applicants who are skilled in sprinkling buzz phrases and keywords throughout their resume, are often favored in hiring. Job-matching algorithms which assess the likelihood of employee retention and success, can be further biased against those who are poor, as Xerox discovered with a now-defunct program they used for evaluating applicants, based on the likelihood of an employee quitting his or her job .

Applicants are also often asked to take computerized personality and cognitive tests (again, based on private algorithms and data sets), which offer questionable predictive value of employee performance, but can be used to illegally exclude those with disabilities, or individuals whose evaluations fall outside of some desired bandwidth (litigation on the legality of these practices is ongoing). With such tests being used to evaluate 60 to 70% of job applicants in the United States, these assessments have a large impact on hiring practices.

What conclusions can we draw from all of this? Are algorithms inherently unfair and prejudiced? Not quite. However, the process surrounding the implementation of these high-impact tools, clearly requires some significant changes.

Morris Hardt, a research scientist at Google, has detailed several sources of algorithmic unfairness. First, he notes, machine learning (that is, algorithms which behave intelligently, and learn from the data they are provided with), will typically reflect the patterns found in data; that is, if there is a “social bias” against any group of people, the algorithm is likely to pick up on and mimic such a pattern (Hardt cites the work of Solon Barocas and Andrew Selbst, who found that algorithms can  “inherit the prejudices of prior decision makers...in other cases, data may simply reflect the biases that persist in society at large.”). Thus, the inherent unfairness of the criminal justice system, or the employee hiring process, towards certain individuals or groups, will be reflected in algorithms which address these arenas.   

Hardt also points out that in assessing data, samples of data concerning underrepresented or disadvantaged groups, are by necessity smaller, and thus less representative, than for the general population. After all, if the premise of big data is that more data can improve predictive value, then less data often result in weaker predictions.

Beyond these issues with data, there is another major issue surrounding the use of algorithms: transparency. Since the mechanics of so many algorithms with a large public impact remain completely secret, we often don’t know whether they are working fairly, or properly.

Are we truly confident that Propublica’s findings regarding the flaws in Northpointe’s recidivism predictions are some sort of anomaly, rather than the norm, in the world of criminal justice algorithms? How certain are we that those formulas which assess the personalities of job applicants, are a fair and accurate reflection of whether a company should consider hiring someone? Do FICO scores, and the (often erroneous) data on which they are based, provide a reasonable snapshot of a borrower’s likelihood of repaying a loan? And if so, could they be made even better?

Fortunately, there are several concrete steps we can take, to overcome this problem. First, we need to carefully define which sorts of algorithms we should be most concerned about. After all, Big Data, and it’s associated algorithms, are used for undertakings ranging from cancer treatment, trading by hedge funds, threat assessments by the US military, and countless other applications in so many fields.

So how do we decide which algorithms ought to be subject to greater scrutiny? Cathy O’Neill, a mathematician and data scientist who recently published an acclaimed book warning of the dangers of algorithms and Big Data, offers a three part test to answer this question. First, is an algorithm high impact, that is, does it affect a large number of people, and carry major consequences for their lives? (Those pertaining to jobs and criminal justice are two examples O’Neil cites). Second, is it algorithm opaque; that is, people who are assessed by these formulas, don’t actually know how their scores are computed (all the examples we have considered thus far meet this criteria). Third, is an algorithm in fact destructive, that is, can it have a major negative impact on a person’s life (again, the aforementioned issues all seem to fit this test)?

What specific steps can we take to limit the potential for algorithms and Big Data to inflict harm? We need to develop rigorous due process and appeals procedures. One promising solution, which was recently implemented by the European Union (taking effect in 2018), requires that any decision based “solely on automated processing” which includes “legal effects” or “similarly significantly affects” an individual, be subject to “suitable safeguards,” including an opportunity to obtain an explanation of an algorithmic decision, and to challenge such decisions.

Here in the United States, comparable legislation, ideally passed at the federal level, is greatly needed. Such a law would first apply O’Neill’s test, to determine whether an algorithmic process warrants greater scrutiny. If it does, then a regulatory body, much like the Federal Communications Commission (FCC), ought to be tasked with providing oversight. Let’s call it the Algorithmic And Data Implementation Commission (AADIC).

How might the AADIC fulfill this mission? Just as with the FCC, a group of commissioners, appointed by the president, and confirmed by Congress, would play a primary role in offering policy guidelines for algorithmic processes generally (analogous to what the FCC did in formulating “net neutrality” rules), and helping determine whether a particular algorithm produces decisions that are fair, accurate and representative. Ideally, at least some of these commissioners would have backgrounds (both academic and commercial)  in fields like data science, statistics, and more generally, the collection and processing of large data sets.

In deliberating on and reaching such decisions, AADIC commissioners (and the public) will be provided with both the underlying formulas, as well as a sample of the data utilized by these algorithms. Commissioners would solicit public comment on the algorithms and data, from both those in support of and opposed to, a particular sort of decisionmaking (similar to amicus briefs to the Supreme Court). Of course, it is crucial for the AADIC to also recieve trusted, impartial advice. Towards this end, the commission would retain it’s own staff of experts, who could assess the effectiveness and overall performance of any data and algorithm sets.

If a majority of commissioners decides that a particular use of algorithms was somehow flawed or problematic, they can veto it’s use for public purposes, and send it back to it’s creators for further improvement and revision. Of course, just as with any government agency, the AADIC requires checks on its’ powers. Just like the FCC, decisions of the AADIC will be challengeable in federal court.

In an era where trust in the federal government is weaker than ever, many will be understandably skeptical of expanding the federal government’s regulatory authority, into yet another sphere.  I too am wary of the gargantuan bureaucracy we find in Washington DC, and certainly don’t see government as a panacea for all the challenges we face. Also, neither the AADIC, nor any other governmental body, to become a crippling roadblock for progress and innovation.

With that said, in this instance, the state must play a prominent role. The scope of algorithmic and data-based decisions in our lives continues to grow unabated, and is in dire need of some rigorous safeguards. While states, public interest groups, and private citizens can all play a positive role here, the authority of the federal government is key to offering the neccessary degree of coordination, oversight, and enforcement, to facilitate fairness, and reduce abuses. In this sense, some algorithms are no different from prescription drugs or securities.

The myriad new possibilities opened up by advancements in Big Data, and algorithmic processes, is nothing short of incredible. From insurance to healthcare to law to transportation, and so many other fields, these tools are remaking entire industries, and bringing an unprecedented degree of insight, efficiency, and cost reduction to our lives. Yet, as we now know, these tools can also be used in a harmful manner, and we must guard against such abuses. The AADIC is a decisive step in that direction.

  




















  ,  

Wednesday, August 10, 2016

The Power Of The Idea List

                                         Credit: Presentation Magazine 
In February 2014, I read a book which changed my life.
I am a longtime fan of James Altucher, a writer, entrepreneur, and former hedge fund manager, who has written several books, and hosted a popular blog, as well as a podcast, where he interviews a number of influential figures, from many different walks of life.
Through these various mediums, Altucher shares his insights from a life filled with more than it’s share of incredible highs and excruciating lows, including multiple occasions in which he sold a company he founded, earning millions of dollars, only to subsequently found himself dead broke, and at times mentally and spiritually shattered.
In Choose Yourself, Altucher explains how he altered this destructive cycle, transforming his life, which allowed him to gain tremendous personal happiness, along with creative, professional and personal freedom. He also draws upon the insights and experiences of others, to most effectively illustrate these points.
While there is much useful advice to take away from this book, I learned something which has stuck with me for every day since I read this book: The importance and impact of idea generation, as a key for finding fulfillment and success in today’s world, and more specifically, one’s personal life.
Altucher argues that we are living in a time of great disruption, as technological change, along with increased productivity, means that companies hire fewer people (and retain them for a shorter period of time), while ideas and individual creativity matters increasingly more often; i.e. there is:
“more disruption in employment, but also greater efficiencies and more opportunities for unique ideas to generate real wealth. You can develop those ideas, execute on them, and choose yourself for success.” (emphasis mine).
As Altucher explains it, we each possess a mental “idea muscle”, that is, a capacity for generating ideas, which will allow us to manifest the sort of creativity that will empower us to succeed in the “choose yourself” era; that is, to become an “idea machine.” However, this muscle, must be exercised and strengthened consistently, to function in it’s optimal state.
How does one ensure that the idea muscle becomes, and subsequently remains, as powerful as possible? Altucher offers a comprehensive workout regimen.
As with so much life advice, however, there are unique portions which work well for each of us. I varied his program to suit my own learning style and personal needs.
In his FAQ on becoming an idea machine, Altucher advocates generating 10 ideas per day, usually around a particular theme. For example, in the past, Altucher has put together idea lists around themes like “10 businesses I can start” or “10 books I can write’ or come up with 10 ways that AirBNB could be improved.
Altucher argues that such intensive idea generation, around a particular theme, pushes us to expand and strengthen our idea machine, or, as Altucher puts it “make my brain literally sweat.” Since the first few ideas are relatively easy to come up with, and subsequent ideas somewhat harder, you will find your brain is vigorously challenged.
As Altucher sees it, after 6 to 12 months of engaging in vigorous ideation, we are able to formulate effective solutions to the situations and problems that we (and those around us, including our family and friends) face, turning one into a “fountain of giving.” This ultimately allows us to contribute more, becoming a “mutant superhero” who lives a life that is “worth remembering forever.”
Not all of these ideas will be good, Altucher notes (most probably won’t stand out as anything special), which is why we must generate a large amount of them.
In forming positive new habits, I find it most effective to start off modestly, but strive for great consistency. Unless I am up against an imminent deadline at work or school, I typically do best when I work on intellectually rigorous tasks, only 5 to 6 days per week, rather than each and every day. Attempting to do something every single day, often creates a feeling that reaching a goal was impossible, which ultimately leads to fatigue and discouragement.
I also decided to write down 2 ideas per day, rather than 10. Why? Simply put, a lack of time. When I began to work on improving my skills around idea generation, I was also waking up early in the mornings to write, learn basic coding (Java, and then Python), read, and meditate, all while working full time. Eventually, a running program was added as well.
I could either attempt to unwaveringly follow Altucher’s exact prescription, and quite possibly fail, or work towards something which could bring real improvement to my life. I chose the latter approach.
How does one go about generating lots of novel ideas? Altucher offers a range of approaches. He argues in favor of regularly skimming chapters from books on several different topics (exposing one to a range of ideas), and actively combining together various existing ideas to form new ones (after all, many ideas are simply the “idea children” of various other concepts, brought together, as Altucher notes was the case in the founding of Google).
One can also engage in activities which activate a different, less frequently used part of our brains (in one instance, Altucher accomplished this by attending a watercolor painting class). Interestingly, Altucher even sees merit in surfing the Internet, and simply being exposed to a range of ideas, which can “plant seeds” for future creativity.
From my end, I would first look for problems in my own life. That is, I would consider something that I found at least somewhat inefficient or challenging, and write down a solution for that particular problem. For example, a few years ago, I was receiving a distractingly large volume of (often irrelevant) work and personal emails; why couldn’t my phone vibrate or ring only when it was an email of importance, rather than, say, Netflix or Amazon? (This was an idea from July 19, 2014).
In other instances, I tried combining concepts I knew of from various areas/disciplines (that is, birthing the “idea children”, whom Altucher wrote about). One instance: Why don’t we utilize shared/crowdsourced knowledge (as found in Wikipedia), to improving people’s spending habits? Perhaps we could arrange for people with similar income levels, anonymously post their monthly spending and saving habits, and thus learn from each other? (This concept dates to June 16, 2015).
Browsing through a range of books and websites, and thus gaining exposure to a variety of ideas, also proved quite helpful in expanding my creative output. Additionally, after I formed a consistent habit of reading 20 minutes fiction, and 20 minutes nonfiction, 4–5 days per week, I saw a spike in the ease and quality of my idea generation.
Case in point: Recently, I read a few articles about discovering and classifying fossils, as well as several pieces around the mechanics of machine learning, artificial intelligence, and big data. I paused and wondered: Perhaps we could apply these techniques to improved classification of fossils? (I conjured that one up on August 2, 2016).
What have I learned, and ultimately gained, from this entire experience?
First off, I now understand that generating new ideas, let alone quality ones, is not easy, especially at first. While some of the concepts I came up with above are at least halfway decent, each useful idea was sandwiched between lots of half-baked, minimally useful musings, which was often the best that I was able to produce on a particular day.
However, as Altucher explains, this is neither unusual, nor in any way a negative outcome. Very few people can realistically be expected to produce mostly exceptional ideas (by definition, most ideas can’t be unusually brilliant). We must generate a substantial volume of material, much of it mediocre, in order to produce a smaller subset of useful material. I’ve come to accept this as a part of the process, and feel neither regret nor shame.
What’s more, as Altucher predicted, the effort which one applies towards becoming an idea machine, helps us become more solution-oriented, in many aspects of our lives. Within just a few months, whenever I ran into challenges, at work or elsewhere, I was certain that some sort of resolution was close at hand. When a client requested that we renegotiate something which had already been agreed upon, or I found myself gaining weaker results from my exercise and diet plans, than I had hoped for, I didn’t panic, or subject myself to additional stress.
Instead, I utilized a calm, methodical approach. First, I asked myself what outcome was ultimately desired, and secondly, what stood in the way of obtaining those results. Lastly, I forced myself to name multiple approaches which might allow me to get to that point. This forced me to break down and actually understand the matter at hand, and in doing so, solutions became increasingly apparent. This methodical but creative approach, was a direct product of my daily idea generation.
I also began to feel both more positive, as well as increasingly confident, in my personal life. As I came up with more ideas, my identity began to shift. I increasingly thought of myself as a capable, creative person, who had solutions, and could bring something of real value to the world. I became more comfortable in speaking out, and offering advice to those who requested it, whether at work or in their personal lives. I adopted the mantra “It will work out. We will find a way.” Simply put, I believe in myself more, and worried less.
When I look back on that day in 2014 when I first read the chapter(s) in Choose Yourself which covered idea generation, I feel an immense sense of gratitude. My decision to internalize Altucher’s advice, on my own terms, and work towards becoming a conscious producer of ideas, has markedly changed my life for the better. I unhesitatingly encourage each of you, to implement your own version of this program, and watch how you transform as a person.

Monday, July 11, 2016

Correlation And Causation In The Age Of Data

I recently came across some unusual, rather interesting information which I’d like to share with you.
Did you know that the marriage rate in Kentucky is closely correlated with the number of people who fall out of fishing boats and drown each year? Or, that the number of letters in the Scripps National Spelling Bee, closely tracks how many people were killed by venomous spiders each year? Now, what if I told you that as the number of Facebook users grew, Greek sovereign debt spiked? How about if we noted that as box office receipts for M. Night Shyamalan movies dropped, so did newspaper sales? Or, on a somewhat morbid note, suppose you learned that swimming pool drowning deaths in a given year, were directly correlated with the number of movies in which Nicolas Cage appeared during that time. Pretty interesting stuff, no?
Now, do we agree that Facebook’s expansion caused the Greek debt crisis? Or, that the increased number of drowning deaths in a particular year, was thanks to Nicolas Cage’s movie appearances? You’ll stop me right there, I’m sure. Does this really makes sense? What relationship is there between the two variables, in any of the scenarios mentioned above? How can we say that the occurrence of one event, actually caused the other to come about?
While some of these statements might sound exceptionally silly, they underscore an important point: Correlation does not necessarily mean causation. Just because two variables somehow move in tandem, doesn’t mean that the occurrence of one variable, in fact brought about a change in the other.
That is, Shyamalan’s less acclaimed movies aren’t killing off the newspaper business. Spelling bees aren’t causing a rash of spider attacks. This might sound like a statement of the obvious, but, in today’s data-rich environment, this basic principle remains as important as ever.
With each passing year, we have access to more data than ever before. According to 2013 findings from Norwegian research firm SINTEF, 90% of all the data available in the world at the time, was created from 2011 to 2013.Other studies found the amount of available digital data is doubling every two years (faster than even Moore’s Law), while some observers believe we’ll see a 4300% increase in annual data generation, by 2020.
This exponential growth in data is driven in part by the proliferation of cellphones, tablets and other electronic devices, but also connected devices(components of the “Internet of Things”, basically, a range of Internet-capable devices that transmit data, including wearable devices, sensors, medical equipment, machine components, and more), as well as increasingly powerful computing tools for analyzing large volumes of complex data (i.e. Big Data). In short, we will enjoy access to more information, about more aspects of life, than ever before.
In making sense of this new knowledge, we are likely to discover countless new correlations, between seemingly disparate pieces of data. Yet, how we interpret this information, remains as critical as ever. In a 2015 piece in Information Age, Ben Rossi cites the example of Google Flu Trends and Google Dengue Trends. Rossi notes that while these tools are intended to detect the spread of the flu and other illnesses, by monitoring increases in Google searches around illness-related phrases, it is well known that people sometimes mindlessly Google a variety of words.
This can inflate the search frequency, and supposed incidence, of these diseases. It is yet another example of confusing correlation and causation, with potentially serious consequences. After all, what if public health agencies began incorrectly directing resources towards one illness, and away from another, more serious disease, based on a misunderstanding caused by Google search trends?.
Writing in technology journal The New Atlantis, statistician Nick Barrowman expands upon some of the challenges posed by the rise of massive sets of data, combined with increasingly powerful computing systems. Thanks to these two developments, correlations may well be “mass produced” such that “many of them will be meaningless.” Barrowman observes that some experts, have heralded the rise of Big Data as eliminating the need for any real understanding of causation, for giving any real thought as to why things happen.
Specifically, Barrowman critiques the work of author and former Wired magazine editor Chris Anderson, who in a widely read 2008 piece argued that: “This is a world where massive amounts of data and applied mathematics replace every other tool that might be brought to bear….who knows why people do what they do? The point is they do it, and we can track and measure it with unprecedented fidelity….correlation supersedes causation, and science can advance even without coherent models, unified theories, or really any mechanistic explanation at all.”
Barrowman acknowledges that correlation data holds value, but argues that without truly considering counterfactual situations, we can’t really confirm whether one event actually caused another. That is, if we believe that A caused B, we can’t truly confirm such causation, unless we ask ourselves what would happen if A had not occurred.
For example, suppose one argues that thanks to a new diet, your neighbor lost 30 pounds. What if this neighbor had never changed his or her diet? What if he or she actually took up jogging, and that was a catalyst behind the weight loss? Without digging further and considering alternatives, we might miscategorize or oversimplify the causes of a particular phenomenon. Barrowman ultimately pushes for a rigorous experimental approach, along with tools like randomization, in order to truly understand whether or not one event actually caused another.
This approach might sound a bit abstract and theoretical, until we consider how oversimplifying correlation and causation, can lead to misdirected public policy. University of Chicago economist Steven Levitt,and his coauthor Stephen Dubner, offered an instructive example of this in Freakonomics. In 2004, then-governor of Illinois, Rod Blagojevich, set forward a plan to mail one book per month, to the home of each newborn child, until that child turned five years old. Blagojevich crafted this initiative, which would cost around $26 million per year, in response to a study which found that children from homes where books were present, earned higher reading test scores.
Of course, as Levitt and Dubner wonder, were these children doing better in school, simply because books were present in the home, or because they were raised in families which valued education, and where intellectual pursuits, including reading, were emphasized by parents? Based on the empirical data before us, the latter is more likely true.
Yet, Blagojevich was prepared to direct tens of millions of dollars per year (this legislation ultimately wasn’t adopted), in pursuit of what was a classic correlation-causation misunderstanding. It isn’t tough to imagine similar misunderstandings, whether in business, public policy, or other crucial arenas, resulting in the adoption of otherwise poorly directed schemes and solutions.
To a certain extent, we are victims of our own brains, that is, our cognitive biases. In his book Thinking Fast And Slow, psychologist and behavioral economist Daniel Kahneman reviews a fascinating study from the late Belgian psychologist Albert Michotte, demonstrating that even infants as young as six months of age, quickly form ideas of causation from seemingly associated visual stimuli, and are caught by surprise when this pattern is disrupted. As Kahneman puts it: “We are evidently ready from birth to have impressions of causality, which do not depend on reasoning about patterns of causation.” In a bid to make sense of our world, we are predisposed, from a very early time in life, to draw conclusions of causation between seemingly associated events.
Kahneman also refers to an instructive anecdote in Nassim Taleb’s bestseller The Black Swan, which further illustrates the human need to find order, through developing narratives of causation. On the day when Saddam Hussein was captured by American forces in Iraq, Bloomberg News displayed a headline stating that US Treasury prices (which fluctuate throughout the day), had risen (which is generally associated with a lower tolerance for risk), because Hussein’s capture might not prevent terrorism. Later in the day, when Treasury prices fell (indicating greater investor receptiveness to risk), Bloomberg posted a new headline, indicating that this was because Hussein’s detention increased the attractiveness of riskier assets.
Rather perplexingly, the same cause (the capture of Saddam Hussein) was used to explain two seemingly opposing events (that is, a rise, followed by a fall, in Treasury prices), which occurred in a short time span. What could explain this seemingly irrational, contradictory approach? As Taleb explains, human beings are predisposed to making sense of information in terms of narratives, and to finding causality in any confusing situation, even when doing so might appear, at second glance, less than prudent.
So, what’s the solution to all of this? In today’s world, vast amounts of data are compiled on virtually every aspect of our existence, and parsed at an unprecedented pace and scale. This offers myriad information and vast potential, in fields ranging from public health and treatment of cancer, to finance, law, urban planning, retail sales, and a million other areas. With each passing day, we are gaining greater insight into actual human existence and behavior.
Not surprisingly, correlations can be spotted just about everywhere. Given this reality, how do we avoid inaccurate findings of causation, which can result in flawed decisions? Simply put, we need to be skeptical. When we find a correlation between two events, or sets of data, we should actively search for alternate explanations, rather than simply accepting causation as a fact. We must also engage in the sort of counterfactual thinking that Barrowman advocated, with focus on experimentation and empirical data. In short, we ought to keep our eyes open, and be wary of drawing sweeping conclusions, purely on the basis of correlations and association.
There must be a conscious effort to avoid assuming that we know too much, and steer clear of falling in love with the power of seductive computer-generated correlations. Big Data is powerful, but it must be applied prudently. With this approach, we can make prudent use of the opportunities which new streams of data present us with, while avoiding our inherent mental blind spots.