BAI (Before AI) “Big Data” was a bugbear and the term was used to describe the acquisition, storage and analysis of large quantities of data. The search giant Google was one of the pioneers in this area and it is developed into an industry worth billions of dollars. While Big Data is obviously still used, it is now usually presented as part of the AI debate. But whether considered on its own or as part of AI, it still raises ethical concerns.
One common use of Big Data is to analyze customer data to make predictions used to target advertising. An infamous old example of this is Target’s pregnancy targeting. This Big Data adventure was a model of inductive reasoning. First, an analysis was conducted of Target customers who had signed up for Target’s new baby registry. The purchasing history of these women was analyzed to find patterns of buying that corresponded to each stage of pregnancy. For example, pregnant women were found to often buy lots of unscented lotion at the start of the second trimester. Once the analysis revealed the buying patterns of pregnant women, Target then applied this information to the buying patterns of women customers. Oversimplifying things, they were essentially using an argument by analogy: inferring that hat women not known to be pregnant who had X,Y, and Z patterns were probably pregnant because women known to be pregnant had X,Y, and Z buying patterns. The women who were tagged as probably pregnant were then subject to targeted ads for baby products and this proved to be a winner for Target, other than some public relations issues.
One interesting aspect of this method is that it does not follow the usual model of predicting a person’s future buying behavior from their past buying behavior. An example of predicting future buying behavior based on past behavior would be predicting that I would buy Gatorade the next time I went grocery shopping because I have bought it consistently in the past. The analysis used by Target and other companies differs from this model by making inferences about the future behavior of customers based on their similarity to customers whose past buying behavior is known. For example, a store might see shifts in someone’s buying behavior that matches other data from people starting to get into fitness and thus predict the person was getting into fitness. The store might then send the person (and others like her) targeted ads featuring Gatorade coupons because their models show that such people buy more Gatorade.
This method also has an interesting Sherlock Holmes aspect to it. The fictional detective was able to use inductive logic (although it was described as deducing) to make impressive inferences from seemingly innocuousness bits of information. The use of Big Data allows for (sometimes) reliable inferences based on what appears to be irrelevant information. For example, likely voting behavior might be inferred from factors such as one’s preferred beverage.
Naturally, Big Data can be used to sell products, including politicians and ideology. It also has non-commercial applications, such as law enforcement. As such, it is hardly surprising that companies gather and analyze data at a relentless and ever growing pace. This certainly is cause for concern.
One ethical concern is that the use of Big Data can impact the outcome of elections. For example, analyzing massive amounts of data allows ads to be crafted and targeted. Given that Big Data is expensive, the data advantage would tend to go to the side with the most money, thus increasing the influence of money on the outcome of elections. Naturally, the influence of money on elections is already a moral concern. While more spending does not ensure victory, there is a clear connection between spending and success.
In any case, Big Data (and now AI) adds yet another tool and expense to political campaigning, thus making it more costly for people to run for office. This, in turn, means that those running for office will need even more money than before, thus making money an even greater factor than in the past. This, obviously enough, increases the ability of those with more money to influence the candidates and the issues. But, as a counterpoint, one could argue that the current age of AI provides Big Data AI tools for a low price and thus makes things “fairer.” As a counter to the counterpoint, one can argue that the best tools and the people who can use them well are still very expensive. But one can argue that the role of Big Data and AI in politics should be addressed by laws.
On the face of it, it would seem unreasonable to require campaigns go without Big Data. After all, it could be argued that this would be tantamount to demanding that campaigns operate in ignorance. However, the concerns about big money buying Big Data to influence elections could be addressed by campaign finance reform, which would be another ethical issue.
One major ethical concern about Big Data is privacy. First, there is the ethical worry that much of the data used in Big Data is gathered without people knowing how the data will be used or that it is even being gathered. For example, if you walk past a neighbor’s smart camera or drive by a Flock camera, data about you is being stolen without your consent and perhaps without you being aware of it. As a side issue, there is the interesting moral question about whether such systems being used to steal data about you would morally warrant your disabling them or even grabbing, for example, the solar panel used to power one, as compensation for their theft. Legally, of course, the answer is obvious—the law is generally against the people rather than protecting them.
While people might know that some information is being collected about them, knowing this and knowing that the data will be analyzed for specific purposes are two different things. As such, it can be argued that private data is obviously being gathered without proper informed consent and this is morally wrong.
The obvious solution is for data collectors to make it clear about what the data will be used for, thus allowing people to make an informed choice regarding their private information. Of course, one problem that will remain is that it is difficult to know what sort of inferences can be made from data. As such, people might think that they are not providing any meaningful private data when they are, in fact, handing over valuable information that can be exploited.
If a business claims that they would be harmed because people would not hand over such information if they knew what it would be used for, the obvious reply is that this hardly gives them the right to deceive to get what they want. However, most businesses need not worry about people deciding not to provide data. While Facebook seems to be dying under the hand of Zuckerberg, it still scoops up data and Tik Tok and Instagram are excellent data collectors.
A second moral concern is that Big Data provides a means of making inferences about private matters, such as pregnancy. While this sort of reasoning is classic induction, Big Data changes the game because of the massive amount of data and processing power available to make these inferences. In short, the analysis of seemingly innocuous data can yield inferences about information that people believe to be private—or at the very least, information they would not think would be appropriate for a company to know. People running companies generally seem that it is right and good for them to know anything they can monetize in the endless extraction quest.
One obvious counter is to argue that privacy rights are not being violated. After all, if the data used does not violate the privacy of individuals, then inferences made from this data do not violate privacy, even if the inferences are about things people think of as private (such as pregnancy). To use an analogy, if I were to spy on someone and learn from this that she was an alcoholic, then I would be violating her privacy. However, if I inferred that she is an alcoholic from publicly available information (like the Vodka bottles spilling from her recycling bin), then I might know something private about her, but I have not violated her privacy.
This counter has some appeal. After all, there is a meaningful and relevant distinction between directly getting private information by violating privacy and inferring private information using public data. To use an analogy, if I get the secret ingredient in someone’s prize secret recipe by sneaking a look at the recipe, then I have acted wrongly. However, if I infer the secret ingredient by tasting the food when I am invited to dinner, then I have not acted wrongly.
A reasonable reply to this counter is that while there is a difference between making an inference that yields private data and getting the data directly, there is also the matter of intent. It is, for example, one thing to infer the secret ingredient simply by tasting it, but it is another to arrange to get invited to dinner specifically so I can get that secret ingredient by tasting the food. To use another example, it is one thing to infer that someone is an alcoholic, but quite another to systematically gather public data to determine whether or not she is an alcoholic. In the case of Big Data, there is clearly intent to infer data that customers have not already voluntarily provided. After all, if the data had been provided, there would be no need to undertake an analysis to get the desired information. Thus, while the means do not involve a direct violation of privacy rights, they do involve an indirect violation—at least in cases in which the data is private (or at least intended to be private).
The solution, which would be difficult to implement, would involve setting restrictions on what sort of inferences can be made from data. And there is the reasonable objection that drawing inferences from data is not a violation of privacy as long as the data used was not itself a violation of privacy rights.
A Philosopher’s Blog is Now on Substack!
You can subscribe and read for free.
