An uncomfortable truth about data -based decisions

Foto del autor

By Jack Ferson

An uncomfortable truth about data -based decisions

A large amount of data does not guarantee a lot of information. Therefore, data -based decisions do not have to be successful.
A few years ago, with the fashion of the digital transformation of companies, the Data-Driven concept: Entity that bases its decisions on data analysis. This led many organizations to capture large volumes. The term remains in force, with events such as Data Driven Day, where companies share success cases in artificial intelligence.

Nowthe data is just the first step; Merit is to turn them into information. Decisions should never be based on data, but on the information we extract from them.

Companies have been able to adopt this approach because data storage has been cheaper. In the 50s, the first hard drive was less than 5 MB and was rented for $ 3,200 per month. Today you can buy a 256 GB card for less than 20 euros. What have companies done? Save and save. Is that good? Not necessarily.

Storing data is cheap, but that does not mean you are storing information. Nate Silver summarizes it in the signal and noise: the volume of data has grown, but the noise signal ratio has decreased. That is to say, An increasing part of the data we store has no patterns that can be learned; They are only erratic behaviors of what we are studying.

Data projects should help make decisions. The bad thing is that the data, by themselves, do not. Data require processing to become information.

A MIT study in 2021 illustrated him showing how the artificial intelligences of the moment deduced which object contained some images. They observed a high dependence at the bottom of the image (a meadow, a road, water, etc.), so thatif they exchanged, the AI failed in their prediction more frequently. That is, what stood out was not to detect objects in images, but in identifying the background and deducting from it what object was in the foreground.

Something can happen in a company that blindly trusts data without interpreting: that your decisions data-driven are wrong. An example is how Nike based its digital marketing on data extracted from its classical commercial strategy. The result: lost new customers, not having learned to go to themand old, for having abandoned them in their new scheme. That is, their data-Driven method oriented to capture customers led them to lose sales.

Part of the solution is to anticipate what you expect to see: use your experience to propose hypotheses and value them with data. It’s about asking appropriate questions. First make a mental model of how they should influence the available factors on the metrics you are studying and then use statistics to see if the data corroborates it. Not doing so can lead you to give good spurious correlations.

A well known as an example is that ice cream sales are related to forest fires. No one occurs to prohibit ice cream sales with the intention of curbing fires because it is known that it is the high temperatures that both cause (and not one to the other). But The data does not know about these relationships: it is the analysts who have to propose them. And if they do not, their analysis can give equally absurd results as that, but in more complex systems. In such cases, it is essential that someone with knowledge of the sector verifies that these results serve for something.

Artificial intelligence is able to train the great current language models with data with a lot of noise. But you are not an artificial intelligence, but someone who seeks to make good decisions. The data has grown as never before. The information, not so much. And that is what really matters.

Deja un comentario