Data Analysis for Beginners: Avoid the 5 Most Common Mistakes
Data analysis is one of the business tools that has changed and developed the most so far in the 21st century. The role of data analyst, barely heard of ten years ago, is now one of the professional profiles whose demand has grown most in companies of all kinds, both large and small. In fact, don't you find it strange to end a day without hearing or reading the words “big data”, “data scientist”, “machine learning” or “artificial intelligence”?
In this article we will clarify, keeping our feet on the ground, what some concepts are and are not, so that you can discover the most common mistakes in data analysis and find a solution for your company or project.
To do so, we interviewed a data-analysis expert, Cristina Campos, data scientist and science communicator. Cristina has a degree in Physics, specialising in Astrophysics, from the University of La Laguna. She received a scholarship from the Canary Islands Astrophysics Institute to study planetary nebulae and later became the IAC’s resident astrophysicist, working on NASA’s Sunrise programme. She subsequently specialised in stochastic mathematics and artificial intelligence, studying Finance at the UOC, Artificial Intelligence at Stanford University and a master’s degree in Mathematics for Financial Instruments at UAB. She worked on stock-market forecasting in the banking sector and now focuses on AI, data analysis and data visualisation at Dainso, which she co-founded.
After speaking with Cristina, you realise that she is passionate about science and communicates her knowledge enthusiastically.
Big data: where do I start to put this chaos in order?
First, we need to clarify what big data is. Put simply, it is the same data-analysis work, but big data involves an enormous quantity of data whose handling requires computing and programming resources beyond everyone’s reach because of its complexity and scale. The good news is that most companies do not have big data; at most, they have large data, which is a different level. The science behind structuring and making use of it is the same regardless of its size: data analysis.
Data analysis aimed at understanding what has already happened in the past—what we might call a descriptive model—is what companies usually want and is far from being affected by what we would call chaos. By contrast, when data analysis is used to generate predictions from past events, the predictive model is different and the possibilities of unforeseen variables increase. Nevertheless, achieving 70% or 80% reliability in a prediction already gives you a great advantage when making business decisions, much more than tossing a coin. Techniques such as machine learning and stochastic mathematics are applied, taking random movements into account; they are widely used, for example, in finance.
In the end, are we going to leave everything in the hands of machines? Where does the human ability to learn and intuit stand when deciding in the face of machine learning or predictive simulations based on data analysis?
Machine learning is precisely an imitation of how the human brain works: it imitates our neural connections to learn from what has already happened and, based on that, make predictions. Data analysis and predictive models are only tools. For example, an experienced salesperson may know which products in a catalogue will sell best, but a large catalogue means missing information. Automating data and analysis will help ensure that sales opportunities are not overlooked.
The power of data analysis is great and scientifically proven, but we non-experts can create false expectations... What are the most common mistakes companies make when handling their data?
- The first mistake is using programmes that were not designed for data analysis and that still present data rudimentarily; in addition, data collection is manual and brings with it a large number of errors.
- The second mistake is how the data is structured: not all data is relevant, and it can introduce noise depending on what we want to measure. As a result, we do not obtain the information we need.
- Third, not setting realistic objectives regarding what can be obtained from data analysis. This is especially true of predictive models, where results that are impossible are sometimes expected. Such expectations are fuelled by the belief that artificial intelligence is superhuman, which it is not at all.
- Another added difficulty is data dispersion. If no software is used that integrates all the data and is designed to keep it synchronised across all processes, as an ERP does, fragmented information works against us.
- Fifth and lastly, a lack of objectivity. If someone is very interested in finding a result through the data, they will find it. It is essential to base the selection, structuring and combination of data on objective criteria and to build a sound model, so that we come as close as possible to understanding reality. This is why it is valuable for people outside the organisation to review the data-analysis work.
Dainso—a company specialising in data analysis, which you co-founded—together with NaN-tic, has created a tool for automating dashboards. The famous dashboards that CEOs like so much. How important is it that information reaches its recipient visually?
It is extremely important for a very simple reason: in their day-to-day work, people who run a business have to focus their efforts in many directions and cannot spend the whole day looking at figures one by one. A dashboard must therefore provide the most important information at a glance as soon as it is opened, so that decisions can be made in time, before things get worse. Being able to anticipate is essential. Presenting information visually saves many hours.
Is it possible that companies are working with tools they do not really understand? AI, Data Science, Machine Learning, Virtual Twins... Could access to information on the internet and intensive crash courses be creating the illusion that we always know what we are doing, when at important moments we may not?
The financial sector, like the medical and education sectors and many others, has changed greatly from our parents’ time to today. Everything is very fast now: someone can take an intensive two-week programming or data-science seminar, but those of us who devote our university years to mathematics, engineering, physics and so on learn how to think in order to solve problems; in data analysis and predictive models there are no shortcuts—it takes time.
That concludes the interview with Cristina Campos, who was kind enough to answer our questions.