Evolution is an unusual word to describe the advancement of the Data Scientist. After all, the definition of evolution is as follows. “The way in which living things change and develop over millions of years”. I’m not claiming that the Homo Habilis could code. But, what we can see is that there is an evolution in the methods, process, and technology used by a Data Scientist.
Many would contest the true beginnings of statistical modeling. But fewer would argue what that evolutionary lifecycle looks like.
The technology and methods have changed. Consider the industrial revolutions of the 19th and 20th centuries. The Renaissance. Ever since the dawn of humankind, we’ve sought to leverage science to improve the world around us.
The beginnings of the Data Science Evolution
Data Science in the form we know it today has only been around since the new millennium. When we saw some statisticians separate themselves from traditional mathematicians and computer scientists.
Data Science in its purest form started as statistics in 800 AD with Iraqi mathematician Al Kindi. He used his own method of statistical analysis for cryptography. Also known as code breaking. His work is the first recognised example of frequency analysis and led the way for others.
Data records grew
During the 1300s Florentine banker, Giovanni Villani began using data. He used extensive records and knowledge of Florence. Using data he could build a comprehensive guide to the city. Data such as population, geography, trade, and education. This is the first documented use of statistics for philanthropic ends.
In the 17th century, John Graunt and William Petty created the first life table after studying the population of London. They were able to calculate that the population of London was somewhere around 384,000. Using only the rates of mortality of London as a marker. They also found the average family size in London during the 17th Century was 8.
These are extraordinarily accurate figures. Despite there being a census in place, there was the mobility of groups in and out of the major cities almost every day. Back then, many residents did not have one fixed abode.
Into the 20th Century…
In the 20th century, statistics became a recognised and prominent field. Used to help quantify the increasingly diverse societies of the 1900s. Karl Pearson and Francis Galton were two revered mathematicians. They studied societal diversity – height, weight, race, hair colour and more.
Galton contributed his knowledge of deviation, correlation and regression analysis. While Pearson pioneered the ‘Pearson product-moment correlation coefficient’ and the ‘Pearson distribution’. They became key in helping to measure a degree of linear dependency. Ronald Fisher built on this research. He wrote the textbooks that defined the academic discipline of statistics.
His most famous work is the 1918 paper, “The Correlation between Relatives on the Supposition of Mendelian Inheritance.” It became one of the cornerstones of statistical academic research. It’s referenced at universities all over the world. Fisher also divided opinion with his work, “The Genetical Theory of Natural Selection.” This looked to prove evolutionary theory using statistics.
Computer Integration
Marvin Minsky and Arthur Samuel pioneered statistics and computer integration. The two men who are arguably the forefathers of Machine Learning. Minsky created the first randomly wired neural network, codenamed SNARC in 1951. While in 1949 Samuel designed a self-learning checkers program. The program worked on a commercial IBM 700 computer. From here on in, the tide began to change. Computers sharing the driving seat with humans in the advancement of statistical analysis. And so, the evolution of the data scientist continued.
The modern day Data Scientist
Fast-forward to the modern day… the profile of the Data Scientist looks incredibly different. One of the Data Scientists who best represents the modern day landscape is Andrew Ng. Andrew is the Chief Data Scientist at Baidu and a Stanford Professor. He is a pioneer of Deep Learning, one of the newest advancements in the world of Machine Learning.
During his time at the head of Google Brain, he and his team developed some of the most intricate and complex deep neural networks in the world. He has taken this research to Baidu. He is helping design their Minwa AI platform. This specialises in image recognition, powered by trained Deep Learning algorithms. His research will help to power many of the visually featured artificial intelligence you will see around the world.
Google is working on many deep learning projects
Another example of the modern day Data Scientist is Demis Hassabis. Demis is a former chess prodigy and neuroscientist, who is head of Google DeepMind, a British AI firm. The program they’ve produced tackles vintage video games and traditional board games. It does so with no human input, again using Deep Learning. It has recently beaten the European champion of the Chinese board game ‘Go’, and (as of date of print!) is currently beating the world champion.
Finally, Sebastian Thrun is a former head of Google X… and was the pioneering mind behind the Google Self-Driving car project. This has seen Google’s fleet drive over 1 million autonomous miles around California. He also worked on implementing probabilistic techniques into robotics. This has since appeared in commercial products such as robot vacuum cleaners. These three leaders are at the top of their respective fields. They are the future of Data Science, despite having an average age of just 42.
A growing jobs market
The evolution of the data scientist spans the field of statistics of over 1200 years. Despite the term only existing since the turn of this century! It is also heralded as ‘The Sexiest Job of the 21st Century’. Which understandably, has created a queue of applicants stretched around the block. Data Scientists have evolved because advancement is the difference between surviving or dying. As technology becomes better, so does it’s autonomy.
The market is currently awash with predictions and concerns about how far autonomy and self-learning robots will go. The only thing that is for certain is that it’s going to be very exciting to observe.
How will the Data Scientist of tomorrow look?
Will the advancement of technology and self-learning programs, mean that Data Scientists will replace themselves?
For the Data Scientist, that would be the cruelest ironic twist of fate.