Turn the flashlight off and the lights on.

The world of data analysis needs to change. Today we are living in darkness, only examining small data segments of interest as hypotheses are generated and then validated. This has worked reasonably well so far, but we are at an inflection point. Data has become too large and complex to be handled with these traditional methods. We need to find methods that permit us to view and understand data sets in their entirety. To understand the value of such methods, imagine wandering into your dark, crowded, messy garage with a flashlight in search of your skis for an upcoming trip. You point the flashlight one way and then another trying to spot your skis, and after enough time of searching, you might just find them behind the box of old clothes. If you enter the same garage looking for the same set of skis, but flip on the lights instead of using your flashlight, you will not only find the skis faster because you can examine the contents of the garage in one glance, but you will also notice your poles, ski pants, and helmet that you didn’t know were located in the garage. These items are just as useful as the skis themselves but could have been overlooked with your flashlight due to the narrow scope of your search. So the question becomes, how do we put lights in every garage so that nothing is overlooked? All my academic work has been focused on topology, the area of mathematics that studies the notion of shape. In the last ten or so years we have found that topology can also be used to discover meaning and knowledge from large and complex data sets. This is done by recognizing that most data sets are equipped with a (higher dimensional) notion of shape, to which topological methods can be applied. By applying Topological Data Analysis (TDA) to data sets ranging from everything to next generation sequencing, financial information, to sensor data, individuals can now discover the subtle nuances of the complete data set, not just the answer to a single question they may have formulated. TDA, in combination with machine learning algorithms, automatically represents data sets as a shape, in the form of a topological network, which presents many of the features associated to shape. This permits the easy interrogation of the data set, and suggests a collection of natural questions based on the shape’s theoretic features. This allows anyone who has generated data to easily analyze the data set and perform tasks such as data segmentation, and the identification of what features characterize the various segment. This new method of analysis is transforming the way individuals are looking at their data. They no longer have to walk into a dark, crowded, messy garage with a flashlight. TDA constitutes a new method of analysis that is transforming the way we understand and mine data. Domain experts and data scientists alike now have the ability to see their complex data as a connected system, easily understanding the relationships each single data point has with all other data points in even the most complex data sets. The world of data analysis needs to change. Today we are living in darkness, only examining small data segments of interest as hypotheses are generated and then validated. This has worked reasonably well so far, but we are at an inflection point. Data has become too large and complex to be handled with these traditional methods. We need to find methods that permit us to view and understand data sets in their entirety. To understand the value of such methods, imagine wandering into your dark, crowded, messy garage with a flashlight in search of your skis for an upcoming trip. You point the flashlight one way and then another trying to spot your skis, and after enough time of searching, you might just find them behind the box of old clothes. If you enter the same garage looking for the same set of skis, but flip on the lights instead of using your flashlight, you will not only find the skis faster because you can examine the contents of the garage in one glance, but you will also notice your poles, ski pants, and helmet that you didn’t know were located in the garage. These items are just as useful as the skis themselves but could have been overlooked with your flashlight due to the narrow scope of your search. So the question becomes, how do we put lights in every garage so that nothing is overlooked? All my academic work has been focused on topology, the area of mathematics that studies the notion of shape. In the last ten or so years we have found that topology can also be used to discover meaning and knowledge from large and complex data sets. This is done by recognizing that most data sets are equipped with a (higher dimensional) notion of shape, to which topological methods can be applied. By applying Topological Data Analysis (TDA) to data sets ranging from everything to next generation sequencing, financial information, to sensor data, individuals can now discover the subtle nuances of the complete data set, not just the answer to a single question they may have formulated. TDA, in combination with machine learning algorithms, automatically represents data sets as a shape, in the form of a topological network, which presents many of the features associated to shape. This permits the easy interrogation of the data set, and suggests a collection of natural questions based on the shape’s theoretic features. This allows anyone who has generated data to easily analyze the data set and perform tasks such as data segmentation, and the identification of what features characterize the various segment. This new method of analysis is transforming the way individuals are looking at their data. They no longer have to walk into a dark, crowded, messy garage with a flashlight. TDA constitutes a new method of analysis that is transforming the way we understand and mine data. Domain experts and data scientists alike now have the ability to see their complex data as a connected system, easily understanding the relationships each single data point has with all other data points in even the most complex data sets.