pipette
ESEspañol

Scientific Data Analysis for Class-Informatics in Computational Taxonomy

Thomas B. Michelon, Hsieh Fushing

PreprintBold claims, read critically

In the authors' words

In this A.I. era, Computational Taxonomy is proposed to study complex systems by analyzing their databases under taxonomic hierarchies abiding the Principle of Science by providing "good explanations". Comparisons among branches or classes are carried out by Scientific Data Analysis (SDA) paradigm that explores all potential associative patterns, including interacting effects of all high orders, and evaluate finite sample precisions for all information pieces individually by effectively making use of all variables' categorical nature. Under each comparison, all confirmed information pieces are collected and displayed along row-axis of a heatmap with all involved study-subjects on the column-axis. Each comparison's heatmap individually characterizes participating classes and study-subjects and simultaneously provides a scientific basis for outlier detection upon all non-participants. All these heatmaps then collectively constitutes so-called Class-informatics that offers good explanations based on characteristic of all classes and study-subjects. Computational Taxonomy's Class-informatics indeed resolves multiple fundamental issues: Tukey's more than 60 years outlier detection problem, issue of self-correction annotation, and a crucial check on assumption of information-content equality between testing and training data sets in Machine Learning. A showcase of Computational Taxonomy is exclusively illustrated on Iris data.

Main resultLimitation the authors admit

Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.