Statistics > Machine Learning
[Submitted on 7 Jun 2014 (this version), latest version 26 Aug 2016 (v3)]
Title:Detecting Feature Interactions in Bagged Trees and Random Forests
View PDFAbstract:Additive models remain popular statistical tools due to their ease of interpretation and as a result, hypothesis tests for additivity have been developed to asses the appropriateness of these models. However, as data continues to grow in size and complexity, practicioners are relying more heavily on learning algorithms because of their predictive superiority. Due to the black-box nature of these learning methods, the increase in predictive power is assumed to come at the cost of interpretability and understanding. However, recent work suggests that many popular learning algorithms, such as bagged trees and random forests, have desireable asymptotic properties which allow for formal statistical inference when base learners are built with subsamples. This work extends the hypothesis tests previously developed and demonstrates that by constructing an appropriate test set, we may perform formal hypothesis tests for additivity amongst features. We develop notions of total and partial additivity and demonstrate that both tests can be carried out at no additional computational cost to the original ensemble. Simulations and demonstrations on real data are also provided.
Submission history
From: Lucas Mentch [view email][v1] Sat, 7 Jun 2014 00:58:30 UTC (41 KB)
[v2] Tue, 11 Nov 2014 22:55:32 UTC (265 KB)
[v3] Fri, 26 Aug 2016 21:10:03 UTC (752 KB)
Current browse context:
stat.ML
References & Citations
Bibliographic and Citation Tools
Bibliographic Explorer (What is the Explorer?)
Litmaps (What is Litmaps?)
scite Smart Citations (What are Smart Citations?)
Code, Data and Media Associated with this Article
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub (What is DagsHub?)
Gotit.pub (What is GotitPub?)
Papers with Code (What is Papers with Code?)
ScienceCast (What is ScienceCast?)
Demos
Recommenders and Search Tools
Influence Flower (What are Influence Flowers?)
Connected Papers (What is Connected Papers?)
CORE Recommender (What is CORE?)
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.