Detecting Feature Interactions in Bagged Trees and Random Forests

Mentch, Lucas; Hooker, Giles

Statistics > Machine Learning

arXiv:1406.1845v1 (stat)

[Submitted on 7 Jun 2014 (this version), latest version 26 Aug 2016 (v3)]

Title:Detecting Feature Interactions in Bagged Trees and Random Forests

Authors:Lucas Mentch, Giles Hooker

View PDF

Abstract:Additive models remain popular statistical tools due to their ease of interpretation and as a result, hypothesis tests for additivity have been developed to asses the appropriateness of these models. However, as data continues to grow in size and complexity, practicioners are relying more heavily on learning algorithms because of their predictive superiority. Due to the black-box nature of these learning methods, the increase in predictive power is assumed to come at the cost of interpretability and understanding. However, recent work suggests that many popular learning algorithms, such as bagged trees and random forests, have desireable asymptotic properties which allow for formal statistical inference when base learners are built with subsamples. This work extends the hypothesis tests previously developed and demonstrates that by constructing an appropriate test set, we may perform formal hypothesis tests for additivity amongst features. We develop notions of total and partial additivity and demonstrate that both tests can be carried out at no additional computational cost to the original ensemble. Simulations and demonstrations on real data are also provided.

Subjects:	Machine Learning (stat.ML); Applications (stat.AP)
Cite as:	arXiv:1406.1845 [stat.ML]
	(or arXiv:1406.1845v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.1406.1845

Submission history

From: Lucas Mentch [view email]
[v1] Sat, 7 Jun 2014 00:58:30 UTC (41 KB)
[v2] Tue, 11 Nov 2014 22:55:32 UTC (265 KB)
[v3] Fri, 26 Aug 2016 21:10:03 UTC (752 KB)

Statistics > Machine Learning

Title:Detecting Feature Interactions in Bagged Trees and Random Forests

Submission history

Access Paper:

References & Citations

1 blog link

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Detecting Feature Interactions in Bagged Trees and Random Forests

Submission history

Access Paper:

References & Citations

1 blog link

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators