Random forests for big data

Villa-Vialaneix, Nathalie,; Genuer, Robin; Poggi, Jean-Michel; Tuleau-Malot, Christine

Repository landing page

Random forests for big data

Authors: Nathalie, Villa-Vialaneix
Robin Genuer
Jean-Michel Poggi
Christine Tuleau-Malot
Publication date: 20 October 2016
Publisher: HAL CCSD

Abstract

International audienceBased on decision trees combined with aggregation and bootstrap ideas, random forests were introduced by Breiman in 2001. They are a powerful nonparametric statistical method allowing to consider in a single and versatile framework regression problems, as well as two-class and multi-class classification problems. Focusing on classification problems, this paper reviews available proposals about random forests in parallel environments as well as about online random forests. Then, we formulate various remarks for random forests in the Big Data context. Finally, we experiment three variants involving subsampling, Big Data-bootstrap and MapReduce respectively, on two massive datasets (15 and 120 millions of observations), a simulated one as well as real world data

Similar works

Full text

Hal-Diderot

oai:HAL:hal-02796431v1

Last time updated on 14/04/2021

This paper was published in Hal-Diderot.

Having an issue?

Is data on this page outdated, violates copyrights or anything else? Report the problem now and we will take corresponding actions after reviewing your request.