| By David Smith | Article Rating: |
|
| November 13, 2012 10:49 AM EST | Reads: |
545 |
By Joseph Rickert
In a recent blog post, David Smith reported on a talk that
Steve Yun and I gave at STRATA in NYC about building and benchmarking Poisson
GLM models on various platforms. The results presented showed that the rxGlm
function from Revolution Analytics’ RevoScaleR package running on a five node
cluster outperformed a Map Reduce/ Hadoop implementation as well as an
implementation of legacy software running on a large server. An alert R user
posted the following comment on the blog:
As a poisson regression was used,
it would be nice to also see as a benchmark the computational speed when using
the biglm package in open source R? Just import your csv in sqlite and run
biglm to obtain your poisson regression. Biglm also loads in data in R in
chunks in order to update the model so that looks more similar to the
RevoScaleR setup then just running plain glm in R.
This seemed like a reasonable, simple enough experiment. So
we tried it. The benchmark results presented at STRATA were done on a 145
million record file, but as a first step, I thought that I would try it on
a 14 million record subset that I already had loaded on my PC, a quad core Dell,
with i7 processors and 8GB of RAM. It
took almost an hour to build the SQLite data base:
# make a SQLite database out of the csv file
library(sqldf)
sqldf("attach AdataT2SQL as new")
file <- file.path(getwd(),"adatat2.csv")
read.csv.sql(file, sql = "create table main.AT2_10Pct as select * from file">
Read the original blog entry...
Published November 13, 2012 Reads 545
Copyright © 2012 SYS-CON Media, Inc. — All Rights Reserved.
Syndicated stories and blog feeds, all rights reserved by the author.
More Stories By David Smith
David Smith is Vice President of Marketing and Community at Revolution Analytics. He has a long history with the R and statistics communities. After graduating with a degree in Statistics from the University of Adelaide, South Australia, he spent four years researching statistical methodology at Lancaster University in the United Kingdom, where he also developed a number of packages for the S-PLUS statistical modeling environment. He continued his association with S-PLUS at Insightful (now TIBCO Spotfire) overseeing the product management of S-PLUS and other statistical and data mining products.< David smith is the co-author (with Bill Venables) of the popular tutorial manual, An Introduction to R, and one of the originating developers of the ESS: Emacs Speaks Statistics project. Today, he leads marketing for REvolution R, supports R communities worldwide, and is responsible for the Revolutions blog. Prior to joining Revolution Analytics, he served as vice president of product management at Zynchros, Inc. Follow him on twitter at @RevoDavid
- Cloud People: A Who's Who of Cloud Computing
- Windows Azure IaaS Reaches General Availability
- New Relic Q1 2013 Blazes Past Growth Targets and Reaches 40,000 Active Customer Accounts
- Portable Experimenter’s Platform, Powered by Raspberry Pi
- Basho Announces Open Source Riak CS and General Availability of Riak CS Enterprise v1.3
- MicroStrategy Announces General Availability of MicroStrategy 9.3.1
- AMAX Launches StorMax(TM) CFS, powered by IBM(R) General Parallel File System(TM) (GPFS(TM))
- MicroStrategy Announces General Availability of MicroStrategy 9.3.1
- CollabNet And UC4 Announce General Availability Of Joint Enterprise DevOps Platform
- Project Floodlight Grows to the World’s Largest SDN Ecosystem; Global Users, Contributors and Partners Innovating Using Open Source SDN
- Mobility News Weekly – Week of March 17, 2013
- New Relic Named Best Place to Work in the Bay Area for Second Year in a Row
- Cloud People: A Who's Who of Cloud Computing
- Windows Azure IaaS Reaches General Availability
- New Relic Q1 2013 Blazes Past Growth Targets and Reaches 40,000 Active Customer Accounts
- Portable Experimenter’s Platform, Powered by Raspberry Pi
- SUSE Receives Common Criteria Security Certifications
- Basho Announces Open Source Riak CS and General Availability of Riak CS Enterprise v1.3
- Appeon Mobile Beta2 - 48 Hours
- Granular Enforcement of Access to File Systems Featured in Latest Release of FoxT ServerControl
- MicroStrategy Announces General Availability of MicroStrategy 9.3.1
- AMAX Launches StorMax(TM) CFS, powered by IBM(R) General Parallel File System(TM) (GPFS(TM))
- MicroStrategy Announces General Availability of MicroStrategy 9.3.1
- CollabNet And UC4 Announce General Availability Of Joint Enterprise DevOps Platform
- Cloud People: A Who's Who of Cloud Computing
- Red Hat Named "Platinum Sponsor" of Virtualization Conference & Expo
- An Introduction to Ant
- Cloud Expo 2011 East To Attract 10,000 Delegates and 200 Exhibitors
- Google Web Toolkit: Finally Java Has Been Put into JavaScript!
- Cloud Expo, Inc. Announces Cloud Expo 2011 New York Venue
- AJAX World RIA Conference News - AJAX & RIA with Server-Side JavaScript
- Early Notes on GoogleApps
- President & CTO of 3tera Speaking Next Week at SYS-CON's Cloud Computing Expo November 19-21 in Silicon Valley
- Rating JRuby, Jython, and Groovy on the Java Platform
- Python Creator Guido van Rossum to Present the Next-Generation Python 3000
- Rackspace Cloud APIs Open Sourced






















