SAS X

0o0 0O

Tal G, creator of the rbloggers.com website, has created a new blog aggregator for SAS language users at http://sas-x.com/

With almost 26 blogs joining there (I suspect many more should join , it seems like a good website to use for analytics users and students.  My favorite SAS Blog is http://statcompute.spaces.live.com/ – its pure code- little anything else.

Related-

SAS MACRO TO CALCULATE PDO (Points to Double Odds) OF A SCORECARD

A SAS MACRO FOR DECISION STUMP

A DEMO OF VECTOR AUTOREGRESSIVE FORECASTING MODEL

 

 

 

Interview Jamie Nunnelly NISS

An interview with Jamie Nunnelly, Communications Director of National Institute of Statistical Sciences

Ajay– What does NISS do? And What does SAMSI do?

Jamie– The National Institute of Statistical Sciences (NISS) was established in 1990 by the national statistics societies and the Research Triangle universities and organizations, with the mission to identify, catalyze and foster high-impact, cross-disciplinary and cross-sector research involving the statistical sciences.

NISS is dedicated to strengthening and serving the national statistics community, most notably by catalyzing community members’ participation in applied research driven by challenges facing government and industry. NISS also provides career development opportunities for statisticians and scientists, especially those in the formative stages of their careers.

The Institute identifies emerging issues to which members of the statistics community can make key contributions, and then catalyzes the right combinations of researchers from multiple disciplines and sectors to tackle each problem. More than 300 researchers from over 100 institutions have worked on our projects.

The Statistical and Applied Mathematical Sciences Institute (SAMSI) is a partnership of Duke University,  North Carolina State University, The University of North Carolina at Chapel Hill, and NISS in collaboration with the William Kenan Jr. Institute for Engineering, Technology and Science and is part of the Mathematical Sciences Institutes of the NSF.

SAMSI focuses on 1-2 programs of research interest in the statistical and/or applied mathematical area and visitors from around the world are involved with the programs and come from a variety of disciplines in addition to mathematics and statistics.

Many come to SAMSI to attend workshops, and also participate in working groups throughout the academic year. Many of the working groups communicate via WebEx so people can be involved with the research remotely. SAMSI also has a robust education and outreach program to help undergraduate and graduate students learn about cutting edge research in applied mathematics and statistics.

Ajay– What successes have you had in 2010- and what do you need to succeed in 2011. Whats planned for 2011 anyway

Jamie– NISS has had a very successful collaboration with the National Agricultural Statistical Service (NASS) over the past two years that was just renewed for the next two years. NISS & NASS had three teams consisting of a faculty researcher in statistics, a NASS researcher, a NISS mentor, a postdoctoral fellow and a graduate student working on statistical modeling and other areas of research for NASS.

NISS is also working on a syndromic surveillance project with Clemson University, Duke University, The University of Georgia, The University of South Carolina. The group is currently working with some hospitals to test out a model they have been developing to help predict disease outbreak.

SAMSI had a very successful year with two programs ending this past summer, which were the Stochastic Dynamics program and the Space-time Analysis for Environmental Mapping, Epidemiology and Climate Change. Several papers were written and published and many presentations have been made at various conferences around the world regarding the work that was conducted as SAMSI last year.

Next year’s program is so big that the institute has decided to devote all it’s time and energy around it, which is uncertainty quantification. The opening workshop, in addition to the main methodological theme, will be broken down into three areas of interest under this broad umbrella of research: climate change, engineering and renewable energy, and geosciences.

Ajay– Describe your career in science and communication.

Jamie– I have been in communications since 1985, working for large Fortune 500 companies such as General Motors and Tropicana Products. I moved to the Research Triangle region of North Carolina after graduate school and got into economic development and science communications first working for the Research Triangle Regional Partnership in 1994.

From 1996-2005 I was the communications director for the Research Triangle Park, working for the Research Triangle Foundation of NC. I published a quarterly magazine called The Park Guide for awhile, then came to work for NISS and SAMSI in 2008.

I really enjoy working with the mathematicians and statisticians. I always joke that I am the least educated person working here and that is not far from the truth! I am honored to help get the message out about all of the important research that is conducted here each day that is helping to improve the lives of so many people out there.

Ajay– Research Triangle or Silicon Valley– Which is better for tech people and why? Your opinion

Jamie– Both the Silicon Valley and Research Triangle are great regions for tech people to locate, but of course, I have to be biased and choose Research Triangle!

Really any place in the world that you find many universities working together with businesses and government, you have an area that will grow and thrive, because the collaborations help all of us generate new ideas, many of which blossom into new businesses, or new endeavors of research.

The quality of life in places such as the Research Triangle is great because you have people from around the world moving to a place, each bringing his/her culture, food, and uniqueness to this place, and enriching everyone else as a result.

Two advantages the Research Triangle has over Silicon Valley are that the Research Triangle has a bigger diversity of industries, so when the telecommunications industry busted back in 2001-02, the region took a hit, but the biotechnology industry was still growing, so unemployment rose, but not to the extent that other areas might have experienced.

The latest recession has hit us all very hard, so even this strategy has not made us immune to having high unemployment, but the Research Triangle region has been pegged by experts to be one of the first regions to emerge out of the Great Recession.

The other advantage I think we have is that our cost of living is still much more reasonable than Silicon Valley. It’s still possible to get a nice sized home, some land and not break the bank!

Ajay– How do you manage an active online social media presence, your job and your family. How important is balance in professional life and when young professional should realize this?

Jamie– Balance is everything, isn’t it? When I leave the office, I turn off my iPhone and disconnect from Twitter/Facebook etc.

I know that is not recommended by some folks, but I am a one person communications department and I love my family and friends and feel its important to devote time to them as well as to my career.

I think it is very important for young people to establish this early in their careers because if they don’t they will fall victim to working way too many hours and really, who loves you at the end of the day?

Your company may appreciate all you do for them, but if you leave, or you get sick and cannot work for them, you will be replaced

. Lee Iacocca, former CEO of Chrystler, said, “No matter what you’ve done for yourself or for humanity, if you can’t look back on having given love and attention to your own family, what have you really accomplished?” I think that is what is really most important in life.

About-

Jamie Nunnelly has been in communications for 25 years. She is currently on the board of directors for Chatham County Economic Development Corporation and Leadership Triangle & is a member of the International Association of Business Communicators and the Public Relations Society of America. She earned a bachelor’s degree in interpersonal and public communications at Bowling Green State University and a master’s degree in mass communications at the University of South Florida.

You can contact Jamie at http://niss.org/content/jamie-nunnelly or on twitter at

STEM is cool

Lady Gaga holding a speech at National Equalit...
Image via Wikipedia

A good video created by my favorite social media people from a company in North Carolina.

STEM is cool (Science Technology Engineering Maths?)

No, Science is not kool aid- it is just COOL. and better paying than watching Justin Bieber or Lady Gaga videos. Get those lazy teenagers out of Glee clubs and back into Science clubs.

The video itself-

Disclaimer- I have no direct or indirect  financial relationship with the creators of this video. I think it is cool people express creativity in positive ways to help their favorite software,company, and even the world. Blah Blah Blah 🙂

Yeah, STEM is cool again.

 

 

Summer School on Uncertainty Quantification

Scheme for sensitivity analysis
Image via Wikipedia

SAMSI/Sandia Summer School on Uncertainty Quantification – June 20-24, 2011

http://www.samsi.info/workshop/samsisandia-summer-school-uncertainty-quantification

The utilization of computer models for complex real-world processes requires addressing Uncertainty Quantification (UQ). Corresponding issues range from inaccuracies in the models to uncertainty in the parameters or intrinsic stochastic features.

This Summer school will expose students in the mathematical and statistical sciences to common challenges in developing, evaluating and using complex computer models of processes. It is essential that the next generation of researchers be trained on these fundamental issues too often absent of traditional curricula.

Participants will receive not only an overview of the fast developing field of UQ but also specific skills related to data assimilation, sensitivity analysis and the statistical analysis of rare events.

Theoretical concepts and methods will be illustrated on concrete examples and applications from both nuclear engineering and climate modeling.

The main lecturers are:
Dan Cacuci (N.C. State University): data assimilation and applications to nuclear engineering

Dan Cooley (Colorado State University): statistical analysis of rare events
This short course will introduce the current statistical practice for analyzing extreme events. Statistical practice relies on fitting distributions suggested by asymptotic theory to a subset of data considered to be extreme. Both block maximum and threshold exceedance approaches will be presented for both the univariate and multivariate cases.

Doug Nychka (NCAR): data assimilation and applications in climate modeling
Climate prediction and modeling do not incorporate geophysical data in the sequential manner as weather forecasting and comparison to data is typically based on accumulated statistics, such as averages. This arises because a climate model matches the state of the Earth’s atmosphere and ocean “on the average” and so one would not expect the detailed weather fluctuations to be similar between a model and the real system. An emerging area for climate model validation and improvement is the use of data assimilation to scrutinize the physical processes in a model using observations on shorter time scales. The idea is to find a match between the state of the climate model and observed data that is particular to the observed weather. In this way one can check whether short time physical processes such as cloud formation or dynamics of the atmosphere are consistent with what is observed.

Dongbin Xiu (Purdue University): sensitivity analysis and polynomial chaos for differential equations
This lecture will focus on numerical algorithms for stochastic simulations, with an emphasis on the methods based on generalized polynomial chaos methodology. Both the mathematical framework and the technical details will be examined, along with performance comparisons and implementation issues for practical complex systems.

The main lectures will be supplemented by discussion sessions and by presentations from UQ practitioners from both the Sandia and Los Alamos National Laboratories.

http://www.samsi.info/workshop/samsisandia-summer-school-uncertainty-quantification

Statistical Analysis with R- by John M Quick

I was asked to be a techie reviewe for John M Quick’s new R book “Statistical Analysis with R” from Packt Publishing some months ago-(very much to my surprise I confess)-

I agreed- and technical reviewer work does take time- its like being a mid wife and there is whole team trying to get the book to birth.

Statistical Analysis with R- is a Beginner’s Guide so has nice screenshots, simple case studies, and quizzes to check recall of student/ reader. I remember struggling with the official “beginner’s guide to R” so this one is different in that it presents a story of a Chinese Army and how to use R to plan resources to fight the battle. It’s recommended especially for undergraduate courses- R need not be an elitist language- and given my experience with Asian programming acumen – I am sure it is a matter of time before high schools in India teach basic R in final years ( I learnt quite a shit load of quantum physics as compulsory topics in Indian high schools- but I guess we didnt have Jersey Shore things to do)

Congrats to author Mr John M Quick- he is doing his educational Phd from ASU- and I am sure both he and his approach to making education simple informative and fun will go places.

Only bad thing- The Name Statistical Analysis with R has atleast three other books , but I guess Google will catch up to it.

This book is here-https://www.packtpub.com/statistical-analysis-with-r-beginners-guide/book

Deduping Facebook

How many accounts in Facebook are one unique customer?

Does 500 million human beings as Facebook customers sound too many duplicates? (and how much more can you get if you get the Chinese market- FB is semi censored there)

Is Facebook response rate on ads statistically same as response rates on websites or response rates on emails or response rates on spam?

Why is my Facebook account (which apparently) I am free to download one big huge 130 mb file, not chunks of small files I can download.

Why cant Facebook use URL shorteners for the links of Photos (ever seen those tiny fonted big big urls below each photo)

How come Facebook use so much R (including making the jjplot package) but wont sponsor a summer of code contest (unlike Google)-100 million for schools and 2 blog posts for R? and how much money for putting e education content and games on Facebook.

Will Facebook ever create an-in  house game?  Did Google put money in Zynga (FB’s top game partner) because it likes

games 🙂 ? How dependent is FB on Zynga anyways?

So many questions———————————————————— so little time

 

Data Visualization using Tableau

Image representing Tableau Software as depicte...
Image via CrunchBase

Here is a great piece of software for data visualization– the public version is free.

And you can use it for Desktop Analytics as well as BI /server versions at very low cost.

About Tableau Software

http://www.tableausoftware.com/press_release/tableau-massive-growth-hiring-q3-2010

Tableau was named by Software Magazine as the fastest growing software company in the $10 million to $30 million range in the world, and the second fastest growing software company worldwide overall. The ranking stems from the publication’s 28th annual Software 500 ranking of the world’s largest software service providers.

“We’re growing fast because the market is starving for easy-to-use products that deliver rapid-fire business intelligence to everyone. Our customers want ways to unlock their databases and produce engaging reports and dashboards,” said Christian Chabot CEO and co-founder of Tableau.

http://www.tableausoftware.com/about/who-we-are

History in the Making

Put together an Academy-Award winning professor from the nation’s most prestigious university, a savvy business leader with a passion for data, and a brilliant computer scientist. Add in one of the most challenging problems in software – making databases and spreadsheets understandable to ordinary people. You have just recreated the fundamental ingredients for Tableau.

The catalyst? A Department of Defense (DOD) project aimed at increasing people’s ability to analyze information and brought to famed Stanford professor, Pat Hanrahan. A founding member of Pixar and later its chief architect for RenderMan, Pat invented the technology that changed the world of animated film. If you know Buzz and Woody of “Toy Story”, you have Pat to thank.

Under Pat’s leadership, a team of Stanford Ph.D.s got together just down the hall from the Google folks. Pat and Chris Stolte, the brilliant computer scientist, realized that data visualization could produce large gains in people’s ability to understand information. Rather than analyzing data in text form and then creating visualizations of those findings, Pat and Chris invented a technology called VizQL™ by which visualization is part of the journey and not just the destination. Fast analytics and visualization for everyone was born.

While satisfying the DOD project, Pat and Chris met Christian Chabot, a former data analyst who turned into Jello when he saw what had been invented. The three formed a company and spun out of Stanford like so many before them (Yahoo, Google, VMWare, SUN). With Christian on board as CEO, Tableau rapidly hit one success after another: its first customer (now Tableau’s VP, Operations, Tom Walker), an OEM deal with Hyperion (now Oracle), funding from New Enterprise Associates, a PC Magazine award for “Product of the Year” just one year after launch, and now over 50,000 people in 50+ countries benefiting from the breakthrough.

also see http://www.tableausoftware.com/about/leadership

http://www.tableausoftware.com/about/board

—————————————————————————-

and now  a demo I ran on the Kaggle contest data (it is a csv dataset with 95000 rows)

I found Tableau works extremely good at pivoting data and visualizing it -almost like Excel on  Steroids. Download the free version here ( I dont know about an academic program (see links below) but software is not expensive at all)

http://buy.tableausoftware.com/

Desktop Personal Edition

The Personal Edition is a visual analysis and reporting solution for data stored in Excel, MS Access or Text Files. Available via download.

Product Information

$999*

Desktop Professional Edition

The Professional Edition is a visual analysis and reporting solution for data stored in MS SQL Server, MS Analysis Services, Oracle, IBM DB2, Netezza, Hyperion Essbase, Teradata, Vertica, MySQL, PostgreSQL, Firebird, Excel, MS Access or Text Files. Available via download.

Product Information

$1800*

Tableau Server

Tableau Server enables users of Tableau Desktop Professional to publish workbooks and visualizations to a server where users with web browsers can access and interact with the results. Available via download.

Product Information

Contact Us

* Price is per Named User and includes one year of maintenance (upgrades and support). Products are made available as a download immediately after purchase. You may revisit the download site at any time during your current maintenance period to access the latest releases.