Twitter Channel for SAS Users

Dear All,

Screenshot-SAS Language Central (sascommunity) on Twitter - Mozilla Firefox

I just created a twitter channel for everthying about the SAS language. (independent of SAS institute as of now)

That’s right -using RSS feeeds, keyword filters, key bloggers AND tweeters and using http://www.twitterfeed.com this is a firehose of information on  everything SAS- includes Google Search and Twitter Search. I am still modifying it before transitioning it – Scandavian Airlines messes up some results for examples.

Anyways if you like to follow tech tweets this is one more-

http://twitter.com/sascommunity

PAW Blog Partner and 15 % off for you

paw09_blog_125

Dear Readers,

If you plan to attend Predictive Analytics World ( Oct20-21) in Washington DC,

Here are the speakers – source

Speakers Washington DC 2009:

Stephen L. Baker, Senior writer, BusinessWeek

Stephen L. BakerStephen L. Baker, author of The Numerati, is a senior writer at BusinessWeek, covering technology. Previously he was a Paris correspondent. Baker joined BusinessWeek in March, 1987, as manager of the Mexico City bureau, where he was responsible for covering Mexico and Latin America. He was named Pittsburgh bureau manager in 1992. Before BusinessWeek, Baker was a reporter for the El Paso Herald-Post. Prior to that, he was chief economic reporter for The Daily Journal in Caracas, Venezuela. Baker holds a bachelor’s degree from the University of Wisconsin and a master’s from the Columbia University Graduate School of Journalism. He blogs at TheNumerati.net and Blogspotting.net, and can be found on Twitter at @stevebaker.


John F. Elder, Ph.D., CEO and Founder, Elder Research, Inc.

Dr. John F. ElderDr. John F. Elder heads a data mining consulting team with offices in Charlottesville, Virginia and Washington DC. Founded in 1995, Elder Research, Inc. focuses on scientific and commercial applications of pattern discovery and optimization, including stock selection, image recognition, text mining, biometrics, drug efficacy, credit scoring, cross-selling, investment timing, and fraud detection.

John obtained a BS and MEE in Electrical Engineering from Rice University, and a PhD in Systems Engineering from the University of Virginia, where he’s an adjunct professor, teaching Optimization or Data Mining. Prior to 13 years leading ERI, he spent 5 years in aerospace defense consulting, 4 heading research at an investment management firm, and 2 in Rice’s Computational & Applied Mathematics department.

Dr. Elder has authored innovative data mining tools, is active on Statistics, Engineering, and Finance conferences and boards, is a frequent keynote conference speaker, and is General Chair of the 2009 Knowledge Discovery and Data Mining conference in Paris. John’s courses on data analysis techniques – taught at dozens of universities, companies, and government labs – are noted for their clarity and effectiveness. Dr. Elder was honored to serve for 5 years on a panel appointed by the President to guide technology for National Security. His book on Practical Data Mining, with Bob Nisbet and Gary Minor, will appear in May 2009.


Usama Fayyad, Ph.D., CEO, Open Insights

Dr. Usama FayyadDr. Usama Fayyad was until recently Yahoo!’s Chief Data Officer and Executive Vice President of Research & Strategic Data Solutions where he was responsible for Yahoo!’s global data strategy, architecting Yahoo!’s data policies and systems, prioritizing data investments, and managing the Company’s data analytics and data processing infrastructure. Fayyad also founded and oversaw the Yahoo! Research organization with offices around the world. Yahoo! Research is building the premier scientific research organization to develop the new sciences of the Internet, on-line marketing, and innovative interactive applications.

Prior to joining Yahoo!, Fayyad co-founded and led the DMX Group, a data mining and data strategy consulting and technology company that was acquired by Yahoo! in 2004. In early 2000, he co-founded and served as CEO of Revenue Science, Inc.(digiMine, Inc.), a data analysis and data mining company that built, operated and hosted data warehouses and analytics for some of the world’s largest enterprises in online publishing, retail, manufacturing, telecommunications and financial services. The company today specializes in Behavioral Targeting and advertising networks. Fayyad’s professional experience also includes five years spent leading the data mining and exploration group at Microsoft Research and building the data mining products for Microsoft’s server division. From 1989 to 1996 Fayyad held a leadership role at NASA’s Jet Propulsion Laboratory (JPL), where his work in the analysis and exploration of scientific databases gathered from observatories, remote-sensing platforms and spacecraft garnered him the top research excellence award that Caltech awards to JPL scientists, as well as a U.S. Government medal from NASA.

Fayyad earned his Ph.D. in engineering from the University of Michigan, Ann Arbor (1991), and also holds BSE’s in both electrical and computer engineering (1984); MSE in computer science and engineering (1986); and M.Sc. in mathematics (1989). He has published over 100 technical articles in the fields of data mining and Artificial Intelligence, is a Fellow of the AAAI and a Fellow of the ACM, has edited two influential books on the data mining and launched and served as editor-in-chief of both the primary scientific journal in the field of data mining and the primary newsletter in the technical community published by the ACM: SIGKDD Explorations.


Eric Siegel, Ph.D., Conference Chair

Eric SiegelThe president of Prediction Impact, Inc., Eric Siegel is an expert in predictive analytics and data mining and a former computer science professor at Columbia University, where he won awards for teaching, including graduate-level courses in machine learning and intelligent systems – the academic terms for predictive analytics. After Columbia, Dr. Siegel co-founded two software companies for customer profiling and data mining, and then started Prediction Impact in 2003, providing predictive analytics services and training to mid-tier through Fortune 100 companies.

Dr. Siegel is the instructor of the acclaimed training program, Predictive Analytics for Business, Marketing and Web, and the online version, Predictive Analytics Applied. He has published 13 papers in data mining research and computer science education, has served on 10 conference program committees, and has chaired a AAAI Symposium held at MIT.

you can register at http://www.predictiveanalyticsworld.com/register.php

Here is the pricing

Pricing
Predictive Analytics World Fall 2009

Includes breakfasts, lunches, priceless networking during coffee breaks, the PAW Reception, and full access to program sessions and sponsor expositions.

Super Early Bird Price
(till June 30)
Early Bird Price
(July 1 – Sept 4)
Regular     Price

Two Day Pass
(Oct 20-21)

$1190 $1390 $1590

Predictive Modeling Methods Workshop
(Oct 22)

$695 $795 $895

Putting Predictive Analytics to Work
(Oct 19)

$695 $795 $895

The discount code I can distribute to you  readers is the following: BLOGDC09 (15% off a two-day pass).You can do the maths…

(Ajay- Nopes I dont get money at all in these activities as blasted by some people
- but I do hope to get some good karma. Have a good time and book now).

PAW is back

The Predictive Analytics world is going to be back in October soon , and all those who missed out the stelar event can start booking now.

Here is the official BR ( blog Release)

Source: http://www.predictiveanalyticsworld.com/blog/wp-trackback.php?p=20

June 5th 2009 10:46 am

Keynotes at October’s PAW: Stephen Baker and Usama Fayyad

Predictive Analytics World, coming October 20-21 to Washington DC, has a great line-up of keynote speakers:

Stephen Baker, author of The Numerati and senior writer at BusinessWeek, where he’s been since 1987. Steve’s book has received a tremendous amount of attention this year. It is a revealing and insightful exploration of the opportunities and pitfalls of applied analytics, and consumer perception thereof.

Usama Fayyad, Ph.D. — CEO, Open Insights and formerly Yahoo!’s Chief Data Officer and Executive Vice President of Research & Strategic Data Solutions. Dr. Fayyad will return as an acclaimed keynote speaker. His keynote at February’s PAW (San Francisco) received extremely strong ratings from conference attendees.

Finally, Eric Siegel, Ph.D., will be kicking off PAW with a reprise of his keynote, “Five Ways to Lower Costs with Predictive Analytics.”

PMML 4.0

There are some nice changes in the PMML 4.0 version. PMML is the XML version for data modeling , or specificallyquoting the DMG group itself

PMML uses XML to represent mining models. The structure of the models is described by an XML Schema. One or more mining models can be contained in a PMML document. A PMML document is an XML document with a root element of type PMML. The general structure of a PMML document is:

  <?xml version="1.0"?>
  <PMML version="4.0"
    xmlns="http://www.dmg.org/PMML-4_0"
    xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" >

    <Header copyright="Example.com"/>
    <DataDictionary> ... </DataDictionary>

    ... a model ...

  </PMML>

So what is new in version 4. Here are some powerful modeling changes. For anyone with any XML knowledge PMML is the way to go.

PMML 4.0 – Changes from PMML 3.2

Associations

  • Itemset and AssociationRule elements are no longer enclosed within a “Choice” element
  • Added different scoring procedures: recommendation, exclusiveRecommendation and ruleAssociation with explanation and example
  • Changed version to “4.0” from “3.2” in the example(s)

BuiltinFunctions

Added the following functions:
  • isMissing
  • isNotMissing
  • equal
  • notEqual
  • lessThan
  • lessOrEqual
  • greaterThan
  • greaterOrEqual
  • isIn
  • isNotIn
  • and
  • or
  • not
  • isIn
  • isNotIn
  • if

Click on Image for better resolution

ClusteringModel

  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

Conformance

  • Changed all version references from “3.2” to “4.0”

DataDictionary

  • No changes

Functions

  • No changes

GeneralRegression

  • Changed to allow for Cox survival models and model ensembles
    • Add new model type: CoxRegression.
    • Allow empty regression model when model type is CoxRegression, so that baseline-only model could be represented.
    • Add new optional model attributes: endTimeVariable, startTimeVariable, subjectIDVariable, statusVariable, baselineStrataVariable, modelDF.
    • Add optional Matrix in Predictor to specify a contrast matrix, optional attribute referencePoint in Parameter.
    • Add new elements: BaseCumHazardTables, EventValues, BaselineStratum, BaselineCell.
    • Add examples of scoring for Cox Regression and contrast matrices.
    • Add new type of distribution: tweedie.
    • Add new attribute in model: targetReferenceCategory, so that the model can be used in MiningModel.
    • Changed version to “4.0” from “3.2” in the example(s)
    • Added reference to ModelExplanation element in the model XSD

GeneralStructure

Header

  • No changes

Interoperability

  • Changed: “As a result, a new approach for interoperability was required and is being introduced in PMML version 3.2.” to “As a result, a new approach for interoperability was introduced in PMML version 3.2.”

MiningSchema

  • Added frequencyWeight and analysisWeight as new options for usageType. They will not affect scoring, but will make model information more complete.

ModelComposition — No longer used, replaced by MultipleModels

ModelExplanation

  • New addition to PMML 4.0 that contains information to explain the models, model fit statistics, and visualization information.

ModelVerification

  • No changes

MultipleModels

  • Replaces ModelComposition. Important additions are segmentation and ensembles.
  • Added reference to ModelExplanation element in the model XSD

NaïveBayes

  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

NeuralNetwork

  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

Output

  • Extended output type to include Association rule models. The changes add a number of new attributes: “ruleFeature”, “algorithm”, “rank”, “rankBasis”, “rankOrder” and “isMultiValued”. A new enumeration type “ruleValue” is added to the RESULT-FEATURE

Regression

  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

RuleSet

  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

Sequence

  • Changed version to “4.0” from “3.2” in the example(s)

Statistics

  • accommodate weighted counts by replacing INT-ARRAY with NUM-ARRAY in DiscrStats and ContStats
  • change xs:nonNegativeInteger to xs:double in several places
  • add new boolean attribute ‘weighted’ to UnivariateStats and PartitionFieldStats elements
  • add new attribute cardinality in Counts
  • Also some very long lines in this document are now wrapped.

SupportVectorMachine

  • Added optional attribute threshold
  • Added optional attribute classificationMethod
  • Attribute alternateTargetCategory removed from SupportVectorMachineModel element and moved to SupportVectorMachine element
  • Changed the example slightly
  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

Targets

  • No changes

Taxonomy

  • Changed: “A TableLocator may contain any description which helps an application to locate a certain table. PMML 3.2 does not yet define the content. PMML users have to use their own extensions. The same applies to InlineTable.” to “A TableLocator may contain any description which helps an application to locate a certain table. PMML standard does not yet define the content. PMML users have to use their own extensions. The same applies to InlineTable.”

Text

  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

TimeSeriesModel

  • New addition to PMML 4.0 to support Time series models

Transformations

  • No changes

TreeModel

  • Changed version to “4.0” from “3.2” in the example(s)
  • Added reference to ModelExplanation element in the model XSD

Sources

http://www.dmg.org/v4-0/GeneralStructure.html

http://www.dmg.org/v4-0/Changes.html

and here are some companies using PMML already

http://www.dmg.org/products.html

I found the tool at http://www.dmg.org/coverage/ much more interesting though (see screenshot).

Screenshot-Mozilla Firefox

Zementis who we have covered in the interviews has played a steller role in bring together this common standard for data mining. Note Kxen model is also highlighted there.

The best PMML convertor tutorial is here

http://www.zementis.com/videos/PMML_Converter_iGoogle_gadget_2_demo.htm

Teratec : High Performance Computing Event

Here is a good HC event.

The Ter@tec’09 Forum
June 30 and July 1st, 2009, Supélec (91- France)


Incidently it is also quite close to KDD conference http://www.decisionstats.com/2009/06/19/conference-of-the-year-kdd-2009/

High performance Simulation and Computing for competitiveness, innovation and employment

© Ter@tec 2008 CEA

The international HPC event
The  Ter@tec annual Forum, created in 2006, is a major occasion of meetings, exchanges and reflection in the field of high performance simulation and computing.

Since the success of its first edition, the Ter@tec Forum has developed and is now organized on two days with plenary conferences, workshops and exhibition.

In 2008, more than 400 international attendees, from research and industry, providers and users, met to review the largest worldwide programs and discuss the perspectives and the major challenges we are facing, both on the technology side and on the user side.

The Forum was recognized as very successful, with high-level presentations and workshops, and the personal participation of Mrs Valérie Pécresse, French Minister for Higher education and Research and Mr Janez PotoČnik, European Commissioner for Science and Research.

Ter@tec 2009, the meeting of the HPC community around the technological and economical aspects of the high performance simulation and computing development.

Source- http://www.teratec.eu/gb/forum/index.html

Conference of the year: KDD 2009

This is one great co9nference you should attend if you have the time and inclination to check out latest advances in the world of Knowledge discovery. While KXEN ( from whom I consult on social madia) is a Gold Sponser- the following posts on workshops, demos and  papers will show you just how much technical stuff as opposed to marketing bullshit and jazz ( as in other confs)  is available in this conference. So pack your bags, and Viva La France for a grueling refreshing course in Knowledge Discovery and Text Mining. Incidentally KXEN intend to show their path breaking cutting edge social network analysis software KSN here.

Disclaimer- I am a social media consultant to KXEN.

KDD2009: Workshops

Abstracts

W1 – Statistical and Relational Learning and Mining in Bioinformatics (StReBio’09)

Jan Ramon, Fabrizio Costa, Christophe Costa Florencio, Joost Kok

Bioinformatics is an application domain where information is naturally represented in terms of relations between heterogenous objects. Modern experimentation and data acquisition techniques allow the study of complex interactions in biological systems. This raises interesting challenges because the amount of data is huge,some information can not be observed, and measurements may be noisy.

The StReBio’09 workshop invites contributions concerning applications of statistical relational learning and mining methods in bio-informatics domains. In particular, the workshop invites both regular papers, problem statements and problem solution papers.

Back to top…

W2 – The 3rd International Workshop on Knowledge Discovery from Sensor Data (SensorKDD-2009)

Olufemi Omitaomu, Auroop Ganguly, Joao Gama, Ranga Raju Vatsavai, Mohamed Medhat Gaber and Nitesh V. Chawla

Wide-area sensor infrastructures, remote sensors, RFIDs, and wireless sensor networks yield massive volumes of disparate, dynamic, and geographically distributed data. The Sensor-KDD 2009 workshop solicits papers that describe innovative solutions in offline data mining and/or real-time analysis of sensor or streaming data. Position papers that describe the challenges and requirements for sensor data based knowledge discovery in high-priority application domains, as well as relevant case studies, are particularly encouraged.

Back to top…

W3 – ACM SIGKDD Workshop on CyberSecurity and Intelligence Informatics (CSI-KDD)

Hsinchun Chen, Marc Dacier, Marie-Francine Moens, Gerhard Paaß, Christopher C. Yang

Computer supported communication and infrastructure are integral parts of modern economy. Their security is of incredible importance to a wide variety of practical domains ranging from Internet service providers to the banking industry and e-commerce, from corporate networks to the intelligence community. Of interest to this workshop are novel knowledge discovery methods addressing this field, e.g. adaptive, active or anticipatory approaches integrating new types of contents and protocols. Equally important are innovative applications demonstrating the effectiveness of data mining in solving real-world security problems.

Back to top…

W4 – Workshop on Visual Analytics and Knowledge Discovery (VAKD ’09)

Kai Puolamäki, Heikki Mannila, Alessio Bertone, Silvia Miksch, Mark A. Whiting, Jean Scholtz

The goal of Visual Analytics is to derive insight from massive, dynamic, ambiguous, and often conflicting data; detect the expected and discover the unexpected; provide timely, defensible, and understandable assessments; and communicate the assessment effectively for action. The goal of this workshop is to raise the awareness of the KDD community for the importance of Visual Analytics and bring together researcher from the underlying fields to bridge the gap between them—to write a KDD research roadmap on Visual Analytics.

Back to top…

W5 – The Third International Workshop on Data Mining and Audience Intelligence for Advertising (ADKDD)

Ying Li, Arun C. Surendran, and Dou Shen

Advertising, especially online advertising, is growing rapidly and brings about large volumes of data along with challenging data mining problems. Following on the success of ADKDD 2007 and 2008, ADKDD 2009 is to be held in Paris France, in conjunction with KDD 2009, to provide a high-level international forum for the academic community and the industry to present the state of the art of algorithms and applications of advertising.

We encourage papers that bring up and formalize new research problems in online advertising, or propose novel data mining techniques for existing problems. We plan to cover (but not restricted to) the following areas: Mining for Ad Relevance and Ranking; Audience Intelligence & User Modeling; Content Understanding; Search Engine Marketing, Optimization (SEMs, SEOs) and Other Topics in Advertising. Accepted papers will be achieved in ACM Digital Library and one or two papers will be recommended to SIGKDD Explorations.

Back to top…

W6 – The 3rd Workshop on Social Network Mining and Analysis (SNA-KDD)

Lee Giles, Prasenjit Mitra, Igor Perisic, John Yen, Haizheng Zhang

(Abstract Coming Soon)

Back to top…

W7 – Human Computation Workshop (HCOMP 2009)

Paul Bennett, Raman Chandrasekar, Max Chickering, Panos Ipeirotis, Edith Law, Foster Provost, Anton Mityagin, Luis von Ahn

Human computation is a new research area that studies the process of channeling the vast internet population to perform tasks or provide data towards solving difficult problems that no known computer algorithms can yet solve perfectly and efficiently, e.g. digitize books, recognize objects in images and songs, translate sentences, summarize news articles, annotate videos etc. The goal of HCOMP 2009 is to bring together academic and industry researchers in a stimulating discussion of existing human computation applications, such as Games With A Purpose (e.g. the ESP game), Mechanical Turk and CAPTCHAs, and future directions of this new subject area.

Included in the workshop are invited talks, presentations, posters, and a demo session where participants are invited to showcase their human computation applications.

Back to top…

W8 – Data Mining using Matrices and Tensors (DMMT’09)

Chris Ding, Tao Li

This workshop will present recent advances in algorithms and methods using matrix and scientific computing/applied mathematics for modeling and analyzing massive, high-dimensional, and nonlinear-structured data. One main goal of the workshop is to bring together leading researchers on many topic areas (e.g., computer scientists, computational and applied mathematicians) to assess the state-of-the-art, share ideas, and form collaborations. We also wish to attract practitioners who seek novel ideas for applications.

Back to top…

W9 – Third Workshop on Data Mining Case Studies and Practice Prize (DMCS)

Gabor Melli, Peter van der Putten, Brendan Kitts

The Data Mining Case Studies Workshop and Practice Prize was established to recognize the very best data mining deployments for the year. Data Mining Case Studies will highlight data mining implementations that have been responsible for a significant and measurable improvement in business operations, advanced scientific discoveries, or provided other benefits to humanity. The best paper will be awarded the Practice Prize. Do you have an outstanding data mining application? This is a unique opportunity to be recognized for your work.

Back to top…

W10 – KDD cup 2009: Fast Scoring on a Large Database (KDDcup09)

Isabelle Guyon, David Vogel

This workshop will discuss the results of the KDD cup 2009. The competition is organized around a large dataset provided by the French telecom company Orange. It is a problem of Customer Relationship Management (CRM), a key element of modern marketing strategies. Orange offered the opportunity to work on a large marketing database to predict the propensity of customers to switch provider (churn), buy new products or services (appetency), or buy upgrades or add-ons proposed to them to make the sale more profitable (up-selling).

Back to top…

W11 – The First ACM SIGKDD Workshop on Knowledge Discovery from Uncertain Data (U’09)

Jian Pei, Lise Getoor, Ander de Keijzer

The First ACM SIGKDD International Workshop on Knowledge Discovery from Uncertain Data (U’09) is to discuss in depth the challenges, opportunities and techniques on the topic of analyzing and mining uncertain data. The theme of this workshop is to make connections among the research areas of probabilistic databases, probabilistic reasoning, and data mining, as well as to build bridges among the aspects of models, data, applications, novel mining tasks and effective solutions. By making connections among different communities, we aim at understanding each other in terms of scientific foundation as well as commonality and differences in research methodology.

Back to top…

KDD-09 Call For Workshop Proposals (Expired)

The ACM KDD-2009 organizing committee invites proposals for workshops to be held in conjunction with the conference. The purpose of a workshop is to provide participants with the opportunity to present and discuss novel research ideas on active and emerging topics of knowledge discovery and data mining. A workshop should also support the interaction and feedback among topic specialists from academia, industry and government.

A workshop may be organized around industrial applications in a particular domain and the challenges this domain poses, such as the Netflix workshop on recommender systems (http://netflixkddworkshop2008.info/).

A workshop may also include a challenge problem, such as the one on time series classification that took place in 2007 (http://www.cs.ucr.edu/~eamonn/SIGKDD2007TimeSeries.html). A session with papers that address a challenge complements the more diverse sessions with regular papers and improves the potential for discussion. Because such challenges require extra time to plan, we may be willing to provide early notice of acceptance.

The organizers of approved workshops are required to announce the workshop and call for papers, gather submissions, conduct the reviewing process and decide upon the final workshop program. They must also prepare an informal set of workshop proceedings to be distributed with the registration materials at the conference. They may choose to form organizing or program committees for assistance in these tasks. The logistics of the workshops will be done with the help from the ACM KDD-2009 organizers.

Back to top…

source-http://www.kdd.org/kdd/2009/workshops.html