An E Book Review

Here is a nice e book I got from my colleagues at The Customer Collective. I really really loved the friendly design making this a very easy e-book to read unlike other self help books. It has tips from 11 top sales experts in how to sell in that specific sector during a recession- and these are not essays but nice bullet point specific action items. Hat tip to the editor and the authors here.

You can download the ebook here

image image

Working Together -Yuuguu

Here is a nice software called Yuuguu. I really liked the software because it enables sharing of screens so we can have secure virtual meetings, it is quite light on processor and memory , and best of all it works across Ubuntu 64 bit, Windows XP, Mac OS .

So if you want to work on a project team that sit across the seven seas and the big pond, and you feel the best way is to talk through a demonstration rather than give  a documentation – then Yuuguu is the right software for you. The worst part of this software — is probably the name. And yes it is free and has a paid version as well.

See www.Yuuguu.com

image

KXEN and a Data Mining Survey

Recently KXEN, the data mining and modeling automation company which has also pioneered social network analytics software came in for a bit of customer love in a data mining survey.

kxen

As per the site 

KXEN’s next generation automated data mining software is a strategic solution for 90% of user organizations and has won their support and praise in a new customer satisfaction survey, the findings of which are revealed today. 

Of almost 2,000 users polled 90% of those responding said the company’s advanced analytics software was strategic to their activity, 87% were highly or very highly satisfied and 85% agreed the software had met or exceeded all of their expectations. The results underscore KXEN’s growing importance in a market traditionally dominated by more costly, harder to use first generation offerings.

KXEN’s analytic software was also highly rated for its simple, clear interface with all respondents agreeing that KXEN solutions were easy to use, and 90% stating its new graphical front end had brought yet more usability benefits.  Confirming these findings, users responding included sales, marketing and other line of business staff as well as specialist analysts, data miners, academics and statisticians.

Turning to the results of using KXEN’s software, 98% of all those responding stated it had improved their overall business with the same number agreeing it had speeded up their data modeling activities. 96% said KXEN had increased the value of predictive analytics in their companies.

Of course there are numerous surveys (including probably the best is from KD Nuggets) and I am trying to find the raw data and samples for this survey as I write. But it is a promising step up for a company I have admired since 2004, when I first tested it, and as late as last year I was building online models with it. Predictably Roger Hadaad whom we interviewed in January 2009 was all praise for his team and its splendid product. Well Done, guys take a bow- it is about time ! A great example of a company that builds innovative analytics quitely without getting into any tangles with open source or business intellgence sentiments.

Ajay- I am a consultant to KXEN for Social Networks Analysis.

terrific Tr.im trims Tweet time

Okay, the title of the post was bad attempt at a haiku. But the tr.im plugin for Firefox is incredible and helps you tweet interesting reading in matter of seconds. More importantly it shows you the analytics behind how many actual users went to that particular tr.im url. While Tr.im is yet another url shortening service like the tinyurl.com and bit.ly services, what makes Tr.im stand out in a terrific manner are the following innovations –

1) User friendly Firefox Plugin that can be downloaded from https://addons.mozilla.org/en-US/firefox/addon/10232/

See the screenshot of the Tr.im panel which conveniently opens on the left. The Statistics can be seen in the separate window ( note the Twitterfox application which is also open on the right – that is a separate application)

2) Analytics for tracking the locations, of people who click on the url and whether they were human or a bot.

3) Seamless Twitter integration even for multiple accounts

So it seems like you will run out of excuses to run away from Twitter soon, and all the additional social network data being generated could really help the next generation of response and online propensity models.

Tr.im that!!

screenshottrim

KXEN – Automated Regression Modeling

I have used KXEN many times for building and testing propensity models. The regression modeling feature of KXEN is awesome in the sense it can make model building very easy to build and deliver.

The KXEN package K2R is the package responsible for this and uses robust regression. A word of the basic mathematical theory behind KXEN’s automated modeling – the technique is called Structural Risk Minimization. You can read more on the basic mathematical technique here or http://www.svms.org/srm/. The following is an extract from the same source.

Structural risk minimization (SRM) (Vapnik and Chervonekis, 1974) is an inductive principle for model selection used for learning from finite training data sets. It describes a general model of capacity control and provides a trade-off between hypothesis space complexity (the VC dimension of approximating functions) and the quality of fitting the training data (empirical error). The procedure is outlined below.

  1. Using a priori knowledge of the domain, choose a class of functions, such as polynomials of degree n, neural networks having n hidden layer neurons, a set of splines with n nodes or fuzzy logic models having n rules.
  2. Divide the class of functions into a hierarchy of nested subsets in order of increasing complexity. For example, polynomials of increasing degree.
  3. Perform empirical risk minimization on each subset (this is essentially parameter selection).
  4. Select the model in the series whose sum of empirical risk and VC confidence is minimal.

Sewell (2006) SVMs use the spirit of the SRM principle.

“Structural risk minimization (SRM) (Vapnik 1995) uses a set of models ordered in terms of their complexities. An example is polynomials of increasing order. The complexity is generally given by the number of free parameters. VC dimension is another measure of model complexity. In equation 4.37, we can have a set of decreasing ?i to get a set of models ordered in increasing complexity. Model selection by SRM then corresponds to finding the model simplest in terms of order and best in terms of empirical error on the data.”
Alpaydin (2004), pages 80-81

Now back to the automated regression modeling.

Robust Regression

(K2R) is a universal solution for Classification, Regression, and Attribute Importance. It enables the prediction of behaviors (nominal targets) or quantities (continuous targets).

Unlike traditional regression algorithms, K2R can safely handle a very high numbers of input attributes (over 10,000) in an automated fashion. K2R provides indicators and graphs to ensure that the quality and robustness of trained models can be easily assessed. K2R graphically displays the attribute importance, which provides the relative importance of each attribute for explaining a given business question. At the same time it gives a clear indication of which attributes either contain no relevant information or are redundant with other attributes.

Benefits: The business value of a data mining project is increased by either training more models or completing the project faster. The ability to train more models allows a larger number of scenarios to be tested at a higher level of granularity. For example, if a direct marketing campaign benefits from separate models trained per region, per customer, segment, per month, the automation of K2R allows all of these models to be trained and safely deployed using the same amount or fewer resources than with traditional tools. learn more

What: K2R is a regression algorithm that allows building models to predict categories or continuous variables.

Why: Traditionally, building robust predictive models required a lot of time and expertise, which prevented companies from using data mining as part of their every day business decisions. K2R makes it easy to build and deploy predictive models in the fraction of the time it takes using classical statistical tools.

How: K2R maps a set of descriptive attributes (model inputs) and target attributes (model output). It uses an algorithm patented by KXEN, which is a derivation of a principle described by V. Vapnik as “Structured Risk Minimization.” Instead of looking for the best performance on a known dataset, K2R automatically finds the best compromise between quality and robustness. The resulting models are expressed as a polynomial expression of the input numbers. The only element specified by the user is the polynomial degree. To improve modeling speed, K2R can also build multi-target models.

Benefits for the business user: K2R allows the business user to easily build and understand advanced predictive models without statistical knowledge. A model can be created in a matter of minutes. Two performance indicators describe model quality (Ki) and model reliability or the ability to produce similar on new data (Kr).

K2R graphically displays the individual variable contribution to the model, which helps to select the most important variables explaining a given business question. At the same time it avoids focusing on data that contains no information.

Models can directly be applied in a simulation mode for a single input dataset predicting the score for an individual business question in real time.

Benefits for the Data Mining expert: K2R frees time for Data Mining professionals to apply their expertise in areas where they add more value instead of spending several days to tune a model. K2R produces results within minutes (less than 15 seconds on a laptop with 50,000 lines and 20 variables).

Here is a case study from the company itself.

Marketing campaign usage scenario

* Send a “Test mailing” to 5000 customers to offer them a new product,
* Collect the results of your test mailing to build a “Training” data set that associates things you know about customers prior to the mailing with the answers to your business question
* Train a model to “predict” the Yes/No answer
* Check the quality and robustness of your model (Ki, Kr)
* Apply the model to the 1,000,000 other customers in your database: this model associates each individual customer with a probability for answering Yes. Because you are using a robust model, the sum of probabilities is a good indicator of how many people will answer yes to this mail
* Send your mailing only to those customers with a high probability to respond positively, or use our built-in profit curves to optimize your return on the campaign

Example: Regression: Dealer evaluation usage scenario

* Collect information about the past performance of your dealers two years ago and associate how much of your product they sold 1 year ago
* Train a model to predict how much a dealer will sell based on the available information
* Check the quality and robustness of the model (Ki, Kr)
* Apply the model to all of your dealers today: the model associates each dealer with an estimation of how many products he will sell,
* Sum up the estimates to predict how much you will sell next year. This is the base line for your sales forecast.

In my next post I would include screenshots on how to build an automated regression model using KXEN.

Ajay Disclaimer- I am a consultant to KXEN for social networks.

Twitterfox- Twitter for the busy people

Here is a nice firefox plugin for people who want to start using Twitter without losing too much time. It sits nicely in one corner and gives gentle tweets – think of it as a big instant messenger, big in terms of number of followers and need to use twitter and busy in terms of time, but very nice and comfortable. The screenshot says it all and all you need to do is start using Firefox and install this from http://twitterfox.net/.

Heavily recommended for non users of Twitter who are curious on what this thing is all about—-

Screenshots courtesy  myself and the gentlepeople at http://twitterfox.net/.

screenshot-decisionstats-e280ba-dashboard-e28094-wordpress-mozilla-firefox

TwitterFox is a Firefox extension that notifies you of your friends’ tweets on Twitter.

This extension adds a tiny icon on the status bar which notifies you when your friends update their tweets. Also it has a small text input field to update your tweets.

Install TwitterFox

If you want to get updates of TwitterFox, feel free to follow @TwitterFox.

New Features and Changes in Version 1.7.7.1

  • Supported Firefox 3.1b3
  • Added a context menu to each tweets which has:
    • Copy
    • Re-tweet
    • Open this tweet in new tab
    • Delete tweet
  • Auto extract is.gd and bit.ly links.
  • Added Mark all as read menu item to main context menu.
  • Increased contrast of background color of read/unread messages.
  • Added in-reply-to-status-id parameter for status update.
  • Added da-DK, th-TH, vi-VN, ar-SA, ar, and kw-GB translations.
  • Bug fixes.

Does Twitter reduce Blogging ?

One more post on Twitter you may sigh, but wait. I am examine Twitter as an economic complementary  or substitute product to Blogging and trying to come up with a mathematical proving rule to dis prove the Null Hypothesis-

Twitter does not affect blogging of individuals or communities as  a whole. or does it ?

Twitter reduces blogging because

  1. Twitter is easier to do. Creating a blog is different ball game.

  2. Tweeting is two way and interactive while Blogging is mostly a one way broadcast.

  3. People respond to Tweets and re tweet them much more than they comment or forward blog posts. This is due to the inherent design of the softwares.

  4. Twitter is chaotic, but so is real life in which human brain processes different information from people like collegues, family, friends and sorts them. Blogging has a structure which helps the reader more than the writer

  5. It is easier to tweet and faster to get your point across than in Blogging.

  6. People allocate a set amount of time for social media activities and personal branding. Now this may be elastic but not totally so. Hence the rise of twitter time in people’ lives would mean lesser time to read and write blogs.

Now to a more quantitative study.

We get statistics from Technocrati – State of the Blogosphere and add in WordPress Stats to boot.

(credit -http://technorati.com/blogging/state-of-the-blogosphere/ )

A chart of total WordPress.com blogs since  launch:

(credit- http://en.wordpress.com/stats/ )

Note new signups can be seen for WordPress.com at http://en.wordpress.com/stats/signups/

Fatigue could be a reason why Twitter is hotting up while Blogging sees steady state growth.

The following figure from Technocrati’s 2008 report sums it best.

http://technorati.com/blogging/state-of-the-blogosphere/who-are-the-bloggers/

But if I compare June 2008 numbers of Blogging Frequency with the 2007 report – I am not able to compare the numbers

(Source -http://technorati.com/blogging/state-of-the-blogosphere/the-how-of-blogging/ )

http://www.sifry.com/alerts/archives/000493.html

It seems that Blog posts did get a boost with the 2008 elections and the current low traffic may simply be due to a lack of issues in Blogosphere. The rise in Twitter traffic is also due to creation of applications by third party providers and this trend has led to Twitter being the number 3 social media site.

Based on the data, it does not seem Twitter reduces Blog posts to a significant degree. After all Twitter is also a great medium to disseminate or spread the word on good blog posts.

It is simply too early to say that Twitter is reducing blogging though there seem clear trends along that line.

What about you ? If you were a blogger, is  your blog post frequency affected by your tweeting activities.