Friday, January 19, 2007

Standard Deviation When lnterpreting Web Analytics


One of my favorite blogs, Good Math, Bad Math has an excellent article describing standard deviation as it relates to the mean of the data.

The mean is more commonly called the average. It's calculated by the sum of the total data points in the population, then divided by the number of data points. A simple example of that would be the data set: (1, 2, 3, 4, 5). The sum of this population equals 15. There are 5 data points in this population, so the mean would be calculated as 15/5=3.

A fancier way to put it would be the following formula:



We can see through web analytics the average visitors per period of time... daily, weekly, monthly. However, measuring by mean alone can be deceptive. The mean doesnt describe some of the more important data sets that are important in determining the meaning of analytics.

The mean wont give you information on the low points, the high points, nor will they tell you the relationship of the mean between the rest of the data. In the earlier data set, the relationship between the data was very easy to determine. In web analytics, those relationships can be a little trickier.

I'll take an example that's near to me. My own web analytics.

Currently, I average 42 page views per day. This means that 42 of my unique pages are viewed... this is not a site visit. My low point is 4 page views in a single day and my highest is 124.

From this, we can tell that there was most likely a spike in my page views at some point. Because the mean is less than twice the largest data point, we can automatically start with that presumption. However, in order to get more information, we must take into account the standard deviation. It is defined as a measure of the spread of its values also, the square root of the variance. (From Wikipedia)

Each differently colored area is the standard deviation. Each section is the same length, but not the same area under the curve. This means that within one standard deviation of the mean, most of the data falls under those data points.

In my case, my standard deviation is calculated as 9.4.

What this means, is that using Chebyshev's Inequality rule,

At least 50% of the values are within 1.4 standard deviations from the mean.
At least 75% of the values are within 2 standard deviations from the mean.
At least 89% of the values are within 3 standard deviations from the mean.
At least 94% of the values are within 4 standard deviations from the mean.
At least 96% of the values are within 5 standard deviations from the mean.
At least 97% of the values are within 6 standard deviations from the mean.
At least 98% of the values are within 7 standard deviations from the mean.
At least 1 - 1/k2 of the values are within k standard deviations from the mean.
When you apply this information to web analytics, one of the things I do is look at the geographic distribution of the users. When I find hubs of higher consumer acitivity, I start getting a clearer idea to who my users are. This could help me target my paid search campaign more accurately, this could let me know that if I provide content, analysis or a blog, a nice mention of something applicable and interesting in their area might be appropriate.

The standard deviation is a powerful method to segment your analytics into greater specificity. When Chebyshev's Inequality shows you that 75% of the data is within two standard deviations, then you have some focused and applicable data to improve your messaging and targeting.

Wednesday, January 17, 2007

More on Evidence Based Managment

I love this cartoon, and as I go through Pfeffer and Suttons' book, there are parts of the book that remind me of this.

This cartoon has been used to entertain math geeks, critique pseudoscience like Intelligent Design, and brilliantly show that knowledge starts with data and ends with something useful (hopefully). But somewhere in the middle is the arduous task of analysis, being wrong, being frustrated, hard work, peer review and all of the un-sexy stuff that gets ignored.

It's kind of a small rant, but it's the "CSI-ification" of science or analysis... they get a clue, then some fancy camera shots and a high tech lab... chemicals and viola! They get the information they need to catch the bad guy.

The miracle is hard work... it's the analysis, it's the brainstorming of lots of smart people trying to figure it out. To call it a miracle is almost an insult... and that's why I love this cartoon

Some Cliches Become Even Truthier

From Eric Mattson's Marketing Blurb, he writes about an interesting study from Susquehanna about the trends of content. In Eric's article, he mentions that in 2005, content sites accounted for 36% of the users' time. However, in 2006, that trend went upwards to 45%.

The interesting point here is where that trend is going to go, and where it will reach a critical mass. In 2007, will the numbers be somewhere from 55 - 60% or will the trend start to slow down, even out and become more static?

Another interesting aspect of this is how do marketers create content for users? Rather than having marketers jump into the latest fad, they actually need to have a strategy that's a little more in depth than slapping together a blog. Garrett French's blog has an excellent market conversation strategy guide.

Understanding that content sites are increasing in popularity, and putting together a strategy to communicate with your users could absolutely bring in business, but even more valuable... a relationship with the content providers and the company itself.

Hard Facts: First Section Review

This book is organized into three sections, for this post, I will be talking about the section called "setting the stage" in Pfeffer and Sutton's book.

Pfeffer and Sutton start the book by describing what evidence based decision making is. It's essentially defined as a process to find the best evidence that you can. Through primary or secondary research, collecting the data and acting on that data. They say that there is an inherent perception blindness when making decisions from what you've always done, what you thought was true, your personal philosophies or beliefs and what ever fad is gripping the business world at the moment. I was impressed with the description of evidence based management or decision making is not a "thing you do", it doesnt have discreet boundaries, it's a process, it's a way to make decisions with a little data to help you out.

One of the examples they explore is how mergers and acquisitions tend to show strain about a month after the merger. Cisco has a great record for mergers because they measure several aspects for its merger targets, not just product or service, not just market share.. but culture as well. They've walked away from deals when the culture didn't match.

Pfeffer and Sutton identify common problems that consistently cause failure.

  • Casual benchmarking
  • Repeating what's worked in the past, or what's worked for others
  • Following deeply held, yet unexamined ideologies
  • Substituting facts for conventional wisdom
Each of these common problems can cause failure in strategies because they either ignore the data that's there, they fail to take into account new, unmeasured data or they assume that their belief is enough to make decisions. One of my favorite quotes (which I have in my quote generator) is David Hume's quote: "A wise man proportions his belief to the evidence". It's not wrong to believe, but in business and strategy, you need to have evidence to support those beliefs.

They provide logical and interesting anecdotes (which, in and of themselves are not evidence) that elucidate some of the principles, and they do an excellent job in documentation and referencing the examples they provide.

In my niche market of "search intelligence", I live and breathe data. Whether it's analytics or mined data, I try to be very careful to either only say what I can prove, or qualify any statement that has more intuition than data.

I just finished the second section, and without finishing it yet, I can say... buy this book. Read it, and let me know what you think. It's good.

Sunday, January 14, 2007

Section by Section Book Review: Pfeffer and Sutton's: Hard Facts

I've started to hit the books hard, books on search, competitive intelligence, analytics, data gathering, analysis and even math. The more I read... the more the world of data opens up into an amazing pattern.

Rather than stay in my little reading hole, As I read the books, I'm going to do a section by section review of what I've learned.

The first book I'm going to take a look at is Jeffery Pfeffer and Robert Sutton's book: Hard Facts - Dangerous Half-Truths & Total Nonsense (BN, Amazon). The book is about examining the pre-suppositions, assumptions and ingrained beliefs that managers, analysts and decision makers face when making decisions about strategies and processes that affect their business.

Pfeffer and Sutton take evidence based management methodologies and deconstruct the myths and assumptions, and they take close aim to the more dangerous half-truths and faddish business mantras. Already, I've read the first few chapters and I've been impressed with the skill in which they dissect some of the all too common axioms and slogans that populate business training.

Stay tuned, I will be doing another post on the book later. So far, I'm enjoying it.

Saturday, January 13, 2007

Privacy...Schmivacy... Balancing Personal Information with Service

A recent eMarketer article - "Is Privacy Overrated?" shares some interesting information about a "Personalization Survey" from ChoiceStream. They report that giving up personal information for more personalized content has a pendular affect.

In 2004, 63% of people who were 18-34 were willing to divulge Demographic information in order to obtain personalized content. In 2005, that number dropped to 47%. In 2006, it rebounded back to 63%. When we're talking about people over 35, the numbers are 2004: 49%, 2005: 46% and 2006: 54%. When we put things into context, 2005 was a pretty sketchy year for personal information, privacy invasion concerns and other geo-political events that made people a little more nervous about giving up their information, but apparently, they got over it and are now giving up more information.

The other part of that survey was about people giving up personalized information to a site so that it could track their clicks and purchases.

For 18-34 year olds, in 2004 it was 48%. This dropped in 2005 to 35% and back up for 2006 to 49%. The over 35 crowd remained a bit more skeptical, in 2004 it was 33%, which took a slight hit in 2005 to 29%, then rallied up to 38%.

The data here is very interesting, while it exposes an interesting behavioral trend, I wanted to take a look at what this means for consumers and marketers.

From the article:

"Consumers are overwhelmed with the vast array of content and choices coming at them every day online. They want guidance, even though they want the freedom to make their own choices and to explore the data if they want to," said Esther Dyson, editor of the blog Release 0.9 and an advisor to ChoiceStream.
While it's an interesting point that consumers are overwhelmed, I don't think that tells the full story.

2005 and '06 had some big stories about identity theft, MySpace stalking, data loss, government domestic observation and other news that really hit at the core of peoples' sense of security. I would find it interesting to research how online marketers changed their tactics to expressely address security concerns to their users, while at the same time, tried to understand what their users wanted so that they could provide the right value to the consumer.

Time will tell if this effect is pendular, we may see these numbers fall again, or they may continue to rise. However, i think that the lesson of privacy, personal information and trust to the online market is something that will be ever-present, and we'll also see if online marketers keep the lessons learned, or will they (if the trust continues to rise) accept the status quo.


Wednesday, January 10, 2007

Defining Search Metrics: Search Engine Presence

In an earlier post, I mentioned (without explicitly defining) the term of "search engine presence". This came from the dissatisfaction I felt when talking to clients about the health of their search engine optimization/ marketing campaigns. All too often, I would hear the same mantra..."I want to be number 1 for the term X" or "Why aren't I number 1 for the term X?"

This felt inherently wrong to me. However, I couldn't really answer their question, nor could I give them a sure-fire way to attain that position. I sometimes felt that I should be glib and say "If I knew that answer, I'd be working at Google as one of their engineers, right?"

At the same time, everyone who's been in search marketing for more than a week knows that it's best to rank well on a variety of keywords, while remaining true to the core goals of the site. After looking at some WebPosition Gold ranking reports, something struck me as odd about them, they gave the ranking reports, but the data it gave seemed too myopic. This is when I started thinking about "presence" as a metric for measuring the health of the search marketing campaign.

I went to Adam Schultz, and I proposed to him a creation of a simple program that we would later called the Competitive Analysis Baseline Reporting tool. This program would take the core, top level keywords from the client's input... adjusted and perfected by some keyword research and take a look at which sites ranked for those keywords. We wanted to get a good look at the entire search spectrum, so we took the top 15 results in Google, the top 10 in MSN and the top 10 in Yahoo!. This way we would get an overview of what I later called the search engine marketspace. It's a capture of data at a specific time of what the marketspace is.

For a practical example, lets take a few keywords... "iphone, apple iphone, ipod phone" for this example, we don't need a lot of keywords because I'm looking at defining, in a practical sense, "presence".

So, with 3 keywords and (15 Google + 10 MSN + 10 Yahoo) we can expect to have a sample size of 105 potential slots for search engine results to appear. When the same company, like Apple shows up across the search engines and at different positions, their presence is counted as 1. Each presence is counted and sorted for the total presence. In this case, the top 1o results is as follows:


1/10/2007

Domains

SE Presence

www.apple.com

7

www.thinksecret.com

7

www.engadget.com

7

www.gizmodo.com

6

www.appleinsider.com

5

www.mobilewhack.com

5

www.everythingiphone.com

5

gizmodo.com

4

en.wikipedia.org

4

news.bbc.co.uk

4

www.businessweek.com

3


For those terms, Apple.com shows up in Google, MSN and Yahoo 7 times, as does Thinksecret.com and Engadget.com. Gizmodo shows up 6 times, and so on. The idea here is not to diminish the actual position, or rank of the site, but emphasize the presence in the overall search marketspace. When your company relies on capturing qualified traffic from search, it's obviously better to have several keywords working for you, rather than focusing on only one keyword. Unfortunately, all too often, SEO/ SEM companies attract the client by either telling the prospect what they want to hear, or implying that ranking on their top keyword is paramount to success.

What we see here is a lack of education and a hype of expectations. When the client is properly educated on the strategies of SEO and SEM, they're more likely to abandon the expectation of the single keyword on top hope and adopt a more gestalt view of the search engines as an environment that changes, evolves and fluctuates. Once they see that, they'll recognize the value of having several keywords that work for them and not just one.