Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Sunday, January 28, 2007

Defining Search Metrics: Search Engine Saturation

In my last articles about search engine metrics, I defined search engine presence as "the number of times a site shows across the search engines for a selected set of keywords". Next, I defined, explained and showed an example of search engine fluctuation as "the natural fluctuation of presence of the same selected set of keywords over time".

This time, I want to talk about the two types of search engine saturation. I'd like to defne search engine saturation as essentially "presence over the total data size". In statistics, the data size is represented by the variable "n". If the data size is a dozen eggs, then n=12. The data size of states in the US is n=50. In this case, we're measuring results over three search engines. "SE" = 3. If we included Ask.com in the results, SE would equal 4. However, at this point, we're only measuring 3 search engines, Google, MSN and Yahoo!.

The next part is the total number of results tallied. Since we're measuring the top 15 ranks in Google, the top 10 in MSN and Yahoo!. With this information, we can calculate the total data size for any data capture.

n= (# of keywords) * (15 Google + 10 MSN + 10 Yahoo!).

in this case:

n= (3)*(35) = 105

To calculate the saturation, you divide each presence by "n" to get the percentage of the search engines the site occupies as a function of the entire marketspace.

Domains

SE Presence

SE Presence

SE Saturation

AVG Saturation

www.apple.com

7

17

16.19%

11.43%

www.engadget.com

7

9

8.57%

7.62%

en.wikipedia.org

4

7

6.67%

5.24%

www.gizmodo.com

6

6

5.71%

5.71%

www.appleinsider.com

5

4

3.81%

4.29%

www.mobilewhack.com

5

4

3.81%

4.29%

www.thinksecret.com

7

3

2.86%

4.76%

news.bbc.co.uk

4

3

2.86%

3.33%

www.businessweek.com

3

3

2.86%

2.86%

appleiphone.blogspot.com

0

3

2.86%

1.43%

www.macworld.com

1

3

2.86%

1.90%

www.everythingiphone.com

5

2

1.90%

3.33%

gizmodo.com

4

1

0.95%

2.38%



Apple.com has a presence of 17, and divided by 105, the search engine saturation equals 16.19%, which represents the market share of the marketspace. This number will fluctuate as the results fluctuate. The saturation measurement is useful as a snapshot of the search engine space and a result of your campaign. However, what's really important is the trend of data. You want your presence to rise and you will want your average saturation to rise as well. The average saturation measures the health of the life of the campaign as the raw numbers of the presence fluctuate. In essence, it's a measurement that you can measure and quantify to see if you're doing well, or if you're trending down.

Each of these metrics have value in and of themselves, however, when taken as a whole, they start to give you a clearer picture of the life of the natural search campaign.

Tuesday, January 23, 2007

Defining Search Metrics: Search Engine Fluctuation

LeeAnn Prescott, the research director for the US markets at Hitwise,
revealed that after Steve Jobs unveiled the iPhone at MacWorld, that the search demand for the iPhone has superceded the search demand for the iPod. It's only a slight coincidence that earlier I defined the concept of "search engine presence" using the example of the iPhone.

That search engine presence measured the top ranking sites, ignoring position, instead focusing on ranking on a variety of terms, that dominate the search engines for a selected sampling of keywords.

Because I was explaining a concept, I chose to use a small sample of keywords that related to a topic that interested me. I used: "iphone, apple iphone, ipod phone". The results were as follows.


1/10/2007

Domains

SE Presence

www.apple.com

7

www.thinksecret.com

7

www.engadget.com

7

www.gizmodo.com

6

www.appleinsider.com

5

www.mobilewhack.com

5

www.everythingiphone.com

5

gizmodo.com

4

en.wikipedia.org

4

news.bbc.co.uk

4

www.businessweek.com

3


These sites ranked for the preceding terms on 10 Jan. However, for the 4 weeks that ended on 20 Jan, Hitwise compiled the following top sites that had traffic from the term "iPhone".

Apple's sites receive over 50% of the traffic from the iPhone search, however, Engadget, by being one of the most trusted sites on consumer electronics and also having the benefit of ranking well across the board for the searches, received the 4th highest amount of traffic. Hitwise and my search engine presence tool measures two different things, I can only measure presence, Hitwise can measure traffic. While I know that I'm comparing apples to oranges, here's where we see the overlap of traffic and presence.

In the 13 days where I took my first measurements, the search engine marketspace has fluctuated. Supporting the idea that traffic and clickthroughs can influence ranking in the search engines. Notice that Apple and Engadget have a significant upsurge in their presence.


1/10/2007

1/23/2007

Domains

SE Presence

SE Presence

www.apple.com

7

17

www.engadget.com

7

9

en.wikipedia.org

4

7

www.gizmodo.com

6

6

www.appleinsider.com

5

4

www.mobilewhack.com

5

4

www.thinksecret.com

7

3

news.bbc.co.uk

4

3

www.businessweek.com

3

3

appleiphone.blogspot.com

0

3

www.macworld.com

1

3

www.everythingiphone.com

5

2

gizmodo.com

4

1


We see EverythingiPhone.com, Gizmodo and ThinkSecret.com drop in their presence, we see Wikipedia increase their presence.

As the search for iPhones increase, as they have... that's created the circumstances for the search engines to continually shuffle the results, shuffle the ranks and shuffle the presence. However, what we have strong evidence for here is that ranks change, presence fluctuates and the environment changes. Therefore, it's vital for any company who has a search engine marketing strategy to keep tabs on the environment for their selected keywords and not just on the position or rank of just a few. If you dont notice the change, it's likely that you could be left behind.

While LeeAnn Prescott and I are measuring two different things, traffic vs. presence. The overlapping data as it applies to the search engines often complement each other and provide a larger picture of what's happening.

Thanks LeeAnn, your article was brilliant, informative and I enjoyed reading every bit of it.

Friday, January 19, 2007

Standard Deviation When lnterpreting Web Analytics


One of my favorite blogs, Good Math, Bad Math has an excellent article describing standard deviation as it relates to the mean of the data.

The mean is more commonly called the average. It's calculated by the sum of the total data points in the population, then divided by the number of data points. A simple example of that would be the data set: (1, 2, 3, 4, 5). The sum of this population equals 15. There are 5 data points in this population, so the mean would be calculated as 15/5=3.

A fancier way to put it would be the following formula:



We can see through web analytics the average visitors per period of time... daily, weekly, monthly. However, measuring by mean alone can be deceptive. The mean doesnt describe some of the more important data sets that are important in determining the meaning of analytics.

The mean wont give you information on the low points, the high points, nor will they tell you the relationship of the mean between the rest of the data. In the earlier data set, the relationship between the data was very easy to determine. In web analytics, those relationships can be a little trickier.

I'll take an example that's near to me. My own web analytics.

Currently, I average 42 page views per day. This means that 42 of my unique pages are viewed... this is not a site visit. My low point is 4 page views in a single day and my highest is 124.

From this, we can tell that there was most likely a spike in my page views at some point. Because the mean is less than twice the largest data point, we can automatically start with that presumption. However, in order to get more information, we must take into account the standard deviation. It is defined as a measure of the spread of its values also, the square root of the variance. (From Wikipedia)

Each differently colored area is the standard deviation. Each section is the same length, but not the same area under the curve. This means that within one standard deviation of the mean, most of the data falls under those data points.

In my case, my standard deviation is calculated as 9.4.

What this means, is that using Chebyshev's Inequality rule,

At least 50% of the values are within 1.4 standard deviations from the mean.
At least 75% of the values are within 2 standard deviations from the mean.
At least 89% of the values are within 3 standard deviations from the mean.
At least 94% of the values are within 4 standard deviations from the mean.
At least 96% of the values are within 5 standard deviations from the mean.
At least 97% of the values are within 6 standard deviations from the mean.
At least 98% of the values are within 7 standard deviations from the mean.
At least 1 - 1/k2 of the values are within k standard deviations from the mean.
When you apply this information to web analytics, one of the things I do is look at the geographic distribution of the users. When I find hubs of higher consumer acitivity, I start getting a clearer idea to who my users are. This could help me target my paid search campaign more accurately, this could let me know that if I provide content, analysis or a blog, a nice mention of something applicable and interesting in their area might be appropriate.

The standard deviation is a powerful method to segment your analytics into greater specificity. When Chebyshev's Inequality shows you that 75% of the data is within two standard deviations, then you have some focused and applicable data to improve your messaging and targeting.