<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>web-analytics on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/web-analytics/</link>
    <description>Recent content in web-analytics on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Fri, 12 Jun 2009 19:14:01 +0000</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/web-analytics/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The Bing effect</title>
      <link>https://www.s-anand.net/blog/the-bing-effect/</link>
      <pubDate>Fri, 12 Jun 2009 19:14:01 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/the-bing-effect/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/assets/flickr-bing_3620265928_o-png.webp&#34; title=&#34;bing.com referral statistics&#34;&gt;&lt;img alt=&#34;bing.com referral statistics&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/flickr-bing_3620265928_o-png.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This graph is the number of referrals Microsoft’s new search engine, &lt;a href=&#34;http://www.bing.com/&#34;&gt;Bing&lt;/a&gt;, sent to my site over the last few days. Looks like the hype is dying out. Though &lt;a href=&#34;http://www.techcrunch.com/2009/06/05/did-bing-just-leapfrog-yahoo-search/&#34;&gt;Bing did leapfrog Yahoo&lt;/a&gt; briefly, that &lt;a href=&#34;http://www.techcrunch.com/2009/06/07/quick-peak-bings-reign-as-2-search-engine-lasted-one-day/&#34;&gt;lasted just one day&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Short notes</title>
      <link>https://www.s-anand.net/blog/short-notes/</link>
      <pubDate>Wed, 20 May 2009 05:15:32 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/short-notes/</guid>
      <description>&lt;p&gt;I&amp;rsquo;m quite busy on a project right now, and don&amp;rsquo;t get time to write long articles. So for a while, I&amp;rsquo;m going to stick to short notes on interesting stuff.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Peter Bregman has a very interesting piece on &lt;a href=&#34;http://blogs.harvardbusiness.org/bregman/2009/05/why-you-should-encourage-weakn.html&#34;&gt;Why You Should Encourage Weakness&lt;/a&gt;. It boils down to a choice: &lt;strong&gt;do you focus on on improving strengths or minimising weaknesses?&lt;/strong&gt; Conventional performance evaluations focus on the latter. I very strongly support Bregman’s view on this. &lt;strong&gt;The weakness isn’t why you hired the person!&lt;/strong&gt; Unless it’s killing the organisation, just leave them to focus on their strengths.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.google.com/analytics/&#34;&gt;Google Analytics&lt;/a&gt; has a fairly interesting &lt;a href=&#34;http://code.google.com/apis/analytics/docs/&#34;&gt;API&lt;/a&gt; that I hadn’t explored until recently. Picked up [Advanced Web Metrics with Google Analytics](&lt;a href=&#34;http://www.s-anand.net/amazon-browser.html#advanced&#34;&gt;http://www.s-anand.net/amazon-browser.html#advanced&lt;/a&gt; web metrics with google analytics) and learnt that you can track outbound clicks, page load times, Javascript events and error logs, almost anything at all using Google Analytics. You can also mirror the logging on your local server using pageTracker._setLocalRemoteServerMode()&lt;/li&gt;
&lt;li&gt;The whole concept of a Sandbox environment seems to be picking up within Google. There’s a &lt;a href=&#34;https://sandbox.google.com/checkout/&#34;&gt;Checkout sandbox&lt;/a&gt;, an &lt;a href=&#34;http://code.google.com/apis/ajax/playground/&#34;&gt;AJAX API playground&lt;/a&gt;, an &lt;a href=&#34;http://adwordsapi.blogspot.com/2009/05/adwords-api-on-app-engine-python.html&#34;&gt;AdWords sandbox&lt;/a&gt;, an &lt;a href=&#34;http://code.google.com/apis/adsense/developer/adsense_api_sandbox.html&#34;&gt;AdSense API sandbox&lt;/a&gt;, the &lt;a href=&#34;http://mapstraction.appspot.com/&#34;&gt;Mapstraction API sandbox&lt;/a&gt;, even an event called &lt;a href=&#34;http://code.google.com/events/io/sandbox.html&#34;&gt;Developer Sandbox&lt;/a&gt;. (After saying Sandbox 6 times, I feel a bit like Hobbes.)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;http://picayune.uclick.com/comics/ch/1992/ch920623.gif&#34;&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Mobile browsing</title>
      <link>https://www.s-anand.net/blog/mobile-browsing/</link>
      <pubDate>Wed, 03 Sep 2008 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/mobile-browsing/</guid>
      <description>&lt;p&gt;When I &lt;a href=&#34;https://www.s-anand.net/blog/attack-of-the-bots/&#34;&gt;analysed my HTTP log last week&lt;/a&gt;, I had another motive: are there enough people accessing my site on a mobile device? Or is it too small at this stage for me to care about?&lt;/p&gt;
&lt;p&gt;Well, have a look at the numbers.&lt;/p&gt;
&lt;table&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Windows&lt;/td&gt;&lt;td&gt;98.4%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Mobile&lt;/td&gt;&lt;td&gt;0.6%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Linux&lt;/td&gt;&lt;td&gt;0.5%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;OS X&lt;/td&gt;&lt;td&gt;0.5%&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Yes, there are more people accessing my site through a mobile device than there are using Linux or OS X. That&#39;s &lt;b&gt;shocking&lt;/b&gt;!&lt;/p&gt;
&lt;p&gt;Now, I&#39;m not saying that this is representative of the rest of the world or anything, but at least it tells me a couple of things.&lt;/p&gt;
&lt;p&gt;Firstly, the whole mobile browsing thing is bigger than I thought it was. I started worrying about this a couple of months ago and got myself a &lt;a href=&#34;http://www.google.com/search?q=HTC+s620&#34;&gt;HTC s620&lt;/a&gt; phone and a &lt;a href=&#34;http://www.google.com/search?q=BlackBerry+8320&#34;&gt;BlackBerry&lt;/a&gt; (for free, through some innovative &lt;a href=&#34;http://en.wikipedia.org/wiki/Social_engineering_(security)&#34;&gt;social engineering&lt;/a&gt; and smooth talking). It really does get pretty useful on the move... which is frankly anywhere outside of the home and the office, and sometimes even within. (It&#39;s handier to read recipes off the HTC than a laptop.) Google had caught on to the whole mobile browsing trend a very long time ago, and are rather well positioned to make use of it. &lt;/p&gt;
&lt;p&gt;Secondly, it means that rather than worrying about my site working on Linux or OS X (i.e. worrying about what plugins to use), I should worry more about it working on mobile devices (i.e. small screen, no Javascript / CSS). &lt;/p&gt;
&lt;p&gt;That&#39;s a fairly big shift in my thinking. Earlier, I had been all for shifting all the processing to client-side Javascript. Now it appears I need to design more towards plain HTML pages generated by Perl / PHP.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dinakaran&lt;/strong&gt; &lt;em&gt;3 Sep 2008 6:23 am&lt;/em&gt;:
Nice to see that people visit your site your mobile. Mobile browsing is a very upcoming trend . It&amp;rsquo;s yet to catch up in India for the telecom infrastructure limitations.If there are better band width&amp;rsquo;s provided by the service providers , mobile browsing will move up to a new level.&lt;br&gt;
&lt;br&gt;
A day when I&amp;rsquo;m able to play 30 minutes of non-interupted podcast through my mobile while travelling down to the office listening to the latest technology discussion is some thing which I dream to come as a reality in the near future.&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>Attack of the bots</title>
      <link>https://www.s-anand.net/blog/attack-of-the-bots/</link>
      <pubDate>Sun, 31 Aug 2008 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/attack-of-the-bots/</guid>
      <description>&lt;p&gt;One out of every 5 hits to my site is from a &lt;a href=&#34;http://en.wikipedia.org/wiki/Internet_bot&#34;&gt;bot&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I spent a fair bit of time this weekend analysing my log file for last month (which runs to gigabytes, and I ended up learning a few things about file system optimisation, but more on that later). 80% of the hits were from regular browsers. 20% were from robots. Here&#39;s a sample of the user-agents:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Mozilla/5.0 (compatible; Yahoo! Slurp; &amp;lt;a href=&amp;#34;http://help.yahoo.com/help/us/ysearch/slurp)&amp;#34;&amp;gt;http://help.yahoo.com/help/us/ysearch/slurp)&amp;lt;/a&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Mozilla/5.0 (compatible; Googlebot/2.1; +&amp;lt;a href=&amp;#34;http://www.google.com/bot.html)&amp;#34;&amp;gt;http://www.google.com/bot.html)&amp;lt;/a&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Mediapartners-Google
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;DotBot/1.0.1 (&amp;lt;a href=&amp;#34;http://www.dotnetdotcom.org/#info&amp;#34;&amp;gt;http://www.dotnetdotcom.org/#info&amp;lt;/a&amp;gt;, crawler@dotnetdotcom.org)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Mozilla/5.0 (Twiceler-0.9 &amp;lt;a href=&amp;#34;http://www.cuill.com/twiceler/robot.html)&amp;#34;&amp;gt;http://www.cuill.com/twiceler/robot.html)&amp;lt;/a&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;msnbot/1.1 (+&amp;lt;a href=&amp;#34;http://search.msn.com/msnbot.htm)&amp;#34;&amp;gt;http://search.msn.com/msnbot.htm)&amp;lt;/a&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;FeedBurner/1.0 (&amp;lt;a href=&amp;#34;http://www.FeedBurner.com)&amp;#34;&amp;gt;http://www.FeedBurner.com)&amp;lt;/a&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Mozilla/5.0 (compatible; attributor/1.13.2 +&amp;lt;a href=&amp;#34;http://www.attributor.com)&amp;#34;&amp;gt;http://www.attributor.com)&amp;lt;/a&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;WebAlta Crawler/2.0 (&amp;lt;a href=&amp;#34;http://www.webalta.net/ru/about_webmaster.html)&amp;#34;&amp;gt;http://www.webalta.net/ru/about_webmaster.html)&amp;lt;/a&amp;gt; (Windows; U; Windows NT 5.1; ru-RU)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Yandex/1.01.001 (compatible; Win16; I)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;You get the idea. The bulk of these are search engines. Over two-thirds of the bot requests were from &lt;a href=&#34;http://help.yahoo.com/l/us/yahoo/search/webcrawler/&#34;&gt;Yahoo Slurp&lt;/a&gt;. Now, this struck me as weird. If I take the top 3 search engines that are sending traffic my way, &lt;/p&gt;
&lt;table&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&amp;#160;&lt;/td&gt;&lt;td&gt;Referral %&lt;/td&gt;&lt;td&gt;Crawl %&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://www.google.com/&#34;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;td&gt;90%&lt;/td&gt;&lt;td&gt;24%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://search.yahoo.com/&#34;&gt;Yahoo&lt;/a&gt;&lt;/td&gt;&lt;td&gt;6%&lt;/td&gt;&lt;td&gt;66%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://search.live.com/&#34;&gt;Microsoft&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3%&lt;/td&gt;&lt;td&gt;0.3%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Others&lt;/td&gt;&lt;td&gt;1%&lt;/td&gt;&lt;td&gt;9%&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The search engine that sends me the most traffic is being reasonably conservative, while Yahoo is just eating up the bandwidth on my site. Actually, this shouldn&#39;t bother me too much. It&#39;s not taking up too much bandwidth, or even CPU usage, given that all the bots put together make up only 20% of my traffic. But somehow... it&#39;s sub-optimal. Inelegant, even.&lt;/p&gt;
&lt;p&gt;So I decided to take a closer look. Just how often are they crawling my site?&lt;/p&gt;
&lt;table&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://search.yahoo.com/&#34;&gt;Yahoo&lt;/a&gt;&lt;/td&gt;&lt;td&gt;&lt;b&gt;&lt;i&gt;Every 5 seconds&lt;/i&gt;&lt;/b&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://www.google.com/&#34;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Every 13 seconds&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://www.dotnetdotcom.org/#info&#34;&gt;DotBot&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Every 9 minutes&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://www.cuill.com/twiceler/robot.html&#34;&gt;Cuill&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Every 9 minutes&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://search.live.com/&#34;&gt;Microsoft&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Every 18 minutes&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://www.FeedBurner.com&#34;&gt;Feedburner&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Every 18 minutes&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://www.attributor.com/&#34;&gt;Attributor&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Every 23 minutes&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;a href=&#34;http://www.yandex.ru/&#34;&gt;Yandex&lt;/a&gt;&lt;/td&gt;&lt;td&gt;Every 27 minutes&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Look at those numbers. Yahoo is hitting my site once &lt;b&gt;&lt;i&gt;every 5 seconds&lt;/i&gt;&lt;/b&gt;. No wonder there&#39;s a help page at Yahoo titled &lt;a href=&#34;http://help.yahoo.com/l/us/yahoo/search/webcrawler/slurp-03.html&#34;&gt;How can I reduce the number of requests you make on my web site?&lt;/a&gt; I followed their advice and set the crawl-delay to 60, so at least it slows down to once a minute. &lt;/p&gt;
&lt;p&gt;Just that one little line change should (hopefully) reduce the load on my site by around 15%.&lt;/p&gt;
&lt;p&gt;As for the other engines, I don&#39;t mind that much in terms of load.&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Google, for all that it crawls every 13 seconds, has faithfully reported that it has only 11% of my site under its index, so I&#39;ve no idea what they&#39;re doing, but I&#39;m not complaining about the traffic that&#39;s coming my way.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&#34;http://www.dotnetdotcom.org/#info&#34;&gt;DotBot&lt;/a&gt;. Today was the first I&#39;d heard of them. Visited the site, and smiled. These guys can do all the crawling of my site that they like, and I hope something interesting comes out of their work.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&#34;http://www.cuill.com/twiceler/robot.html&#34;&gt;Cuill&lt;/a&gt;, sends me 0.2% of my traffic, but it&#39;s a new search engine, I&#39;m happy to give it time.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&#34;http://search.live.com/&#34;&gt;Microsoft&lt;/a&gt;&#39;s OK, sends me a tiny stream of traffic.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&#34;http://www.FeedBurner.com&#34;&gt;Feedburner&lt;/a&gt; is just pinging my RSS feed every 18 minutes. &lt;/li&gt;
  &lt;li&gt;&lt;a href=&#34;http://www.attributor.com/&#34;&gt;Attributor&lt;/a&gt; and &lt;a href=&#34;http://www.yandex.ru/&#34;&gt;Yandex&lt;/a&gt; I&#39;m hearing of for the first time, again. Not too much load on a system, so that&#39;s OK.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;p&gt;What&#39;s amazing is the sheer number of bots out there. Last month, I counted over 600 distinct &lt;a href=&#34;http://en.wikipedia.org/wiki/User_agent&#34;&gt;user-agent&lt;/a&gt; strings just representing bots. So it&#39;s true. The Web is &lt;a href=&#34;http://en.wikipedia.org/wiki/Semantic_web#Purpose&#34;&gt;no longer just for humans&lt;/a&gt;. We do need a &lt;a href=&#34;http://en.wikipedia.org/wiki/Semantic_Web&#34;&gt;Semantic Web&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Dhar&lt;/strong&gt; &lt;em&gt;31 Aug 2008 9:09 pm&lt;/em&gt;:
Hmmm, curious as to:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Why the bots should crawl your site every 5 seconds or so?&lt;/li&gt;
&lt;li&gt;How you can find out how much of your site has been indexed by Google.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Cheers,
D.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;S Anand&lt;/strong&gt; &lt;em&gt;1 Sep 2008 12:11 am&lt;/em&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;I think Yahoo&amp;rsquo;s crawler is aggressive in any case. My site doesn&amp;rsquo;t seem to be an exception: there are a lot of threads discussing this problem.&lt;/li&gt;
&lt;li&gt;Google&amp;rsquo;s webmaster tools tells you how many URLs have been indexed from your sitemap.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>Most bookmarked pages</title>
      <link>https://www.s-anand.net/blog/most-bookmarked-pages/</link>
      <pubDate>Mon, 15 Jan 2007 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/most-bookmarked-pages/</guid>
      <description>&lt;p&gt;These are the most bookmarked pages on my site:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/&#34;&gt;My home page&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/excel_tips.html&#34;&gt;Excel tips&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/calvin/&#34;&gt;Calvin &amp;amp; Hobbes quotes&lt;/a&gt; (I typed them all)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/Indian_Torrents.html&#34;&gt;Indian torrents&lt;/a&gt; (I have a search engine for Indian torrents)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/Tamil_transliterator.html&#34;&gt;Tamil Transliterator&lt;/a&gt; (Lets you type Tamil in English)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/Tamil_songs_by_Ilayaraja_in_1980s.html&#34;&gt;Tamil songs quiz&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/Movie_quote_quiz.html&#34;&gt;Movie quote quiz&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/my_best_links.html&#34;&gt;My best links&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.s-anand.net/top_10_lists.html&#34;&gt;Top 10 lists&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;But this post is not about these links.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s about how I found this out.&lt;/p&gt;
&lt;p&gt;Think about it&amp;hellip; &lt;strong&gt;how could I know what pages have been bookmarked&lt;/strong&gt;? The browser doesn&amp;rsquo;t send any information about bookmarks.&lt;/p&gt;
&lt;p&gt;Some months ago, I moved away from Google Analytics mainly to have more control over tracking visitors. Among other things, &lt;strong&gt;I track referrers&lt;/strong&gt;. When you click on any page and go to another one, the second page knows the first page you came from. That first page is the referrer.&lt;/p&gt;
&lt;p&gt;So I know every page people clicked on to get to my site. Usually it&amp;rsquo;s Google. Sometimes it&amp;rsquo;s someone&amp;rsquo;s blog. Sometimes, it&amp;rsquo;s blank.&lt;/p&gt;
&lt;p&gt;The blank referrers either indicate that the browser has blocked the referrer page, or that the person didn&amp;rsquo;t visit any page before mine. The former is rare (less than 1%). So, realistically, &lt;strong&gt;a blank referrer has either the person typed in the URL, or bookmarked it&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;To make sure, I did a quick survey over the weekend. Those who came to my page without a referrer saw a survey form, asking where they came from. Almost all of them had either bookmarked or typed my page. The typists went directly to &lt;a href=&#34;http://www.s-anand.net/&#34;&gt;my home page&lt;/a&gt;. All the other links you see above are bookmarks.&lt;/p&gt;
&lt;p&gt;So all I had to do was count the number of hits for each page with blank referrers. That&amp;rsquo;s the list above.&lt;/p&gt;
&lt;p&gt;Absence of information can be a powerful indicator too.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Chitra&lt;/strong&gt; &lt;em&gt;16 Jan 2007 7:11 am&lt;/em&gt;:
Oho&amp;hellip;how many people did the list comprise of?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;S Anand&lt;/strong&gt; &lt;em&gt;16 Jan 2007 12:56 pm&lt;/em&gt;:
I had over 4,000 unique IP addresses, but I don&amp;rsquo;t believe that&amp;rsquo;s even representative of visitors, thanks to most ISPs using dynamic IPs. I get about 80 direct IP hits per day, so I guess the number of bookmarks is around twice that.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rupesh&lt;/strong&gt; &lt;em&gt;17 Jan 2007 9:41 am&lt;/em&gt;:
Would you get to know if someone has subscribed to your site feed this way?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;S Anand&lt;/strong&gt; &lt;em&gt;17 Jan 2007 12:06 pm&lt;/em&gt;:
Actually, I use Feedburner to track subscriptions to my site. But no, using this I can&amp;rsquo;t always tell if it&amp;rsquo;s a feed subscription. People coming through live bookmarks can&amp;rsquo;t be distinguished from normal bookmarks. But people using feed readers (Google Reader / Bloglines) can be identified by their referrer.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shreyas&lt;/strong&gt; &lt;em&gt;19 Jan 2007 5:46 am&lt;/em&gt;:
Apologies for a completely off-topic comment but have you taken off the &amp;lsquo;Starred Items on Google Reader&amp;rsquo; link? I kind of enjoyed reading the ones you had a couple of weeks back!&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>Demographics prediction from online behaviour</title>
      <link>https://www.s-anand.net/blog/demographics-prediction-from-online-behaviour/</link>
      <pubDate>Mon, 26 Jun 2006 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/demographics-prediction-from-online-behaviour/</guid>
      <description>&lt;p&gt;Microsoft adCenter Labs has a &lt;a href=&#34;http://adlab.microsoft.com/DPUI/DPUI.aspx&#34;&gt;demographics prediction engine&lt;/a&gt;. Based on a person&#39;s search queries and web sites visited, it can predict their gender and age.&lt;/p&gt;
&lt;p&gt;So I tried that on parts of the body, to see what men were interested in vs women.&lt;/p&gt;
&lt;table class=&#34;numbers lines&#34;&gt;
&lt;tr&gt;&lt;th&gt;topic&lt;/th&gt;&lt;th&gt;male&lt;/th&gt;&lt;th&gt;female&lt;/th&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;hair&lt;/td&gt;&lt;td&gt;25%&lt;/td&gt;&lt;td&gt;75%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;eyes&lt;/td&gt;&lt;td&gt;33%&lt;/td&gt;&lt;td&gt;67%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;cheek&lt;/td&gt;&lt;td&gt;33%&lt;/td&gt;&lt;td&gt;67%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;hands&lt;/td&gt;&lt;td&gt;33%&lt;/td&gt;&lt;td&gt;67%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;lips&lt;/td&gt;&lt;td&gt;36%&lt;/td&gt;&lt;td&gt;64%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;ears&lt;/td&gt;&lt;td&gt;39%&lt;/td&gt;&lt;td&gt;61%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;fingers&lt;/td&gt;&lt;td&gt;40%&lt;/td&gt;&lt;td&gt;60%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;forehead&lt;/td&gt;&lt;td&gt;42%&lt;/td&gt;&lt;td&gt;58%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;nose&lt;/td&gt;&lt;td&gt;43%&lt;/td&gt;&lt;td&gt;57%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;neck&lt;/td&gt;&lt;td&gt;46%&lt;/td&gt;&lt;td&gt;54%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;beard&lt;/td&gt;&lt;td&gt;55%&lt;/td&gt;&lt;td&gt;45%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;moustache&lt;/td&gt;&lt;td&gt;58%&lt;/td&gt;&lt;td&gt;42%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;leg&lt;/td&gt;&lt;td&gt;60%&lt;/td&gt;&lt;td&gt;40%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;palm&lt;/td&gt;&lt;td&gt;61%&lt;/td&gt;&lt;td&gt;39%&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;toe&lt;/td&gt;&lt;td&gt;64%&lt;/td&gt;&lt;td&gt;36%&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;
&lt;p&gt;While I can understand men being more interested in beards and moustaches (perhaps even legs), why are they far more interested in toes than women?&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dhar&lt;/strong&gt; &lt;em&gt;28 Jun 2006 2:10 am&lt;/em&gt;:
Hey Anand, a quick question.&lt;br&gt;
Are you using Google Spreadsheets? Yesterday, I realized that after the initial enthu, I have not visited the site. This makes me wonder if Google Spreadsheets in itself is a powerful enough solution to attract the masses.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dhar&lt;/strong&gt; &lt;em&gt;28 Jun 2006 2:12 am&lt;/em&gt;:
As pointed out by someone on Slashdot, try using the URL &lt;a href=&#34;http://www.google.com/&#34;&gt;http://www.google.com/&lt;/a&gt; as the input and you will get &amp;ldquo;Female Oriented with following Confidence: Female 1.00, Male 0.00&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;S Anand&lt;/strong&gt; &lt;em&gt;28 Jun 2006 2:45 pm&lt;/em&gt;:
I do use Google spreadsheets. Just see my next post.&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>Search queries to my site</title>
      <link>https://www.s-anand.net/blog/search-queries-to-my-site/</link>
      <pubDate>Thu, 06 Apr 2006 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/search-queries-to-my-site/</guid>
      <description>&lt;p&gt;On a related note, 60% of the search queries that lead to my site this year were &lt;a href=&#34;https://www.s-anand.net/blog/calvin/&#34;&gt;Calvin and Hobbes quotes&lt;/a&gt;. &amp;ldquo;i can&amp;rsquo;t help but wonder what kind of desperate straits would drive a man to invent this thing.&amp;rdquo; topped the list (Calvin referring to a yo-yo), with &lt;a href=&#34;http://www.google.com/search?q=%22i+always+catch+these+trick+questions%22&#34;&gt;i always catch these trick questions&lt;/a&gt; following closely.&lt;/p&gt;
&lt;p&gt;People searching for Excel related stuff were next (20%): &lt;a href=&#34;http://www.google.com/search?q=excel+%22indirect%28address%28%22&#34;&gt;excel indirect(address(&lt;/a&gt;, &lt;a href=&#34;http://www.google.com/search?q=row%28%29+excel+offset+address&#34;&gt;row() excel offset address&lt;/a&gt; and the like.&lt;/p&gt;
&lt;p&gt;A few were also looking for me by name or school (10%).&lt;/p&gt;
&lt;p&gt;The last 10% ranged from the puzzling to the bizarre, including these gems.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;http://www.google.com/search?q=michalengelo+hidden+skull&#34;&gt;michalengelo hidden skull&lt;/a&gt;. Probably looking for the alleged hidden skull in Last Judgement.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.google.com/search?q=googlemail+access+between+england+and+india&#34;&gt;googlemail access between england and india&lt;/a&gt;. Why? Did he think there wouldn&amp;rsquo;t be any?&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.google.com/search?q=greenwich+meridien+time+for+india&#34;&gt;greenwich meridien time for india&lt;/a&gt;. This is usually the same in Greenwich and in India. Rest of the world too.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.google.com/search?q=origin+of+monkey+in+fez&#34;&gt;origin of monkey in fez&lt;/a&gt;. What?&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.google.com/search?q=address+of+sexy+girl+in+ahmedabad&#34;&gt;address of sexy girl in ahmedabad&lt;/a&gt;. But why &lt;strong&gt;my&lt;/strong&gt; site?&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;http://www.google.com/search?q=address+of+tool+makers+for+converting+html+to+xml+in+chennai&#34;&gt;address of tool makers for converting html to xml in chennai&lt;/a&gt;. Do they have a license to convert?&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Sai&lt;/strong&gt; &lt;em&gt;11 Apr 2006 4:55 pm&lt;/em&gt;:
Really funny! What kind of metics/tools do you have to know what searches lead to your site?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;S Anand&lt;/strong&gt; &lt;em&gt;11 Apr 2006 7:44 pm&lt;/em&gt;:
I use Google Analytics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ritzkini&lt;/strong&gt; &lt;em&gt;14 Apr 2006 5:37 pm&lt;/em&gt;:
address of sexy girl in ahmedabad !! :))&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Krishna&lt;/strong&gt; &lt;em&gt;15 Apr 2006 5:57 pm&lt;/em&gt;:
Thats a real funny set of queries. But why your site? ;-)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;http://www.s-anand.net/blog/the-calvin-and-hobbes-search-takedown/&#34;&gt;The Calvin and Hobbes search Takedown | s-anand.net&lt;/a&gt;&lt;/strong&gt; &lt;em&gt;21 May 2010 12:02 pm&lt;/em&gt; &lt;em&gt;(pingback)&lt;/em&gt;:
[&amp;hellip;] also increased traffic to my site, which was a bit disconcerting. I didn’t want to attract attention. In 2007, I [&amp;hellip;]&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>No site has linked to my page</title>
      <link>https://www.s-anand.net/blog/no-site-has-linked-to-my-page/</link>
      <pubDate>Fri, 31 May 2002 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/no-site-has-linked-to-my-page/</guid>
      <description>&lt;p&gt;Almost &lt;a href=&#34;http://www.google.com/search?as_lq=www.geocities.com%2Froot_node&#34;&gt;no site has linked to my page&lt;/a&gt;. (I sort of like that a bit. It means the 30-odd hits a day that I get are from people who TYPE my URL.)&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Wednesday is popular</title>
      <link>https://www.s-anand.net/blog/wednesday-is-popular/</link>
      <pubDate>Sun, 29 Jul 2001 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/wednesday-is-popular/</guid>
      <description>&lt;p&gt;Wednesday seems the most popular day for visiting my site. While I get 14 visitors a day on average, I seem to get 22 visitors on Wednesdays.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Weblog analyser</title>
      <link>https://www.s-anand.net/blog/weblog-analyser/</link>
      <pubDate>Thu, 01 Feb 2001 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/weblog-analyser/</guid>
      <description>&lt;p&gt;I&amp;rsquo;m using &lt;a href=&#34;http://www.analog.cx/&#34;&gt;Analog&lt;/a&gt; for analysing the IIMB webserver stats. Anyone with enthu, find a better log analyzer and mail me its name.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Popular sites at IIMB</title>
      <link>https://www.s-anand.net/blog/popular-sites-at-iimb/</link>
      <pubDate>Mon, 10 Jul 2000 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/popular-sites-at-iimb/</guid>
      <description>&lt;p&gt;Whose sites are the most popular on this server? Last month, the top 10 sites were admissions, placement, Prof. Srinivasan&amp;rsquo;s, and Prof. Bandi&amp;rsquo;s. I ran &lt;a href=&#34;http://www.netstore.de/Supply/http-analyze/&#34;&gt;http-analyze&lt;/a&gt; to get these results.&lt;/p&gt;
</description>
    </item>
  </channel>
</rss>
