<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>probability on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/probability/</link>
    <description>Recent content in probability on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 11 Oct 2010 14:53:01 +0000</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/probability/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Bayes’ Theorem</title>
      <link>https://www.s-anand.net/blog/bayes-theorem/</link>
      <pubDate>Mon, 11 Oct 2010 14:53:01 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/bayes-theorem/</guid>
      <description>&lt;p&gt;I’ve tried understanding &lt;a href=&#34;http://en.wikipedia.org/wiki/Bayes&#39;_theorem&#34;&gt;Bayes’ Theorem&lt;/a&gt; several times. I’ve always managed to get confused. Specifically, I’ve always wondered why it’s better than simply using the average estimate from the past. So here’s a little attempt to jog my memory the next time I forget.&lt;/p&gt;
&lt;p&gt;Q: A coin shows 5 heads when tossed 10 times. What’s the probability of a heads?&lt;br&gt;
A: It’s not 0.5. That’s the most likely estimate. The probability distribution is actually:&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/assets/bayesian1.webp&#34;&gt;&lt;img alt=&#34;dbeta(x,5,5)&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/bayesian1.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;That’s because you don’t really know the probability with which the coin will throw a heads. It could be any number p. So lets say we have a probability distribution for it, f(p).&lt;/p&gt;
&lt;p&gt;Initially, you don’t know what this probability distribution is. So assume they’re all the same – a flat function: f(p) = 1&lt;a href=&#34;https://www.s-anand.net/blog/assets/bayesian2.webp&#34;&gt;&lt;img alt=&#34;dbeta(x,1,1)&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/bayesian2.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Now, given this, let’s say a heads falls on the next toss. What’s the revised probability distribution? It’s:&lt;/p&gt;
&lt;p&gt;f(p) ← f(p) * probability(heads | x) / probability(heads) = 1 * (x^1 * (1-x)^0) / 1 = x&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/assets/bayesian3.webp&#34;&gt;&lt;img alt=&#34;dbeta(x,2,1)&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/bayesian3.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Let’s say the next is again a heads. Now it’s&lt;/p&gt;
&lt;p&gt;f(p) ← f(p) * probability(heads | x) / probability(heads) = x * (x^1 * (1-x)^0) / 1 = x^2&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/assets/bayesian4.webp&#34;&gt;&lt;img alt=&#34;dbeta(x,3,1)&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/bayesian4.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Now if it’s a tails, it becomes:&lt;/p&gt;
&lt;p&gt;f(p) ← f(p) * prob(tails | x) / prob(tails) = x^2 * (x^0 * (1-x)^1) / 1 = x^2 * (1-x)&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.s-anand.net/blog/assets/bayesian5.webp&#34;&gt;&lt;img alt=&#34;dbeta(x,3,2)&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/bayesian5.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;… and so on. (This happens to be a called a Beta distribution.)&lt;/p&gt;
&lt;p&gt;Now, instead of this being the probability of heads, it could be the probability of a person having blood pressure, or a document being spam. As you get more data, the probability distribution of the probability keeps getting revised.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Ram&lt;/strong&gt; &lt;em&gt;18 Jan 2012 1:59 pm&lt;/em&gt;:
Hi Anand,
f(p) ← f(p) * probability(heads | x) / probability(heads) = 1 * (x^1 * (1-x)^0) / 1 = x
What I understand from the above notation is f(p) stands for probability distribution function (which is used recursively), x stands for the probability of getting a head (which is unknown) and 1-x stands for the complement of x.
Can you please explain why the distribution function is defined as you have done? i.e. , f(p) ← f(p) * probability(heads | x) / probability(heads)&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>A sense of proportion</title>
      <link>https://www.s-anand.net/blog/a-sense-of-proportion/</link>
      <pubDate>Fri, 14 May 2010 08:44:30 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/a-sense-of-proportion/</guid>
      <description>&lt;p&gt;A quote from &lt;a href=&#34;http://www.youtube.com/watch?v=YczGZCTPkRE#t=20m24s&#34;&gt;David Heinemeier Hansson&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;So the problem is, a lot of business managers and especially business owners, they have no sense of probability. They can&amp;rsquo;t fathom that concept. So They treat the probability of 1 to 10 trillion as the same as a 1 to a 100. And like, &amp;ldquo;We&amp;rsquo;ve got to deal with this 1 to a trillion probability, because, what if it happens?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;No! Doesn&amp;rsquo;t matter! I mean, don&amp;rsquo;t care.&lt;/p&gt;
&lt;p&gt;So as soon as that sense of probability spreads, that people can treat that reasonably, I think all this nonsense just goes away.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This lack of proportion, sadly, is at the heart of my every day problems. (Just watch the video!)&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;elbin&lt;/strong&gt; &lt;em&gt;14 May 2010 5:49 pm&lt;/em&gt;:
Money matters!&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>Tamil spelling corrector</title>
      <link>https://www.s-anand.net/blog/tamil-spelling-corrector/</link>
      <pubDate>Tue, 01 May 2007 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/tamil-spelling-corrector/</guid>
      <description>&lt;p&gt;The Internet has a lot of tamil song lyrics in English. Finding them is not easy, though. Two problems. The lyrics are fragmented: there&amp;rsquo;s no one site to search them. And Google doesn&amp;rsquo;t help. It doesn&amp;rsquo;t know that alaipaayudhe, alaipaayuthe and alaipayuthey are the same word.&lt;/p&gt;
&lt;p&gt;This is similar to the problem I faced with &lt;a href=&#34;https://www.s-anand.net/blog/hindi-songs-online/&#34;&gt;tamil audio&lt;/a&gt;. The solution, as before, is to make an index, and provide a search interface that is tolerant of English spellings of Tamil words. But I want to go a step further. &lt;strong&gt;Is it possible to display these lyrics in Tamil?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;My &lt;a href=&#34;https://www.s-anand.net/blog/Tamil_transliterator.html&#34;&gt;Tamil Transliterator&lt;/a&gt; does a lousy job of this. Though it&amp;rsquo;s tolerant of mistakes, it&amp;rsquo;s ignorant of spelling and grammer. So,&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;kanda nal muthalai kathal peruguthadi&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&amp;hellip; becomes&amp;hellip;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;kanda nal muthalai kathal peruguthadi&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&amp;hellip; when in fact we want&amp;hellip;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;kanda naaL muthalaay kaathal peruguthadi&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(If you&amp;rsquo;re viewing this on an RSS reader, check &lt;a href=&#34;https://www.s-anand.net/blog/tamil-spelling-corrector/&#34;&gt;my post&lt;/a&gt; to see what I mean.)&lt;/p&gt;
&lt;p&gt;I need an automated Tamil spelling corrector. Reading &lt;a href=&#34;http://www.norvig.com/spell-correct.html&#34;&gt;Peter Norvig&amp;rsquo;s &amp;ldquo;How to Write a Spelling Corrector&amp;rdquo;&lt;/a&gt; and actually having understood it, I gave spelling correction in Tamil a shot.&lt;/p&gt;
&lt;p&gt;Norvig&amp;rsquo;s approach, in simple terms, is this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Get a dictionary&lt;/li&gt;
&lt;li&gt;Tweak the word you want to check (add a letter, delete one, swap 2 letters, etc.)&lt;/li&gt;
&lt;li&gt;Pick all tweaks that get you to a valid word on the dictionary&lt;/li&gt;
&lt;li&gt;Choose the most likely correction (most common correct word, adjusted for the probability of that mistake happening)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Making a dictionary is easy. I just need lots of Tamil literature, and then I pick out words from it. For now, I&amp;rsquo;m just using the texts in &lt;a href=&#34;http://www.tamil.net/projectmadurai&#34;&gt;Project Madurai&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Tweaking the word to check is easy. Norvig&amp;rsquo;s article has a working code example.&lt;/p&gt;
&lt;p&gt;Picking valid tweaks is easy. Just check against the dictionary.&lt;/p&gt;
&lt;p&gt;The tough part is choosing the likely correction. For each valid word, I need the probability of having made this particular error.&lt;/p&gt;
&lt;p&gt;Let&amp;rsquo;s take an example. I&amp;rsquo;ve spelt kathal. A list of valid tweaks to this word include: kal, kol, kadal, kanal, and kaadhal. For each of these, I need to figure out how often the valid tweaks occur, and the probability that I typed kathal when I really meant one of these tweaks. This is what such a calculation would look like:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Tweak&lt;/th&gt;
					&lt;th&gt;Frequency&lt;/th&gt;
					&lt;th&gt;Probability of typing kathal&lt;/th&gt;
					&lt;th&gt;Product&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;kal&lt;/td&gt;
					&lt;td&gt;1&lt;/td&gt;
					&lt;td&gt;0.04&lt;/td&gt;
					&lt;td&gt;0.04&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;kol&lt;/td&gt;
					&lt;td&gt;4&lt;/td&gt;
					&lt;td&gt;0.02&lt;/td&gt;
					&lt;td&gt;0.08&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;kadal&lt;/td&gt;
					&lt;td&gt;10&lt;/td&gt;
					&lt;td&gt;0.1&lt;/td&gt;
					&lt;td&gt;1.0&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;kanal&lt;/td&gt;
					&lt;td&gt;1&lt;/td&gt;
					&lt;td&gt;0.01&lt;/td&gt;
					&lt;td&gt;0.01&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;kaadhal&lt;/td&gt;
					&lt;td&gt;6&lt;/td&gt;
					&lt;td&gt;0.25&lt;/td&gt;
					&lt;td&gt;1.50&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Once we have this, we can see that kaadhal is the right one &amp;ndash; it has the maximum value (1.50) in the last column, where we multiply the frequency and the probability.&lt;/p&gt;
&lt;p&gt;(You probably realise how small my dictionary is, looking at the frequencies. Well, that&amp;rsquo;s how big &lt;a href=&#34;http://www.tamil.net/projectmadurai&#34;&gt;Project Madurai&lt;/a&gt; is. But increasing the size of a dictionary is a trivial problem.)&lt;/p&gt;
&lt;p&gt;Anyway, getting the frequency is easy. How do I get the probabilities, though? That&amp;rsquo;s what I&amp;rsquo;m working on right now.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;SDW&lt;/strong&gt; &lt;em&gt;1 May 2007 9:01 am&lt;/em&gt;:
Hi there, with regards to &amp;lsquo;Thendral Ennai Mutham Itathu&amp;rsquo; as described on &lt;a href=&#34;http://www.s-anand.net/blog/classical-ilayaraja/&#34;&gt;http://www.s-anand.net/blog/classical-ilayaraja/&lt;/a&gt;, the female voice was not SPS but B.S. Sasirekha. fyi please.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ரவிசங்கர்&lt;/strong&gt; &lt;em&gt;4 May 2007 5:33 pm&lt;/em&gt;:
நல்ல முயற்சி, அணுகுமுறை.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ravi&lt;/strong&gt; &lt;em&gt;11 May 2007 2:36 pm&lt;/em&gt;:
Good article. I am very interested in text analysis, phrase extraction too (and statistical machine translation, if you will :) Is not probability just the number of occurences of each of these divided by total number of occurence? Maybe I am missing something. Did you also consider using Vikatan or Kumudam article data in addition to Project Madurai to get contemporary word counts and probabilities? It is one of my projects to create a word frequency list for Tamil. Now that I see another active soul, I might start the project :) Python or ruby I am wondering&amp;hellip;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;kooliip&lt;/strong&gt; &lt;em&gt;1 May 2007 12:00 pm&lt;/em&gt;:
meaning for saroja sama nikalo&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GANESH&lt;/strong&gt; &lt;em&gt;1 May 2007 12:00 pm&lt;/em&gt;:
I LOVE YOU SUJA CHELLAM&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;mailto:hajaroz@yahoo.com&#34;&gt;hajaroz@yahoo.com&lt;/a&gt;&lt;/strong&gt; &lt;em&gt;1 May 2007 12:00 pm&lt;/em&gt;:
please give me that spelling checking software&amp;hellip;.it is very good and it is new one&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;v.arul&lt;/strong&gt; &lt;em&gt;14 Aug 2008 6:39 am&lt;/em&gt;:
Give me the correct tamil spelling for unnippaga&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;shamu&lt;/strong&gt; &lt;em&gt;10 Sep 2008 12:17 am&lt;/em&gt;:
தமிழ்லில் எழுதுவதற்கு முய்ற்சி&lt;br&gt;
ந்ன்றி&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;nita&lt;/strong&gt; &lt;em&gt;29 Dec 2009 3:39 pm&lt;/em&gt;:
hi i would like to know if there are any resources online which would help me to check if my spellings in tamil typing are right? i am not very confortable with tamil so i can use the help of spell check, i would be most greatful! kindly help, thanks:)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;prabhu&lt;/strong&gt; &lt;em&gt;13 Aug 2011 8:43 am&lt;/em&gt;:
Hi anand,
Do you have any tamil spell checker tool online. If yes, please provide the link asap. Thanks in advance.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;thamizh selvan&lt;/strong&gt; &lt;em&gt;19 Mar 2011 4:09 pm&lt;/em&gt;:
i want to know the exact spelling of thamizh selvan in tamil . whether a connecting letter will come between thamizh and selvan or not .. thank you&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;chandroo&lt;/strong&gt; &lt;em&gt;19 Aug 2012 6:02 am&lt;/em&gt;:
I need a Tamil spell checking software.Please tell me.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;saranraj&lt;/strong&gt; &lt;em&gt;1 Jan 2013 9:00 am&lt;/em&gt;:
Please let me know the correct spelling for the tamil word &amp;ldquo;Nadathanum&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;https://github.com/arcturusannamalai/open-tamil&#34;&gt;Muthu&lt;/a&gt;&lt;/strong&gt; &lt;em&gt;29 Apr 2015 5:37 am&lt;/em&gt;:
You may do easy text processing of Tamil content using open-tamil library in Python.
$ pip install open-tamil
To convert your data into Tamil letters (not Unicode code-points) you can type in Python,
&lt;blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;import tamil
letters = tamil.utf8.get_letters( data )
Thanks,
-Muthu&lt;/p&gt;
&lt;/blockquote&gt;
&lt;/blockquote&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>Not all distributions are normal</title>
      <link>https://www.s-anand.net/blog/not-all-distributions-are-normal/</link>
      <pubDate>Wed, 19 Jul 2006 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/not-all-distributions-are-normal/</guid>
      <description>&lt;p&gt;14 years ago, I was introduced to the process of normalising grades. Professors &amp;ldquo;fit&amp;rdquo; students&amp;rsquo; marks into a &lt;a href=&#34;http://en.wikipedia.org/wiki/Normal_distribution&#34;&gt;normal distribution&lt;/a&gt; and assign grades based on that. (I still don&amp;rsquo;t know how they do it).&lt;/p&gt;
&lt;p&gt;Since then, I&amp;rsquo;ve encountered normalising a lot. My performance at work is normalised. I normalise my song ratings and movie ratings. I&amp;rsquo;ve normalised all kinds of things at work: lead-time of delivery of fans, movements in savings account balances, calls to a call centre, demand for a resource&amp;hellip; you name it.&lt;/p&gt;
&lt;p&gt;(What I mean by normalising is, I find the mean and standard deviation, and assume that it&amp;rsquo;s a normal distribution with that mean and standard deviation. For things under my control, like movie ratings, I revise the ratings to fit a normal distribution.)&lt;/p&gt;
&lt;p&gt;In fact, I normalise everything I encounter by default.&lt;/p&gt;
&lt;p&gt;A few years ago, I started feeling uncomfortable about this. I&amp;rsquo;ve now figured out why &lt;strong&gt;normalising is bad &amp;ndash; at least when done blindly like I do&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;First, let&amp;rsquo;s explore why normalising is good. &lt;strong&gt;Normalising eliminates biases.&lt;/strong&gt; If the Prof in Section A grades higher than the Prof in Section B, normalising takes care of it. If a Prof is extremist (more A&amp;rsquo;s as well as F&amp;rsquo;s), normalising takes care of it. If a Prof is skewed (lots below average, few extremely high above average), normalising takes care of it.&lt;/p&gt;
&lt;p&gt;Eliminating biases makes sense if Section A is fundamentally like Section B. It&amp;rsquo;s not better, nor more extremist, nor more skewed. If the sections are large enough and picked randomly, this assumption is correct. If Section A represents the smarter half, or people born in the second half of the year, or people from the Western states, or any other non-random selection, this need not be correct.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;An aside: You may wonder why people born in the second half of the year is non-random. If school admissions start in September, and admissions start when you&amp;rsquo;re 3 years old, kids born in September will be nearly 4 years old when they join. Kids born in August will be between just over 3 years. That one-year difference, to a three-year old, is HUGE. For example, you will find a &lt;a href=&#34;http://www.11v11.co.uk/news/archives/152-Birth-Date-Bias-in-Football-An-AFS-Special-Survey.html&#34;&gt;birth date bias in football&lt;/a&gt;, with most premiership players being born in the months of September - November.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Normalising goes a step further than eliminating bias, however. &lt;strong&gt;Normalising forces a &lt;strong&gt;normal&lt;/strong&gt; distribution&lt;/strong&gt;. This would be right if the underlying data is normally distributed. But if not, we may be making a mistake by force-fitting.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&#34;http://www.statisticalengineering.com/central_limit_theorem.htm&#34;&gt;Central Limit Theorem&lt;/a&gt; says that if you add up random variables, you get a normal distribution. &lt;strong&gt;Provided it&amp;rsquo;s a large sample, variables are independent, and each has a finite standard deviation&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This means that &lt;strong&gt;many things you get by adding random variables are normally distributed&lt;/strong&gt;. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Number of heads when you toss a coin (add up each coin toss)&lt;/li&gt;
&lt;li&gt;Average age of an army platoon (add up each soldier&amp;rsquo;s age)&lt;/li&gt;
&lt;li&gt;Terminus-to-terminus time for a bus (add up the time between each stop)&lt;/li&gt;
&lt;li&gt;Price movement of an stock exchange index (add up each stock&amp;rsquo;s price movement)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But a lot of &lt;strong&gt;real-life data is NOT normally distributed&lt;/strong&gt;. The usual reasons are:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;It&amp;rsquo;s not the sum of random variables&lt;/li&gt;
&lt;li&gt;It doesn&amp;rsquo;t satisfy the central limit theorem (independence, large sample, finite standard deviations)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Here are some non-normal distributions that are NOT the sum of random variables:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Soldier&amp;rsquo;s age within an army platoon&lt;/strong&gt;. What random variables could you add up? You&amp;rsquo;ll probably find a lot of people at age 18, because that&amp;rsquo;s the minimum age. A little fewer at age 19 &amp;ndash; last year&amp;rsquo;s recruits. Far less at age 20 &amp;ndash; 2 years minimum service accomplished. Certainly not a normal distribution.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Price movement of a single stock&lt;/strong&gt;. What random variables could you add up? You&amp;rsquo;ll find that there are far larger price movements than a normal distribution predicts.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are some non-normal distributions that don&amp;rsquo;t satisfy the central limit theorem. (These are, in fact, things I said were normally distributed earlier. You see? It&amp;rsquo;s easy to think things are normal, but in reality they&amp;rsquo;re not.)&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The terminus-to-terminus time for a bus&lt;/strong&gt;. The number of bus stops is quite small. More importantly, the time between stops isn&amp;rsquo;t independent. If there&amp;rsquo;s a traffic jam, an entire section of the route will take more time. If there&amp;rsquo;s a delay between point 2 to 3, it&amp;rsquo;s likely that there&amp;rsquo;ll be a delay between points 1-2 and 3-4 as well.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The price movement of a stock exchange index&lt;/strong&gt;. The price movement of stocks follows a power-law distribution, which does not have finite standard deviations. Also, the price movements are not independent.&lt;/li&gt;
&lt;li&gt;See more &lt;a href=&#34;http://www.qualityamerica.com/knowledgecente/articles/PYZDEKnonnormal.html&#34;&gt;non-normal distributions&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Summary: &lt;strong&gt;Don&amp;rsquo;t assume that anything you see is a normal distribution&lt;/strong&gt;. It usually isn&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ll shortly talk about what happens when you assume something&amp;rsquo;s a normal distribution, when it really is not.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;comments&#34;&gt;Comments&lt;/h2&gt;
&lt;!-- wp-comments-start --&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&#34;http://demesos.blogspot.com/&#34;&gt;Wil&lt;/a&gt;&lt;/strong&gt; &lt;em&gt;18 Nov 2010 9:14 pm&lt;/em&gt;:
Great article - you brought it to the point!&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- wp-comments-end --&gt;
</description>
    </item>
    <item>
      <title>Knowing when to stop</title>
      <link>https://www.s-anand.net/blog/knowing-when-to-stop/</link>
      <pubDate>Wed, 21 Jun 2006 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/knowing-when-to-stop/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;http://plus.maths.org/issue3/marriage/index.html&#34;&gt;Mathematics, marriage and finding somewhere to eat&lt;/a&gt; has a simple solution to all these problems. Whether you&amp;rsquo;re hiring someone, or picking a partner, or finding a house &amp;ndash; or any problem that requires you to pick the best among N choices &amp;ndash; here&amp;rsquo;s the rule.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Scan the first 37% of choices. Then pick the first one that&amp;rsquo;s better than anything you&amp;rsquo;ve seen so far.&lt;/p&gt;
&lt;/blockquote&gt;
</description>
    </item>
    <item>
      <title>Card trick</title>
      <link>https://www.s-anand.net/blog/card-trick/</link>
      <pubDate>Wed, 14 Apr 2004 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/card-trick/</guid>
      <description>&lt;p&gt;Interesting card trick using the &lt;a href=&#34;http://www.sciencenewsforkids.org/pages/puzzlezone/muse/muse1003.asp&#34;&gt;Kruskal Count&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>More on expectations higher than expected</title>
      <link>https://www.s-anand.net/blog/more-on-expectations-higher-than-expected/</link>
      <pubDate>Thu, 03 Oct 2002 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/more-on-expectations-higher-than-expected/</guid>
      <description>&lt;p&gt;After reading my post on the ET article mentioning &amp;ldquo;&lt;a href=&#34;http://economictimes.indiatimes.com/cms.dll/articleshow?artid=23700601&#34;&gt;expected to see a higher than expected rise&lt;/a&gt;&amp;rdquo;, a certain &lt;a href=&#34;http://www.geocities.com/prachi_deuskar/&#34;&gt;CA gold-medallist friend&lt;/a&gt; of mine wrote back this obscure note that I refuse to understand:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;hellip; if you take it literally it is not possible. To put it more technically, something called a &lt;strong&gt;law of iterated expectation&lt;/strong&gt; comes to play. Today&amp;rsquo;s expectation of tomorrow&amp;rsquo;s expectation about what will happen day after is just today&amp;rsquo;s expectation of what will happen day after.&lt;/p&gt;
&lt;p&gt;I think what they mean by &amp;ldquo;&amp;hellip;&amp;rdquo; is that the revised expectation is higher than the earlier expectation.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;She is in the habit of being right, so I had better accede.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Coincidences are common</title>
      <link>https://www.s-anand.net/blog/coincidences-are-common/</link>
      <pubDate>Sun, 14 Oct 2001 12:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/coincidences-are-common/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;http://www.csicop.org/si/9809/coincidence.html&#34;&gt;Coincidences are common&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
  </channel>
</rss>
