Monday, May 11, 2015

Poll Entrée: UK 2015




Just a few days before May 7 the UK 2015 Election Day, I was lucky enough to watch the celebrated US poll aggregator Nate Silver in action in Britain on BBC Television. As you very well know, poll aggregators are websites that take the published polls results and combine them statistically to report their own election predictions.


According to Wikipedia, notable poll aggregators are Real Clear Politics; Electoral-vote.com; Princeton Election Consortium; FiveThirtyEight (founded by Nate Silver); Pollster.com or Huffpost Pollster; the political blog Talking Points Memo; Votamatic; Frontloading HQ; Election Projection; and Politics by the Numbers.


Nate Silver's prediction on BBC Television was that though the Conservative Party will beat the Labor Party the share of parliamentary seats will be so close that it will result in a hung-parliament. As someone writes: "... after near 30 minutes of mind numbingly boring footage of basically a caravan driving around the UK, finally has Nate Silver state his forecast conclusion that put the Conservatives on 283 seats, Labour 270, SNP 48, Lib Dems 24 and UKIP on just 1 seat.  Then he states the obvious that no major party even with the Lib Dem support could form a majority:"


Living in UK and having seen something like this polling scene in 2010 elections, ordinary folks may very well have a gut-feeling that the Tories will win the most seats but couldn't get a majority by themselves. So it's nothing new. In fact, this latter part of their gut-feeling may as well have been influenced by the predictions of the pollsters.


Meanwhile, one UK polling company, Survation, said they played too safe and threw away the only would be correct prediction. It claimed that its election eve poll had been close to the final result with the Conservatives on 37% and Labour on 31% (the final results were 36.9% to 30.4%). According to its CEO, "The results seemed so 'out of line' with all the polling conducted by ourselves and our peers -- what poll commentators would term an 'outlier' -- that I 'chickened out' of publishing the figures -- something I'm sure I'll always regret".


As it happened, some of the election predictions, according to one pollster were:



Source:
Market Oracle 

Market Oracle

May2015.com

Electoralcalculus.co.uk

ElectionForecast.co.uk

The Guardian

28th Feb
26th Apr
26th Apr
27th Apr
27th Apr
Conservative
296
272
279
286
274
Labour
262
271
282
267
270
SNP
35
55
47
48
54
Lib Dem
30
26
18
24
27
UKIP
5
3
1
1
3
Others
22
22
22
22
22


Predictions from other sources, for example, as given in Wikipedia, were not much different from these. In terms of the vote share percentages, mostly it was like one or two percentage points over the Labor in favor of the Conservative party.
The exit poll result given at 10 PM at the end of voting on the Election Day gives a more realistic number of seats for the Conservatives, but still not enough to give a majority.


Why couldn't all the best pollsters predict the majority for the Conservatives? And not even by the exit poll?


One of the biggest problems for pollsters in the recent times is that the response rate for pre-election polls was getting really low. So that getting a traditional poll done on a good representative sample of voters is getting well-nigh impossible or getting prohibitively expensive. Some hails internet polling as used by YouGov and others as possible remedy for such flaws of the traditional polls.

But, YouGov’s poll on Thursday night, conducted after votes had been cast, was baffling. It again showed that the parties were neck-and-neck (How 'shy Tories' confounded the polls and gave David Cameron victory, Jessica Eglot, The Guardian, 8 May 2015). Besides, it didn't do well recently in the 2014 Scottish independence referendum. One poll by YouGov created a stir when it published a 2% lead for "Yes" and the actual outcome was an 11% lead for "No" (Election 2015: Hold on, this isn't what you said would happen, BBC newsbeat).

Generally, it would be more difficult to get the polls right in the political system prevailing in UK than in the US. On the FiveThirtyEight website Nate Silver said "When there are only two major candidates, the choice isn’t very complicated. ... UK has become less and less of a two-party system. While the Conservatives and Labour collectively accounted for about 90 percent of the vote through the election of 1970, they’ll be down to somewhere in the neighborhood of 65 percent to 70 percent this year."  


He was worried that "The World May Have A Polling Problem":


Consider what are probably the four highest-profile elections of the past year, at least from the standpoint of the U.S. and U.K. media:


Perhaps it’s just been a run of bad luck. But there are lots of reasons to worry about the state of the polling industry. Voters are becoming harder to contact, especially on landline telephones. Online polls have become commonplace, but some eschew probability sampling, historically the bedrock of polling methodology. And in the U.S., some pollsters have been caught withholding results when they differ from other surveys, “herding” toward a false consensus about a race instead of behaving independently. There may be more difficult times ahead for the polling industry.


The British Polling Council (an association of UK polling organizations) declared on 7 May that it will set up an independent inquiry to examine "the possible causes of this apparent bias" in the UK 2015 election polls and make recommendations for future polling.


As reported by CNN, Council president John Curtice -- Professor of Politics at Strathclyde University – said that while polls should be judged on their percentages rather than seats and while it could well be true that many had fallen within their margins of error, an inquiry was still needed. The polls had been accurate on the SNP, Liberal Democrats and Greens -- but they all had an error in the same direction, he said. And that "The reason an inquiry has been set up is that actually the industry collectively clearly underestimated the Conservative lead over Labour".


What would happen if the UK2015 election system were proportional representation than first-past-the-post as it was done? This is what The Electoral Reform Society, a campaign group, arrived at through their calculation using the D'Hondt method of converting votes to seats.



Well, this is interesting, but the real challenge for pollsters in the present electoral system in UK is another kind of vote conversion. It is the need to convert estimates of votes received by parties to estimates of seats in the parliament. So even if you could get the number of votes within the margin of error you planned for, you may still get the error in projected parliamentary seats that hurts.


Meanwhile, International New York Times labeled pollsters as British Election's Other Losers (Dan Bilefsky, 8 May, 2015). It cited Alberto Nardelli, the Guardian's data editor, who said there was no simple explanation to what went wrong with the polling.
“It could be simply that people lied to the pollsters, that they were shy or that they genuinely had a change of heart on polling day,” he said. “Or there could be more complicated underlying challenges within the polling industry, due, for example, to the fact that a diminishing number of people use landlines or that Internet polls are ultimately based on a self-selected sample.”
It also cited Korteweg who thinks that polling disasters are becoming a trend in UK:
Rem Korteweg, a senior research fellow at the Center for European Reform in London, said that British pollsters were going through a particularly bad period. He cited the referendum last year on Scottish independence, when pollsters’ predictions of a neck-and-neck performance for the no and yes camps were upended by results in which 55 percent of Scots voted against becoming independent compared with 45 percent in favor.
“This isn’t the first time in 18 months when the polls got it wrong,” he said. “This is starting to become a trend in this country.”
Mr. Korteweg attributed the pollster’s failings in this latest election to the fact that voters often give socially desired responses during polling, only to behave differently when they vote. “People say who they are voting for with their heart and then vote with their wallets,” he said.
And who even suggested the possibility of having exploited dark fears in politics:


He said that the last-minute surge by the conservatives could also be explained by the fact that the Conservatives had adroitly exploited fears among voters that the Labour Party would be able to govern only in coalition with the Scottish National Party, which wants an independent Scotland.


What really happened in the UK 2015 General Elections was that the Conservatives got a majority since 1992, and the Labor Party suffered their worst defeat since 1987 (United Kingdom general election, 2015, Wikipedia).


Party
Leader
Votes
Votes %
Seats
Seats %
Conservative Party
David Cameron
11,334,920
36.9%
331
50.9%
Labour Party
Ed Miliband
9,344,328
30.4%
232
35.7%
UK Independence Party
Nigel Farage
3,881,129
12.6%
1
0.2%
Liberal Democrats
Nick Clegg
2,415,888
7.9%
8
1.2%
Scottish National Party
Nicola Sturgeon
1,454,436
4.7%
56
8.6%
Green Party
Natalie Bennett
1,154,562
3.8%
1
0.2%
Democratic Unionist Party
Peter Robinson
184,260
0.6%
8
1.2%
Plaid Cymru
Leanne Wood
181,694
0.6%
3
0.5%
Sinn Féin
Gerry Adams
176,232
0.6%
4
0.6%
Ulster Unionist Party
Mike Nesbitt
114,935
0.4%
2
0.3%
Social Democratic & Labour Party
Alasdair McDonnell
99,809
0.3%
3
0.5%
Others
N/A
349,487
1.1%
1
0.2%



Monday, May 4, 2015

Exit Polls etc


Exit polls work like this (EXPLAINER: How Exit Polls Work, Robin Sproul, Nov. 1, 2014, ABC News):

How conducted:

Interviewers stand outside polling places in precincts that are randomly selected. They attempt to interview voters leaving the polling place at specific intervals (every third or fifth voter, for example).

Voters who agree to participate in the poll fill out a short questionnaire and place it in a ballot box. Interviewers phone in results three times during the day.
When a voter refuses to participate, the interviewer notes the gender and approximate age and race of that voter. In this way, the exit poll can be statistically corrected to make sure all voters are fairly represented in the final results.

Questions asked:

... typical(ly) ... Who they just voted for in key races; What opinions they hold about the candidates and important issues; The demographic characteristics of the voter

 

Example of a 2014 exit poll issue question:


How worried are you about the direction of the nation’s economy in the next year?
• Very worried
• Somewhat worried
• Not too worried
• Not at all worried

 

Are exit polls accurate?


... like any other survey, are subject to sampling errors. Before news organizations report any exit poll results or make projections ... they compare results to pre-election polls, past precinct voting history, and have statisticians and political experts carefully review the data.
After the polls close the exit poll results are weighted using the actual vote count to make the data more accurate. Even projections that are made without any actual vote data are not based solely on the results of exit polls.

Accounting for the early votes or by mail:

 

In the 2012 election, just over one third of Americans voted before Election Day, using some form of absentee or early voting. Capturing information about these voters is challenging, but it is critical to report accurate information about all voters.

In states with high numbers of absentee/early voters, telephone polls are conducted to reach those voters. Data from these telephone polls are combined with the exit poll data to provide a complete portrait of all voters.

Reporting exit poll results:

On Election Day, there is a strict quarantine on any news coming from the early waves of exit poll data until 5:00 p.m. ET. By about 5:45 p.m., some initial demographic information about voter turnout will be available on ABCNews.com.

Winners will not be projected until polls are closed, so announcements come on a state-by-state basis as individual state polls close. Information will be constantly updated throughout the evening on ABCNews.com and on all ABC News programs.
The following are the hard questions needed to ask, not only by the media, but also by the public to critically understand a poll result (20 Questions A Journalist Should Ask About Poll Results, http://www.ncpp.org/?q=node/4):

1. Who did the poll?
2. Who paid for the poll and why was it done?
3. How many people were interviewed for the survey?
4. How were those people chosen?
5. What area (nation, state, or region) or what group (teachers, lawyers, Democratic voters, etc.) were these people chosen from?
6. Are the results based on the answers of all the people interviewed?
7. Who should have been interviewed and was not? Or do response rates matter?
8. When was the poll done?
9. How were the interviews conducted?
10. What about polls on the Internet or World Wide Web?
11. What is the sampling error for the poll results?
12. Who’s on first?
13. What other kinds of factors can skew poll results?
14. What questions were asked?
15. In what order were the questions asked?
16. What about "push polls"?
17. What other polls have been done on this topic? Do they say the same thing? If they are different, why are they different?
18. What about exit polls?
19. What else needs to be included in the report of the poll?
20. So I've asked all the questions. The answers sound good. Should we report the results?

Pollsters like to use the word "scientific" liberally. You find the terms scientific polling or scientific sample of eligible voters or voters in their writings constantly. That may be because pollsters mostly deal with the public and so they favor the term scientific as that would be more appealing over the probability sample which is the standard description used in mainstream sample survey literature.

"A fundamental tenet of scientific measurement is that the measuring device is standardized over different objects being measured" [Groves 1989, as cited by Bishop and Mockabee, Survey Practice, Vol 4, No 6(2011)]. It follows that any question should mean essentially the same thing to all respondents. A question should be worded so that every one should be answering the same question. Without such standardization you can't know if your measurements are reliable and valid.

Although most public opinion researchers have certainly become sensitive to the effects of variations in question wording and context, they have, with few exceptions, been much less attentive – if not oblivious – to how the meaning-and-interpretation of survey questions can vary across respondents and over time even when the wording and context of the question itself remains identical. How well, for question? What, if anything, do we know about how respondents interpret the question: “Do you approve or disapprove of the way Barack Obama is handling his job as president?” Or how he’s handling the economy? What does “handling his job as president” actually mean to them? What does “the economy” mean? Does it vary across respondents and over time?

On the other hand, the pollsters' obsession with the qualifier "scientific" may be traceable to sociological roots. According to sociologist Howard S. Becker (Criticism of polls and surveys in American social science, March 24, 2015, Observatoire des Sondages), in addition to the for-profit commercial branch of polling, there is the other and more ambitious branch, grown out of academic survey research, that tried to make a scientific social science:

We can better understand today’s polls if we see them as part of a larger movement, designed to create a “scientific” social science, whose two connected but distinguishable branches collaborated in the shared effort to legitimate a style of research that came to be known, variously, as survey research or polling. One branch grew out of the interest of businesses, and the advertising agencies they supported, in finding out what their audiences and customers wanted so that they could make larger profits. The other grew out of the statistical tradition in sociology which, ... wanted to prove that sociology and related disciplines studying contemporary society were “real sciences” like physics and chemistry, capable of producing demonstrably true generalizations and laws by using the rigorous methods of measurement and statistical and mathematical analysis of those sciences. ...



For the question of whether this ambition of sociologists in America succeeded, Becker says:

Many thought, and still think, that they succeeded, and the proof can be seen in the pages of the major American sociological journals, where studies in this style provide the vast majority of articles.
But the victory was never complete. ... And then the whole enterprise lost substantial ground as a result of the 1948 presidential election in the United States. ...

This event caused a serious reconsideration of the many problems of doing accurate polling and making predictions that withstood the test of reality. The credibility of the whole operation was being openly doubted. This failure of polling and survey methods affected both the commercial interests of the big polling organizations and the aspirations and continued existence of academic research organizations and individual scholars.

The predictive failure of the election polls had important consequences. In the time between 1936 and 1948, polling had become a large business, which made its profits by doing surveys designed to help commercial enterprises—manufacturers, advertisers, radio networks, Hollywood studios—guess what the buying public would respond to in a way that would make money for them. Election studies had become what they have remained [2], the one kind of survey study whose accuracy can be assessed by comparison with the events it is meant to predict.

The resilient polling industry and the social scientists have taken such failures to heart and have been looking for richer techniques to improve their analyses, predictions, and estimates ever since. Methodologically polling started out with "straw polls" before 1936.

□       Literary Digest had successes with straw polls from 1916 to 1932.
□       Literary Digest (with 2.3 million in straw sample) ousted by Gallup in 1936 with his 3000 respondents in quota sample.
□       Later, quota sampling gave way to random sampling or more precisely probability sampling. Today every lookup for the definition of "scientific polling" gives something like "Scientific polling consists of surveying a random sample of the population in order to obtain statistically significant results for an upcoming vote or election".

There could be many shades of scientific polling as far as random selection goes and some may not be so scientific. In a recent publication by Open Society Foundation (From Novelty to Normalcy: Polling in Myanmar’s Democratic Transition, March 15, 2015, available at http://www.opensocietyfoundations.org/sites/default/files/polling-myanmar-democratic-transition-20150318.pdf ) talking about survey sampling in Myanmar we've read that:

"At the local level in Myanmar, households are selected by a random process, much as they are in countries with a longer research tradition. Most Myanmar research companies we spoke with described techniques that are standard in international research: direct multi-stage sampling of geographical areas, then sampling of spots (villages or other small geographic units), then sampling of dwellings by a random walk procedure, then identification of qualified individuals in the dwelling by means of screening questions, and then selection of an individual ... ". (pp. 13-14)

If so the random process would be marred by the random walk procedure which is not that standard. About that, Designing Household Survey Samples: Practical Guidelines (ST/ESA/STAT/SER.F/98, United Nations, 2005) said:

22. Another type of non-probability sampling that is widely used is the so-called “random walk” procedure at the last stage of a household survey. The technique is often used even if the prior stages of the sample were selected with legitimate probability methods. The illustration below shows a type of sampling that is a combination of random walk and quota sampling. The latter is another non-probability technique in which interviewers are given quotas of certain types of persons to interview.

Example
To illustrate the method, interviewers are instructed to begin the interview process at some random geographic point in, say, a village, and follow a specified path of travel to select the households to interview. It may entail either selecting every nth household or screening each one along the path of travel to ascertain the presence of a special target population such as children under 5 years old. In the latter instance each qualifying household would be interviewed for the survey until a pre-determined quota has been reached.

And it went on to describe the technique which appears essentially to be to avoid the listing of households. However, the verdict was that in practice a random sample may not be realized as its supporters claimed because:

24. ... It usually fails due to (a) interviewer behaviour and (b) the treatment of nonresponse households including those that are potentially non-response. It has been shown in countless studies that when interviewers are given control of sample selection in the field, biased samples result.

And it is more likely to be biased:

25.  ... With the quota sample approach, persons who are difficult to contact or unwilling to participate are more likely to be underrepresented than would be the case in a probability sample. In the latter case interviewers are generally required to make several callbacks to households where its members are temporarily unavailable. ...
 
In addition to the evolution of scientific polls out of ad hoc straw polls, Hillygus observed these new trends in polling (The Evolution of Election Polling in the United States, Public Opinion Quarterly, Vol. 75, No. 5, 2011, http://poq.oxfordjournals.org/ content/75/5/962.full.pdf#page=1&view=FitH):

[1] "In forecasting the election, statistical models and prediction markets appear to be viable alternatives to polling predictions, especially early in the campaign."
[2] "In understanding voting behavior, surveys are increasingly replaced by experimental designs or alternative measures of attitudes and behaviors."
[3] "In campaign strategy polls are increasingly second fiddle to massive databases from voter files and consumer databases, changing the campaign messages that we see."  

She also observed that:

With the proliferation in polls, we have also seen greater variability in the methodologies used and the quality of the data. The lack of transparency about those methodologies has contributed to skepticism about the industry. Coupled with changes in technology and the information environment, it is perhaps no wonder that polls have lost some of their luster.

All in all, for those of us who were not yet captivated by the magical allure of polling, it would be wise to give hard looks at messages of both the polling pundits and their critics. Besides, we should as well keep our eyes wide open for the up and coming technologies and new trends.

A prominent Myanmar historian once said that the purpose of learning history is to prevent ourselves from becoming dumb asses. Learning technology and science is perhaps to make our hearts to be able to defy stresses or ignore temptations a bit longer, so that our heads may have a bit more time to think.



Sunday, April 26, 2015

Elections and Polls


How does an election poll work as a business, for example, in the US?
"For pollsters, there's no money in asking questions about elections and releasing the numbers to the media. They do it as a marketing tool to attract clients who want to know what people think about, say, shampoo."

The fact is that commercial polls are where money is and election polls, if paid for by news sources like newspapers, radio and TV, are not profitable, and are often operated at a loss.

Opinion polls and market research has been done for a long time. But there is no real way to verify even some simple information like what percentage of the people are using this brand of shampoo or that brand of toothpaste. Forget about asking the opinions of the people if they prefer this policy or that policy or if they like this or that political party, or the government, especially if you are at a place where people still needs to get used to speaking out.

Predicting elections results correctly is about the only way for pollsters to show they know their stuff. But it is not always easy to get the predictions right. A classic example is the 1936 Roosevelt-Landon presidential election in the US. The Literary Digest conducted a postal opinion poll aiming to reach 10 million people, a quarter of the electorate. After tabulating 2.4 million returns they predicted that Landon would win by a convincing 55 per cent to 41 per cent. But the actual result was that Roosevelt crushed Landon by 61 per cent to 37 per cent. In contrast, a small survey of 3000 interviews conducted by the opinion poll pioneer George Gallup came much closer to the final vote, forecasting a comfortable victory for Roosevelt.

With this success Gallup went on to establish the “American Institute of Public Opinion (AIPO)” with the goal “impartially to measure and report public opinion on political and social issues of the day without regard to the rightness and wisdom of the views expressed.” It was the real beginning of the claim by pollsters, past and present, that polls "can measure the true will of the people and that, through the polls, the people get a real voice in between elections and on all kinds of issues." Now there are reasons to suspect if that is not too much of a publicized or idealized view.
                                                                                                                          
For the history and development of opinion polls we will have to look at the history of this industry in the US. It has its beginnings there and it still has its biggest presence there. Forties in the last century was the time institutionalization and professionalization of public opinion research made headway. In 1941 the first university institute National Opinion Research Center at the University of Chicago was founded followed by American Association for Public Opinion Research, the first professional/academic association in 1946. The World Association for Public Opinion Research followed one year later. In 1948 the Public Opinion Quarterly the first academic journal was published. Development of opinion polls in Europe was somewhat delayed because of the effects of the World War.

Yet 12 years after Gallup's acclaim, the most famous failure was the polls predicting that Republican Thomas Dewey would beat incumbent Democratic president Harry Truman in the 1948 election. Not only Gallup, but two other major pollsters Crossley and Roper were wrong too.


Candidate
Party
Electoral Votes
Percent Popular
Votes
Final
Gallup
Estimate
Final
Roper
Estimate
Final
Crossley
Estimate
Harry Truman
Democrat
303
49.6%
44.5%
38%
45%
Thomas Dewey
Republican
189
45.1%
49.5%
53%
50%

"Between 1956 and 2004 the US presidential elections showed an average deviance of only 1.9 percent (based on Mosteller method 3, one of several statistical ways how to calculate the margin of error ... But there have been major disasters for the pollsters in many countries, e.g. the US presidential elections of 1980, the British parliamentary election of 1992, or the German parliamentary election of 2005", and US presidential elections again in 2012. Even with these exceptions the polls were correct in predicting election outcomes most of the time. This is really not surprising, if we note that in mature democracies people have little reasons to lie as to whom they'll vote for and given that the sample of eligible voters is truly representative, and if the voter turnout is big enough, the numbers will always prove to be correct. So, the explanations for prediction failures must find fault with things other than the problem of the deceiving citizenry in those countries.

When an incumbent and an inspiring candidate meet in the presidential elections the Americans have a simple explanation for the results: the people let the good one stay and kick out the bum. Maybe some pollsters don't believe in such simple formulas or else they find it much harder to distinguish a good one and a bum than people do.

The International New York Times (Why Polls Can Sometimes Get Things So Wrong,
July 3, 2014) explained:

The science of polling is sound, but if you ask the wrong group of people your poll questions, you can get the wrong answers. Think of it this way: An arrow shot by an expert marksman has some chance of hitting the target depending on the wind, the distance and any number of other things, but if the marksman aims at the wrong target, those other things have nothing to do with why the arrow misses.

In 2012 Obama-Romney presidential elections, Gallup was wrong again and gave four factors that reduced the accuracy of its polling in a 17-page report. According to Huffington Post, June 5, 2013 they were:

Misidentification of Likely Voters. ... using a procedure developed in the 1950s ... Last year, this likely voter model moved Gallup's estimate of the margin separating Obama and Romney 4 points in Romney's direction.
Under-Representation of Regions. Gallup also weights its data by a variety of factors ... effectively undersampling states that vote more Democratic.

Faulty Representation of Race and Ethnicity. ...Gallup in recent years has used an unusual method to ask about race that distorted the racial composition of its samples when the data were eighted. ...This led to a disproportionate number of people who said they were multiracial, and that in turn distorted the weighting procedure, effectively giving too much weight to some white voters.

Nonstandard Sampling Method. Before 2011, Gallup had selected phone numbers using random digit dialing, or RDD, which calls randomly generated numbers. This is the procedure that most national media polls have used for decades. ... Gallup made a significant change in 2011, when it dropped the RDD methodology for its landline sample, using instead numbers randomly selected from those listed in residential telephone directories. ... But the change came with a downside: Not everyone who has a landline has a listed number. Although Gallup's initial research indicated that cell phone calls would cover the difference, they didn't: The listed sample turned out to be older and more heavily Republican than the RDD sample.
It is evident that all things being equal, exit-poll, that is, polls taken at a sample of the voting stations of a sample of voters after they have voted, would be much more accurate. But exit polling is not allowed everywhere and for example Singapore and New Zealand have a complete ban. Less accuracy aside, even when exit polls are allowed, pre-election polls are also in much demand. Because "they are the basis for campaign strategy by candidates, parties, and interest groups. They are the primary tool that academics and journalists use to understand voting behavior."

It is correct to say that in US, the first day after elections is the beginning of the season for next rounds of election polling. The only alertness, tenacity, and single mindedness comparable with that in the case of our citizens seems to be their search for cramming masters for their children as soon as a high-school completion exam (which also doubles as the university entrance exam) is over. Maybe we could expect a different kind of polling and marketing industry based on this to develop as we now see a flurry of activities by polling pundits trying to make inroads into Myanmar.


                                                                  



Thursday, April 2, 2015

Bhagavad Gita and Paradigms


What do they have in common? Bhagavad Gita and Paradigms?
Gita was thought to be written by Sage Ved Vyasa some time in the fifth to second century BCE according to Wikipedia, while Thomas Kuhn published his "The Structure of Scientific Revolutions" in 1962. I used to have both of them in my small collection. But no more.


Kuhn's book I had was the first edition, a paper back which I remember as having its cover mostly yellow. Now thanks to Wikipedia I could show what it looked like and I am proud to be one of the owners of first edition of this highly acclaimed work. Can't remember from which bookshop I picked it up. It was in Yangon.

It was in the mid sixties, I think. When I first saw its title I thought it must be about how to launch successful revolutions and at that time, I must note, that a lot of people were interested in revolutions. At that time also, I remember being quite skeptical about revolutions that are scientific because I'd read about scientific method by one Mr. Singer. Singer said when you say "a scientific boxer" you are misusing what is meant by the true scientific method. Those were the thoughts occurring before I picked up the book and read the description about the content on the book cover. Then I saw that it was on revolutions in science.

Bhagavad Gita when I discovered it, was a small, thin, slim book in the Bogyoke Aung San Museum in Yangon. It was in a small bookcase with glass doors and I could clearly read the title on the jacket. I am not sure, if the bookcase was in the bedroom beside a simple bed or in some other place. One of these days I would visit the museum to see if the book is still there.


When was this visit? I have forgotten completely. But if I were able to find my own book which was exactly like Bogyoke's I would have known when it was because I would sign all my books on their title pages with dates. The fact was that not too long after my visit to the museum I was lucky to find Bhagavad Gita, the same one as translated by Sir Edwin Arnold with a yellowish jacket and a red hardcover. I think, I picked up this book from the City Book Club and if I remember it right it was in the Pansodan Street, Yangon.

This book must have been misplaced. I've been looking for it for the last two days and still looking. Of the Kuhn's book, I had given it to my cousin years ago. Now that he had passed away and his family had moved to US and UK I would never have the chance to find out when I bought it. What I can say now is that from it I was able to get some bits of understanding on paradigm shifts in the sciences. Before I discovered Kuhn's book, I'd been reading quite a bit on history and philosophy of science. It was in the early sixties and during that time I found myself a year or so away from my studies at Rangoon University for some reasons of my own. In my recollection as of now of Kuhn's central idea, before re-reading him, is that while the scientists worked within an accepted idea (paradigm) of their times, some would find results that are at odds with the current paradigms. First, these will be recorded as footnotes in their papers. Later the inconsistencies grew too large to be accommodated within the current paradigms and new paradigms (revolutions) arise because of necessity.

As for Bhagavad Gita, I know it was a 700 verses long part of the great Indian epic, Mahabharata. I tried reading it, but never got past the first few verses! Nevertheless it was a great joy that I owned an exact copy of what Bogyoke had owned. What were his thoughts reading Bhagavad Gita? Where or from whom he'd heard about it? What was this thirtyish young Statesman contemplating in relation to the new nation he had been working on? We may never know.


Anyway, this internet age has provided me with the option of retrieving the texts of The Structure of Scientific Revolutions and Bhagavad Gita for free. It has compensated me for the texts I'd lost, but would never replace the physical objects that held them and the memories and the pride that goes with them.

The preface by Edwin Arnold in his translation of Bhagavad Gita is mostly excluded in many of the download sources on the Web, but included in the version produced as Project Gutenberg etext at http://livros.universia.com.br/?dl_name=The-Song-celestial-or-Bhagabad-gita-from-the-Mahaharata-being-a-discourse-de-Autor-Desconhecido.pdf.

This famous and marvellous Sanskrit poem occurs as an episode of the Mahabharata, in the sixth—or "Bhishma"—Parva of the great Hindoo epic. It enjoys immense popularity and authority in India, where it is reckoned as one of the ``Five Jewels,"—pancharatnani—of Devanagiri literature.  ... the question of its date, which cannot be positively settled. It must have been inlaid into the ancient epic at a period later than that of the original Mahabharata, ... The weight of evidence, however, tends to place its composition at about the third century after Christ;


Encyclopedia Britannica online gives a concise description of Bhagavad Gita:

Bhagavadgita, ( Sanskrit: “Song of God”) an episode recorded in the great Sanskrit poem of the Hindus, the Mahabharata. It occupies chapters 23 to 40 of Book VI of the Mahabharata and is composed in the form of a dialogue between Prince Arjuna and Krishna, an avatar (incarnation) of the god Vishnu. Composed perhaps in the 1st or 2nd century CE, it is commonly known as the Gita.

On the brink of a great battle between warring branches of the same family, Arjuna is suddenly overwhelmed with misgivings about the justice of killing so many people, some of whom are his friends and relatives, and expresses his qualms to Krishna, his charioteer—a combination bodyguard and court historian. Krishna’s reply expresses the central themes of the Gita. He persuades Arjuna to do his duty as a man born into the class of warriors, which is to fight, and the battle takes place. ...


Now, after reading the Bhagavad Gita page of Wikipedia I wonder if Bogyoke had not been reading Bhagavad Gita and thinking about independence movement the way Indian nationalists had been inspired by Gita.

At a time when Indian nationalists were seeking an indigenous basis for social and political action, Bhagavad Gita provided them with a rationale for their activism and fight against injustice. Among nationalists, notable commentaries were written by Bal Gangadhar Tilak and Mahatma Gandhi, who used the text to help inspire the Indian independence movement. ... No book was more central to Gandhi's life and thought than the Bhagavad Gita, which he referred to as his "spiritual dictionary". ... Mahatma Gandhi expressed his love for the Gita in these words:

I find a solace in the Bhagavadgītā that I miss even in the Sermon on the Mount. When disappointment stares me in the face and all alone I see not one ray of light, I go back to the Bhagavadgītā. I find a verse here and a verse there and I immediately begin to smile in the midst of overwhelming tragedies – and my life has been full of external tragedies – and if they have left no visible, no indelible scar on me, I owe it all to the teaching of Bhagavadgītā.


Back to Kuhn, February 2012 marked the 50th anniversary of the publication of "The Structure of Scientific Revolutions" and John Naughton wrote in The Guardian (Thomas Kuhn: the man who changed the way the world looked at science, 19 August 2012, http://www.theguardian.com/science/2012/aug/19/thomas-kuhn-structure-scientific-revolutions):

Fifty years ago this month, one of the most influential books of the 20th century was published by the University of Chicago Press. Many if not most lay people have probably never heard of its author, Thomas Kuhn, or of his book, The Structure of Scientific Revolutions, but their thinking has almost certainly been influenced by his ideas. The litmus test is whether you've ever heard or used the term "paradigm shift", which is probably the most used – and abused – term in contemporary discussions of organisational change and intellectual progress. A Google search for it returns more than 10 million hits, for example. And it currently turns up inside no fewer than 18,300 of the books marketed by Amazon. It is also one of the most cited academic books of all time. So if ever a big idea went viral, this is it.
... Before Kuhn, in other words, we had what amounted to the interpretation of scientific history, in which past researchers, theorists and experimenters had engaged in a long march, if not towards "truth", then at least towards greater and greater understanding of the natural world.

"Whig history is the approach which presents the past as an inevitable progression towards ever greater liberty and enlightenment, culminating in modern forms of liberal democracy and constitutional monarchy. ... The term is also used extensively in the history of science ... which focuses on the successful chain of theories and experiments that led to present-day science, while ignoring failed theories and dead ends" (Whig history, Wikipedia). In this context, Kuhn was able to accommodate the failed theories and dead ends as well as uncertainties and phenomenal successes of science in a coherent way in his more realistic version of the history of science.

Kuhn's version of how science develops differed dramatically from the Whig version. Where the standard account saw steady, cumulative "progress", he saw discontinuities – a set of alternating "normal" and "revolutionary" phases in which communities of specialists in particular fields are plunged into periods of turmoil, uncertainty and angst. These revolutionary phases – for example the transition from Newtonian mechanics to quantum physics – correspond to great conceptual breakthroughs and lay the basis for a succeeding phase of business as usual. The fact that his version seems unremarkable now is, in a way, the greatest measure of his success. But in 1962 almost everything about it was controversial because of the challenge it posed to powerful, entrenched philosophical assumptions about how science did – and should – work.

... What Kuhn had run up against was the central weakness of the Whig interpretation of history. By the standards of present-day physics, Aristotle looks like an idiot. And yet we know he wasn't. Kuhn's blinding insight came from the sudden realisation that if one is to understand Aristotelian science, one must know about the intellectual tradition within which Aristotle worked. One must understand, for example, that for him the term "motion" meant change in general – not just the change in position of a physical body, which is how we think of it. Or, to put it in more general terms, to understand scientific development one must understand the intellectual frameworks within which scientists work. That insight is the engine that drives Kuhn's great book.

Nevertheless, it is nothing strange that Kuhn's great book has its share of serious criticisms. Some would be strong, some quite mild. Of the first variety, I felt, what were mentioned under the section "Criticisms" in the Wikipedia page "The Structure of Scientific Revolutions" would qualify. I would say Greg Radick's criticisms recounted in "The Guardian" (Beyond our Kuhnian inheritance, Rebekah Higgit, 28 August 2012) at http://www.theguardian.com/science/the-h-word/2012/aug/28/thomas-kuhn were a milder example. I'm sure there is much more than what is touched here by me, second hand and a bit outdated.

Now, what do they have in common? Bhagavad Gita and Paradigms?
Personally, I could just say I used to have a book of Sir Edwin Arnold's translation of Gita and the first edition of Thomas Kuhn's book on scientific revolutions or the progress of the paradigms. What they have in common was a common owner, me; nothing simpler than that.

Nostalgically I've downloaded additional five versions in English of the Gita, and Kuhn's book.

If you are interested, for Gita try:

Get Kuhn's "The Structure of Scientific Revolutions" at:



Wednesday, March 18, 2015

Forest to Trees, the Ecological Fallacy


There's is a saying "can't see forest (or wood) for the trees" which meant to warn us that focusing too much on details could make you miss the whole picture, overall impression, or key point. What about turning that upside down and talk about individual trees by just looking at the forest?


Years ago, I had known a community development initiative by one UN agency in Myanmar, which engaged a local consulting firm to select a number of poorest communities in some selected areas in Myanmar. As far as I know the contractor devised some scoring scheme for Village-Tracts which are the lowest areal units in the administrative system in Myanmar. A Village-Tract normally contained a number of villages (hamlets) and a village is usually treated as a community for the purpose of community initiatives in Myanmar. Quite by accident I came to know at least one concrete instance of the problem of using Village-Tracts (collection of communities) to identify the poorest communities for targeting community development assistance.  

That time, I was working on a terminal assessment for a community forestry initiative assisted by an INGO. I visited an independent project located in central Myanmar that was doing research in organic farming and providing training on that subject. The researcher in residence told me that he had been approached by the UN agency, which I have mentioned, to administer the distribution of aid to the poorest communities which the UN agency had identified. The researcher was an expatriate who have been living and working on site for ten years. He knew the area well and he wasn't impressed with the "poorest communities" that had been identified. So he insisted that he and his team would discard this list and would freshly determine the poorest communities on their own if they were to distribute the aid. And happily, he was allowed to go on working with his improved targeting exercise and the distribution of aid.

Lesson: it is a clear case of ecological fallacy.

The ecological fallacy is a logical error of interpretation that involves deriving conclusions about the nature of individuals solely on an analysis of group data. (It's Fallacy Friday: Ecological Fallacy, February 27, 2015, PHI KAPPA Literary Society)

An ecological fallacy (or ecological inference fallacy) is a logical fallacy in the interpretation of statistical data where inferences about the nature of individuals are deduced from inference for the group to which those individuals belong. (Wikipedia)

The ecological fallacy refers to the incorrect assumption that the relationships between variables observed at the aggregated, or ecological-level, are the same at the individual-level. (Ecological Fallacy: Concepts, Causes and Solutions, Wei at. al, University of Manitoba, January 10, 2010)

Ecological fallacy is a fallacy in research, wherein you draw an inference about a group, and incorrectly attribute that inference to any individual in that group. (Explanation of Ecological Fallacy in Research with Examples, Neha B Deshpande in Buzzle, March 9, 2015). Illustration below:


Pollet et. al. cites a case due to which ecological fallacy became well known (http://www.willem.maartenfrankenhuis.nl/wp-content/uploads/2013/06/Pollet-et-al.-2014-human-nature1.pdf ):

The term “ecological fallacy” became well known after William Robinson (1950) used U.S. census data to test hypotheses related to immigration and literacy ... hypothesis: those who are literate are more likely to migrate, and therefore proportions of immigrants within a state will be positively related to literacy rates in those states. (He) found evidence for such a positive relationship between the average literacy of U.S. states and the proportion of immigrants living in those states. However, at the individual level, immigrants were less likely to be literate than native individuals ... The positive state-level relationship between proportion of immigrants and literacy rates might have arisen because immigrants tended to settle in states with higher literacy levels, perhaps because these states afforded better economic opportunities or were otherwise more tolerant of immigrants. Thus, literacy levels are higher in some states despite, rather than because of, lower literacy among immigrants. The state-level literacy statistics at the aggregate level did not accurately reflect the literacy of immigrants (and, indeed, portrayed a pattern that was opposite to the individual-level pattern). In sum, then, the ecological fallacy is committed when group-level relationships are assumed to reflect individual-level relationships. The fallacy can occur when group aggregates are incorrectly assumed to be representative of individuals within those groups, or when macrolevel relationships are governed by processes that are unrelated to those hypothesized to operate at the individual level. In the Robinson (1950) study on literacy, for example, the scores at state level were assumed to represent literacy of immigrants and nonimmigrants equally, whereas at the individual level, immigrants were less likely to be literate than non-immigrants.

In our example of the targeting exercise for the poorest villages think about a particular village-tract that is excluded because it is not poor enough in terms of the village-tract level poverty score.  Consider the case that this village-tract contains one village for which if the poverty score were taken at the village level it would be poorer than any of the village in the village-tracts included in the target. Here we could easily see how a village could be wrongly excluded while another could be wrongly included in its place because of the ecological fallacy arising from taking village-tracts instead of villages.  

The Ecological Fallacy entry in Wikipedia contains an interesting example on election for governor of Washington, USA:

The ecological fallacy was discussed in a court challenge to the Washington gubernatorial election, 2004 in which a number of illegal voters were identified, after the election; their votes were unknown, because the vote was by secret ballot. The challengers argued that illegal votes cast in the election would have followed the voting patterns of the precincts in which they had been cast, and thus adjustments should be made accordingly. An expert witness said this approach was like trying to figure out Ichiro Suzuki's batting average by looking at the batting average of the entire Seattle Mariners team, since the illegal votes were cast by an unrepresentative sample of each precinct's voters, and might be as different from the average voter in the precinct as Ichiro was from the rest of his team. The judge determined that the challengers' argument was an ecological fallacy and rejected it.

In my previous post "Big data: problems of correlation, bias, and machine learning" I showed the graph of statistically significant correlation between chocolate consumption per capita and number of Nobel laureates in a country to illustrate the fact that "correlation does not imply causation". The information for it came from a paper published in the New England Journal of Medicine in 2012. It claimed that chocolate consumption could enhance cognitive function. The basis for this conclusion was that the number of Nobel Prize laureates in each country was strongly correlated with the per capita consumption of chocolate in that country.  However a recent paper by Velickovic in Scientific American (What Everyone Should Know about Statistical Correlation: A common analytical error hinders biomedical research and misleads the public, January-February 2015 - http://www.americanscientist.org/issues/pub/2015/1/what-everyone-should-know-about-statistical-correlation/99999) the author and his commentators were unsure of if the authors of the New England Journal of Medicine article made a blunder or if they were writing it tongue-in-cheek.

The interesting point, however, is that apart from being educational as illustrating correlation does not imply causation, the analysis in the New England Journal of Medicine article could be seen as an example of ecological fallacy. Velickovic writes:

... the authors fell into an ecological fallacy, when a conclusion about individuals is reached based on group-level data. In this case, the authors calculated the correlation coefficient at the aggregate level (the country), but then erroneously used that value to reach a conclusion about the individual level (eating chocolate enhances cognitive function). Accurate data at the individual level were completely unknown: No one had collected data on how much chocolate the Nobel laureates consumed, or even if they consumed any at all. I was not the only one to notice this error. Many other scientists wrote about this case of erroneous analysis. Chemist Ashutosh Jogalekar wrote a thorough critique on his Scientific American blog The Curious Wavefunction , and Beatrice A. Golomb of University of California, San Diego, even tested this hypothesis with a team of coauthors, pointing out that there is no link.