Blog Paper

Posted in Biographical, The Universe and Stuff with tags , , on April 12, 2016 by telescoper

I don’t often blog about my own research. To be honest that’s partly because I don’t get much time to do any. Fortunately, however, I have an excellent postdoctoral research assistant (Dipak) and some excellent collaborators. Anyway, I just heard yesterday that the following paper has been accepted for publication in the Journal of Cosmology and Astroparticle Physics (JCAP):

Munshi

It’s not exactly a light read – it’s 32 pages long – but at least it gives the non-cosmology readers of this blog an idea of my research interests. Hopefully it won’t be too long before we can apply techniques such as those described in the above paper to real data!

Hopefully also in future I’ll be able to persuade my co-authors to submit to the Open Journal of Astrophysics!

Fear, Risk, Uncertainty and the European Union

Posted in Politics, Science Politics, The Universe and Stuff with tags , , , , , , , , , on April 11, 2016 by telescoper

I’ve been far too busy with work and other things to contribute as much as I’d like to the ongoing debate about the forthcoming referendum on Britain’s membership of the European Union. Hopefully I’ll get time for a few posts before June 23rd, which is when the United Kingdom goes to the polls.

For the time being, however, I’ll just make a quick comment about one phrase that is being bandied about in this context, namely Project Fear.As far as I am aware this expression first came up in the context of last year’s referendum on Scottish independence, but it’s now being used by the “leave” campaign to describe some of the arguments used by the “remain” campaign. I’ve met this phrase myself rather often on social media such as Twitter, usually in use by a BrExit campaigner accusing me of scaremongering because I think there’s a significant probability that leaving the EU will cause the UK serious economic problems.

Can I prove that this is the case? No, of course not. Nobody will know unless and until we try leaving the EU. But my point is that there’s definitely a risk. It seems to me grossly irresponsible to argue – as some clearly are doing – that there is no risk at all.

This is all very interesting for those of us who work in university science departments because “Risk Assessments” are one of the things we teach our students to do as a matter of routine, especially in advance of experimental projects. In case you weren’t aware, a risk assessment is

…. a systematic examination of a task, job or process that you carry out at work for the purpose of; Identifying the significant hazards that are present (a hazard is something that has the potential to cause someone harm or ill health).

Perhaps we should change the name of our “Project Risk Assessments” to “Project Fear”?

I think this all demonstrates how very bad most people are at thinking rationally about uncertainty, to such an extent that even thinking about potential hazards is verboten. I’ve actually written a book about uncertainty in the physical sciences , partly in an attempt to counter the myth that science deals with absolute certainties. And if physics doesn’t, economics definitely can’t.

In this context it is perhaps worth mentioning the  definitions of “uncertainty” and “risk” suggested by Frank Hyneman Knight in a book on economics called Risk, Uncertainty and Profit which seem to be in standard use in the social sciences.  The distinction made there is that “risk” is “randomness” with “knowable probabilities”, whereas “uncertainty” involves “randomness” with “unknowable probabilities”.

I don’t like these definitions at all. For one thing they both involve a reference to “randomness”, a word which I don’t know how to define anyway; I’d be much happier to use “unpredictability”.In the context of BrExit there is unpredictability because we don’t have any hard information on which to base a prediction. Even more importantly, perhaps, I find the distinction between “knowable” and “unknowable” probabilities very problematic. One always knows something about a probability distribution, even if that something means that the distribution has to be very broad. And in any case these definitions imply that the probabilities concerned are “out there”, rather being statements about a state of knowledge (or lack thereof). Sometimes we know what we know and sometimes we don’t, but there are more than two possibilities. As the great American philosopher and social scientist Donald Rumsfeld (Shurely Shome Mishtake? Ed) put it:

“…as we know, there are known knowns; there are things we know we know. We also know there are known unknowns; that is to say we know there are some things we do not know. But there are also unknown unknowns – the ones we don’t know we don’t know.”

There may be a proper Bayesian formulation of the distinction between “risk” and “uncertainty” that involves a transition between prior-dominated (uncertain) and posterior-dominated (risky), but basically I don’t see any qualititative difference between the two from such a perspective.

When it comes to the EU referendum is that probabilities of different outcomes are difficult to calculate because of the complexity of economics generally and the dynamics of trade within and beyond the European Union in particular. Moreover, probabilities need to be updated using quantitative evidence and we don’t actually have any of that. But it seems absurd to try to argue that there is neither any risk nor any uncertainty. Frankly, anyone who argues this is just being irrational.

Whether a risk is worth taking depends on the likely profit. Nobody has convinced me that the country as a whole will gain anything concrete if we leave the European Union, so the risk seems pointless. Cui Bono? I think you’ll find the answer to that among the hedge fund managers who are bankrolling the BrExit campaign…

 

 

This is not a spiral

Posted in Art on April 11, 2016 by telescoper

Talking of art, feast your eyes on this image from a really interesting website called CosmusUp:

Spiral

Picture courtesy of CosmosUp

It’s a stunningly convincing optical illlusion. You can see how it works if you draw an horizontal line through the centre. The small black and white squares at the corners of the larger ones are aligned differently on the inside and outside of each ring and alternate in orientation as you go around. This creates a pattern of black squares that appears to bend, creating the impression of an anti-clockwise spiral.

Why? You endeavoured to embroil me with weomen…

Posted in History with tags , on April 10, 2016 by telescoper

Here’s a post about an episode in the life of Sir Isaac Newton which I first came across when reading about Samuel Pepys. Many assume that Newton’s behaviour was a result of mental illness on his part, but that’s by no means clear. I can think of many possible reasons why he might have acted the way he did, including that he just found the behaviour of other people too perplexing…

corpusnewtonicum's avatarCorpus Newtonicum

Why. It is a word that I frequently entertain when I study Isaac Newton. There is no scientist about whom so much is written, yet I feel that we only know so little about the man. Most Newton biographers provide us with detailed descriptions of his life and works, using the abundance of source materials available: Newton’s correspondence, descriptions by himself and others of various episodes of his life, Trinity College and Cambridge University attendance records, and so on. Every biographer, in his own way, tries to understand some of the more poignant moments in Newton’s life. Likewise, many struggle.

View original post 1,034 more words

Constructed Universe

Posted in Art, The Universe and Stuff with tags , , on April 10, 2016 by telescoper

I saw this interesting piece “Constructed Universe” (1983) by Daniel Faust from the Metropolitan Museum of Art  via Twitter and it intrigued me enough to share it here, although some of you might think it’s just a load of balls.

image

Captain Black

Posted in Uncategorized on April 8, 2016 by telescoper

Too busy to blog today, so here’s a picture of Captain Black from the popular TV series Captain Scarlet.

image

If I understand my Twitter feed correctly he was on Newsnight tonight, using the pseudonym Conrad. Earth men, your time is at an end.

What does “Big Data” mean to you?

Posted in The Universe and Stuff with tags , , , , on April 7, 2016 by telescoper

On several occasions recently I’ve had to talk about Big Data for one reason or another. I’m always at a disadvantage when I do that because I really dislike the term.Clearly I’m not the only one who feels this way:

say-big-data-one-more-time

For one thing the term “Big Data” seems to me like describing the Ocean as “Big Water”. For another it’s not really just the how big the data set is that matters. Size isn’t everything, after all. There is much truth in Stalin’s comment that “Quantity has a quality all its own” in that very large data sets allow you to do things you wouldn’t even try with smaller ones, but it can be complexity rather than sheer size that also requires new methods of analysis.

Planck_CMB_large

The biggest event in my own field of cosmology in the last few years has been the Planck mission. The data set is indeed huge: the above map of the temperature pattern in the cosmic microwave background has no fewer than 167 million pixels. That certainly caused some headaches in the analysis pipeline, but I think I would argue that this wasn’t really a Big Data project. I don’t mean that to be insulting to anyone, just that the main analysis of the Planck data was aimed at doing something very similar to what had been done (by WMAP), i.e. extracting the power spectrum of temperature fluctuations:

Planck_power_spectrum_origIt’s a wonderful result of course that extends the measurements that WMAP made up to much higher frequencies, but Planck’s goals were phrased in similar terms to those of WMAP – to pin down the parameters of the standard model to as high accuracy as possible. For me, a real “Big Data” approach to cosmic microwave background studies would involve doing something that couldn’t have been done at all with a smaller data set. An example that springs to mind is looking for indications of effects beyond the standard model.

Moreover what passes for Big Data in some fields would be just called “data” in others. For example, the Atlas Detector on the  Large Hadron Collider  represents about 150 million sensors delivering data 40 million times per second. There are about 600 million collisions per second, out of which perhaps one hundred per second are useful. The issue here is then one of dealing with an enormous rate of data in such a way as to be able to discard most of it very quickly. The same will be true of the Square Kilometre Array which will acquire exabytes of data every day out of which perhaps one petabyte will need to be stored. Both these projects involve data sets much bigger and more difficult to handle that what might pass for Big Data in other arenas.

Books you can buy at airports about Big Data generally list the following four or five characteristics:

  1. Volume
  2. Velocity
  3. Variety
  4. Veracity
  5. Variability

The first two are about the size and acquisition rate of the data mentioned above but the others are more about qualitatively different matters. For example, in cosmology nowadays we have to deal with data sets which are indeed quite large, but also very different in form.  We need to be able to do efficient joint analyses of heterogeneous data structures with very different sampling properties and systematic errors in such a way that we get the best science results we can. Now that’s a Big Data challenge!

 

The Distribution of Cauchy

Posted in Bad Statistics, The Universe and Stuff with tags , , , , , on April 6, 2016 by telescoper

Back into the swing of teaching after a short break, I have been doing some lectures this week about complex analysis to theoretical physics students. The name of a brilliant French mathematician called Augustin Louis Cauchy (1789-1857) crops up very regularly in this branch of mathematics, e.g. in the Cauchy integral formula and the Cauchy-Riemann conditions, which reminded me of some old jottings aI made about the Cauchy distribution, which I never used in the publication to which they related, so I thought I’d just quickly pop the main idea on here in the hope that some amongst you might find it interesting and/or amusing.

What sparked this off is that the simplest cosmological models (including the particular one we now call the standard model) assume that the primordial density fluctuations we see imprinted in the pattern of temperature fluctuations in the cosmic microwave background and which we think gave rise to the large-scale structure of the Universe through the action of gravitational instability, were distributed according to Gaussian statistics (as predicted by the simplest versions of the inflationary universe theory).  Departures from Gaussianity would therefore, if found, yield important clues about physics beyond the standard model.

Cosmology isn’t the only place where Gaussian (normal) statistics apply. In fact they arise  fairly generically,  in circumstances where variation results from the linear superposition of independent influences, by virtue of the Central Limit Theorem. Thermal noise in experimental detectors is often treated as following Gaussian statistics, for example.

The Gaussian distribution has some nice properties that make it possible to place meaningful bounds on the statistical accuracy of measurements made in the presence of Gaussian fluctuations. For example, we all know that the margin of error of the determination of the mean value of a quantity from a sample of size n independent Gaussian-dsitributed varies as 1/\sqrt{n}; the larger the sample, the more accurately the global mean can be known. In the cosmological context this is basically why mapping a larger volume of space can lead, for instance, to a more accurate determination of the overall mean density of matter in the Universe.

However, although the Gaussian assumption often applies it doesn’t always apply, so if we want to think about non-Gaussian effects we have to think also about how well we can do statistical inference if we don’t have Gaussianity to rely on.

That’s why I was playing around with the peculiarities of the Cauchy distribution. This distribution comes up in a variety of real physics problems so it isn’t an artificially pathological case. Imagine you have two independent variables X and Y each of which has a Gaussian distribution with zero mean and unit variance. The ratio Z=X/Y has a probability density function of the form

p(z)=\frac{1}{\pi(1+z^2)},

which is a Cauchy distribution. There’s nothing at all wrong with this as a distribution – it’s not singular anywhere and integrates to unity as a pdf should. However, it does have a peculiar property that none of its moments is finite, not even the mean value!

Following on from this property is the fact that Cauchy-distributed quantities violate the Central Limit Theorem. If we take n independent Gaussian variables then the distribution of sum X_1+X_2 + \ldots X_n has the normal form, but this is also true (for large enough n) for the sum of n independent variables having any distribution as long as it has finite variance.

The Cauchy distribution has infinite variance so the distribution of the sum of independent Cauchy-distributed quantities Z_1+Z_2 + \ldots Z_n doesn’t tend to a Gaussian. In fact the distribution of the sum of any number of  independent Cauchy variates is itself a Cauchy distribution. Moreover the distribution of the mean of a sample of size n does not depend on n for Cauchy variates. This means that making a larger sample doesn’t reduce the margin of error on the mean value!

This was essentially the point I made in a previous post about the dangers of using standard statistical techniques – which usually involve the Gaussian assumption – to distributions of quantities formed as ratios.

We cosmologists should be grateful that we don’t seem to live in a Universe whose fluctuations are governed by Cauchy, rather than (nearly) Gaussian, statistics. Measuring more of the Universe wouldn’t be any use in determining its global properties as we’d always be dominated by cosmic variance

The Insignificance of ORB

Posted in Bad Statistics with tags , , , on April 5, 2016 by telescoper

A piece about opinion polls ahead of the EU Referendum which appeared in today’s Daily Torygraph has spurred me on to make a quick contribution to my bad statistics folder.

The piece concerned includes the following statement:

David Cameron’s campaign to warn voters about the dangers of leaving the European Union is beginning to win the argument ahead of the referendum, a new Telegraph poll has found.

The exclusive poll found that the “Remain” campaign now has a narrow lead after trailing last month, in a sign that Downing Street’s tactic – which has been described as “Project Fear” by its critics – is working.

The piece goes on to explain

The poll finds that 51 per cent of voters now support Remain – an increase of 4 per cent from last month. Leave’s support has decreased five points to 44 per cent.

This conclusion is based on the results of a survey by ORB in which the number of participants was 800. Yes, eight hundred.

How much can we trust this result on statistical grounds?

Suppose the fraction of the population having the intention to vote in a particular way in the EU referendum is p. For a sample of size n with x respondents indicating that they hen one can straightforwardly estimate p \simeq x/n. So far so good, as long as there is no bias induced by the form of the question asked nor in the selection of the sample which, given the fact that such polls have been all over the place seems rather unlikely.

A little bit of mathematics involving the binomial distribution yields an answer for the uncertainty in this estimate of p in terms of the sampling error:

\sigma = \sqrt{\frac{p(1-p)}{n}}

For the sample size of 800 given, and an actual value p \simeq 0.5 this amounts to a standard error of about 2%. About 95% of samples drawn from a population in which the true fraction is p will yield an estimate within p \pm 2\sigma, i.e. within about 4% of the true figure. In other words the typical variation between two samples drawn from the same underlying population is about 4%. In other other words, the change reported between the two ORB polls mentioned above can be entirely explained by sampling variation and does not at all imply any systematic change of public opinion between the two surveys.

I need hardly point out that in a two-horse race (between “Remain” and “Leave”) an increase of 4% in the Remain vote corresponds to a decrease in the Leave vote by the same 4% so a 50-50 population vote can easily generate a margin as large as  54-46 in such a small sample.

Why do pollsters bother with such tiny samples? With such a large margin error they are basically meaningless.

I object to the characterization of the Remain campaign as “Project Fear” in any case. I think it’s entirely sensible to point out the serious risks that an exit from the European Union would generate for the UK in loss of trade, science funding, financial instability, and indeed the near-inevitable secession of Scotland. But in any case this poll doesn’t indicate that anything is succeeding in changing anything other than statistical noise.

Statistical illiteracy is as widespread amongst politicians as it is amongst journalists, but the fact that silly reports like this are commonplace doesn’t make them any less annoying. After all, the idea of sampling uncertainty isn’t all that difficult to understand. Is it?

And with so many more important things going on in the world that deserve better press coverage than they are getting, why does a “quality” newspaper waste its valuable column inches on this sort of twaddle?

Sonnet No. 98

Posted in History, Poetry with tags , , , on April 5, 2016 by telescoper

It’s been a while since I posted any of Shakespeare’s sonnets. A brief mention on the radio this morning that William Shakespeare died 400 years ago this month convinced me to rectify that omission and, since it is April, I thought I’d put up this one, No. 98. As with the rest of the first 126 of these poems, it is addressed by the poet to a “fair youth”, i.e. from an older man to a younger one. These sonnets deal with such themes as love, beauty, mortality, absence and longing, framed by the affectionate relationship between two men of very different ages:

From you have I been absent in the spring,
When proud-pied April, dressed in all his trim,
Hath put a spirit of youth in every thing,
That heavy Saturn laughed and leaped with him.
Yet nor the lays of birds, nor the sweet smell
Of different flowers in odour and in hue
Could make me any summer’s story tell,
Or from their proud lap pluck them where they grew.
Nor did I wonder at the lily’s white,
Nor praise the deep vermilion in the rose;
They were but sweet, but figures of delight
Drawn after you, you pattern of all those.
Yet seemed it winter still, and you away,
As with your shadow I with these did play.