Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Tuesday, January 27, 2026

AutoArt folder distribution

 kw: analytical projects, art generation, ai art, statistics, statistical distributions, lognormal, scale free

I began using art generating software in November 2022, when DALL-E2 became available. Since then, I've enjoyed having a series of art generating "engines" available, including numerous engines (called "models") in the aggregators Leonardo AI and OpenArt. As often as I can, I generate images for this blog; in some cases, I download images I find on the Internet. However, my primary artistic pastime is creating images of things and scenes I imagine.

Just in the past few days I was inspired by a heavy snowfall to find short poems about snow, and use them to create wintry images. This image was drawn by Nano Banana Pro, under the Leonardo AI umbrella with "None" as the style; that is, native NB Pro. The aspect ratio was set to 16:9. It displays the entire poem, something NB Pro can do better than any other art engine I have found. The prompt was "Watercolor painting evoked by a poem:", followed by the text of the poem "The First Snow" by Charlotte Zolotow.


The image is particularly evocative in shifting to an exterior view as the window dissolves. I suspect there are a number of images that use this device in the training material for NB Pro.

When I made signed versions of this and several others that were generated in the same session, to be included in a folder for a "screen saver" slide show, I began thinking about the various numbers of different image types I've created in the past three-plus years. Last year I went through my (poorly organized) folder stack of "AutoArt" and reorganized it into 35 categories, each in its own folder. To date, there are 1,472 signed images in 35 folders containing between two and 405 images. My inner statistician began to stir…

The image below shows two analyses of the statistical distribution of the numbers of files in these folders.

Charts like these make it quite evident which statistical treatment is appropriate to a particular set of data. I'll explain what these charts mean and how they were created.

"Scale Free" is a type of power law distribution related to the Pareto distribution. It is easy to analyze, which makes it popular. To analyze a series of numbers graphically in Microsoft Excel:

  • Enter the numbers in column B, starting in cell B2.
  • Put an appropriate header in cell B1
  • Highlight these data (B1:B36 in this case)
  • Sort from largest to smallest, using the Sort & Filter section under Editing in the Ribbon.
  • Enter 1 in cell A2 and 2 in A3.
  • Put a header in cell A1; I usually put "N".
  • Highlight cells A2 and A3.
  • Double-click the fill handle at the lower right of A3. This will fill the rest of the column with numbers in order, as far as the data goes in column B. In this case, we get numbers from 1 to 35.
  • Highlight these two columns to the end of data. In this case, from A1 to B36.
  • In the Ribbon, use Insert and in the Charts section, select the icon showing scattered dots with axes; this is X-Y Chart.
  • The title of the chart is whatever the header text is in B1. Edit as you wish.
  • Double-click one of the axes to open the Format dialog.
  • Click Logarithmic Scale near the bottom of the menu.
  • Click the other axis and also click Logarithmic Scale. This is now a log-log chart.

The result will be similar to the upper chart. Now for the lognormal analysis, beginning with these two columns of numbers:

  • Insert a new column between A and B; this is the new column B.
  • In cell B1 enter a header such as "Prob.". You are going to create a probability axis.
  • In cell B2 enter this formula (where the largest number in column A is 35):

=NORM.S.INV((A2-0.5)/35)

  • Double-click the fill handle at the lower right of A2 to fill the column with the formula.
  • Highlight the data in B and C (B1 to C36 in this case).
  • Use Insert as before to create an X-Y Chart.
  • Edit the chart title.
  • Note that the vertical axis is now centered above the zero. 
  • Assuming the Format dialog is still open, click the horizontal axis.
  • In the middle of the menu in the section "Vertical Axis Crosses", click the bubble at "Axis Value".
  • Enter "-3".
  • Click the vertical axis and click Logarithmic Scale. This is now a log-probability chart.
  • If you want the markers to be a different color, click one of them. The Format Data Series menu appears at the right.
  • Select the icon of a paint bucket pouring paint.
  • Click the Marker tab
  • For both Fill and Border, select the color you want.

This will be similar to the lower chart. For the data I used, the chart shows the points scattered approximately along a straight line. By contrast, in the upper chart there is a definite downward bend. In a log-log chart such a shape is diagnostic that the distribution is not scale free, but is more likely to be lognormal, or even normal (Gaussian). In this case, the second chart shows that lognormal is a good model of the data distribution.

This is an illustration of the Theory of Breakage, formally described by A.N. Kolmogoroff in 1941. When an area is divided (US state or county areas are good examples), the distribution is lognormal. When a sheet of glass is broken, the weights of the pieces also have a lognormal distribution (I've done this experiment). Some recent publications claim that a theory of breakage produces a power law distribution, but this is false. Certain phenomena in nature tend to be normally distributed. The classic example is the height of adult men, or of women (but not both) in a population, such as the residents of a particular town or county. However, most phenomena produce groups of measurements that are lognormally distributed, in which the logarithm of the quantity being measured is distributed as a normal, or Gaussian, curve.

I could go further into this, but this is enough for the purpose of this post.

Sunday, June 14, 2020

Stats for Blogger - New Look

kw: blogs, blogging, statistics, analytics

I leave my Blogger tab at the Stats section, so when I start the browser that is the first thing I see. A few weeks ago Google announced it was upgrading Blogger. I had the chance to try it out, and I saw that in a few months nearly everyone will be transitioned anyway. I generally like it. It operates more like a mobile app, and I can live with that. As to the statistics, however, I wasn't so pleased.

The nice four-pane Overview is gone. One may gather them piecemeal, and more items are available than before. However, I liked the hour-by-hour detail in the Weekly view of activity. Now it is daily, unless you look only at the past 24 hours.

I primarily look at total volume and the worldwide distribution, particularly when I'm commenting my amusement with spider scanning from Russia or Asia. As it happens, today there was a new spike in the weekly chart that mostly occurred yesterday. Here are the charts for the past week and month:


For reference, prior to the Covid-19 pandemic, for the past few years my daily Views were between 30 and 50, except when a spidering operation was hitting the blog 100 times or more daily. The bump after June 8 is likely to be another round of spidering. The Countries charts (7 days and then 30) give me a clue a clue or two:


Not many people even know where Turkmenistan is. The 7 day chart shows it clearly; it is the third country to the northwest of India. That small country was the source of some earlier spidering, and it appears to be the source of the "bump" in views over the past few days. 

Interestingly, another section of the Stats usually shows that Chrome is most frequently used, but during the past week and the past day, it has been MSIE, which is how Internet Explorer is reported by Blogger Stats. 

There is an option to see even more using Google Analytics, which I haven't delved into. Perhaps there I can see or build a chart showing hourly views over a week or longer, but that a project for another day.




Saturday, May 09, 2020

The usual is less usual than we think

kw: analysis, statistics, black swans, statistical distributions

Alternative Title: Thousand-year Floods Occur Too Often

Just to set expectations ahead of time: The aim of this article is to analyze activity for a stock in the U.S. stock market. However, the phenomena of interest here apply in numerous areas, including flood control, a fun place to begin.

This picture shows the Sorlie Bridge linking Grand Forks, ND with East Grand Forks, MN, during the flood of 1997. It's what happens when you design for a 100-year flood, and it has been 170 years since the last big flood. The peak flow was only 5% greater than the design criterion.

What would a 1,000-year flood look like? Sooner or later it'll happen. Some interesting statistics are involved.

Here we have a log-Pearson plot of gauged "floods" on a river in Australia. AEP is "Annual Exceedance Probability". I was taught how to make such plots during a summer project in graduate school. The data are the highest flow recorded every time the river overflows its inner banks. There are 46 such events charted here. I infer from the rightmost datum that the record goes back about 80 years. A "flood" as thus defined will happen every year or two.

Extending the coordinates of this chart I estimate that a 1,000-year flood would have a peak flow of between 350 and 400 on the scale shown (I infer m3/s; in ft3/s: 12,000-14,000), not quite twice the flow of a 100-year flood.

Quantitative hydrology doesn't extend back even 200 years. How can we validate the use of log-Pearson analysis in extrapolating for 1,000-year floods? There are also at least two other analysis methods in use. How could we validate any of them?

Nearly forty years ago, a fellow student of mine in graduate school worked out such a validation, for at least one "creek" in the Black Hills area. In this view from Google Maps, the creek mouth is off the left. The image width is about a quarter mile.

The boulders seen here are not glacial erratics. The student, Bill, measured lichen patches on hundreds of the boulders, choosing only those larger than he was. He was able to discern, by the size of the largest patch of a certain kind of lichen on each boulder, when they were last tumbled down the creek and washed out onto the plain, sometimes as far as a couple of miles. He determined the outwash area for twenty large floods from that creek. I saw a map he made of those outwash areas.

References on lichenometry helped him obtain approximate dates for the floods. All had occurred in the past 5,000 years. As a student of hydrology, he was able to determine the flow volume needed to move a car-size boulder a half mile to a mile from the creek mouth. Here's the kicker: a chart of historical hydrology measurements on that creek indicate that every one of the twenty floods was at least a 1,000-year flood. But they were actually something like 250-year floods! (or, actually, scattered out from 250 years and upward.) Such data didn't fit into the historical analysis. To be clear about what follows, the log-Pearson analysis is not related to the Normal Distribution, nor to Lognormal statistics.

Statistical analysis in many areas depends heavily on the Normal Distribution, AKA Gaussian. As it happens, this is frequently a very good model. However, even the textbook example, human height, isn't as "normal" as it might seem. In this blog post by John D. Cook, he shows that extremes of height and of shortness are not well predicted by a normal distribution with a standard deviation of 2.5 inches, the criterion that fits well for about 95% of the data for men and for women (which must be analyzed separately). By that measure, there should be no more than one man taller than seven feet, or shorter than 4'-5", nor one woman taller than 6'-9", or shorter than 4'-1", among our entire world population of more than 7 billion. But the NBA is full of seven-footers (more than 40, as counted in 2018), and there are numerous "little people", particularly in entertainment, ranging from four feet down to just above two feet tall. The shortest adult woman on record measured nineteen inches. Data such as these fit better a "fat tailed" distribution, for which more extreme values are more frequent than a Normal analysis would predict.

For events such as floods, it is likely that the excess frequency of extra-large floods is due to a different climate regime from the usual. For height measurements on people, different populations with distributions account for some of the "excess" variation, while medical issues such as glandular gigantism or dwarfism also account for some. These kinds of considerations indicate that we need to study a broader scope of phenomena. If an analysis doesn't include everything relevant, our model is incomplete. But what of "surprises" in a data set that is thought to be well-behaved?

These are examples of black swans. I reviewed The Black Swan by N.N. Taleb thirteen years ago. The book's theme is that "uncommon" events aren't so uncommon. He applied findings such as these to investing. You may recall there had been a significant bear market a few years earlier, starting in 2000. The one that followed the book by a year (2008-09) made Taleb look like a prophet. Now we are in the middle of another one.

One claim in The Black Swan is that daily moves on stock issues follow a Cauchy distribution. It is described as the ratio of the sine of one uniform random variable divided by the cosine of another. Mathematically, that is the tangent of a uniform random variable over the interval [-π/2, π/2], but not including the end points, for which its value is infinite. The Cauchy distribution is typically described as a tangent function over the half circle.

That was something I could test. The Cauchy distribution is very fat-tailed or, viewed another way, it middle part very skinny compared to a normal distribution. It looks like this. We'll see how a normal curve looks in a moment.

Does the distribution of daily closing values of stocks in the market look like this? I chose the most stable stock in America, AT and T, to test this. I used data for the past 37 years (everything available through Yahoo! Finance). I used what they call "Adjusted Close", which factors in the values of stock splits and dividends. I analyzed the entire period, and two sub-periods, the ten years from 1/1/1984 to 12/31/1993 and the two years from 1/1/2018 to 12/31/2019. Note: I'll call this company "ATT" to avoid the ampersand, which is not well behaved in HTML text.

This is the shape of the distribution of daily moves, expressed in percent, and two normal distribution curves. As the legend indicates, the blue curve shows the recent two years, the green curve shows the early ten years (which includes the crash of 1987), and the red curve shows all 37 years. The dashed magenta line is a normal curve with the same standard deviation as the 37-year curve, and the thin black line is a normal curve that fits the bulk of the data, but not the tails. I hope you can see that a normal curve looks broader than the Cauchy curve in the chart just above. The range -10% to +10% doesn't quite encompass all the data; there are three data points further to the right and one further to the left…out of nearly 9,200 daily motions.

It may take some peering, but it is also visible that the black curve outside the range [-2%,+2%] lies below the red, green, and blue curves. Those small amounts are the "fat" in the fat tails of those three distributions. Their very similar shapes also illustrate why I call ATT so stable. Through thick and thin it preserves a certain character of daily responses to market pressures. I do not intend to compare ATT to another stock here; that is for another day, another post.

It will perhaps be easier to see the implications of these data by looking at cumulative distribution functions (CDF's).

The x-axis of this chart is in units of standard deviation. The straight, black dashed line shows how a normal distribution would plot as a straight line. The ATT data going above at the left and below at the right are the fat tails of their distribution.

Looking at the two charts above, one may suspect that the ATT CDF is not as fat-tailed as a Cauchy distribution. The next chart shows this:

The x- and y- axes are in different units because I didn't normalize them. However, the very flat shape of the curve between -2 and +2 standard deviation units shows that the tails are much fatter than those for the ATT data.

Stock market motions, at least for this stock, are thus not as extremely variable as a Cauchy-distributed function.

I must at this point make an aside regarding lognormal distributions. My geologic background as an undergraduate emphasized sedimentology. One tool used to study sediments such as a sandy soil is to shake a sample through a stack of sieves and make a frequency distribution of the result, by weight. The sieves are carefully crafted to separate particles, at each step, that are 1.414 (i.e., √2) smaller in diameter. The frequency distribution so produced is a lognormal analysis. A frequency plot for a well-developed sediment from a single source will look like a normal curve, except that the x-axis is the logarithm of particle diameter. Over the years I found that many natural phenomena that have a wide range are lognormally distributed. We can check for lognormality by charting on log-probability coordinates.

I decided to see if the distribution of stock price daily moves might be related to a lognormal distribution. I squared the ATT data; thus all the results are positive, no longer both up and down. This is a CDF of the result, with the vertical axis logarithmic.

If the data closely followed the black dashed line, they would be lognormally distributed. The data charted at 1E-08 were actually zero, but that cannot be shown on a log-probability chart. What is the big dip leading up to them? This is a consequence of the size of a penny! I call it the "penny effect". During part of the interval stocks were traded only in increments of 1/16 of a dollar, then in increments of certain fractions of a penny, and since then in whole cents.

If the stock is trading for $10, and the next day the closing price is $10.01, the difference between the two is 0.1%, or 0.001 (10.01/10.00 - 1). Squaring this produces 0.000001. When the adjusted closing price is closer to $5, as it was in the early 1980's, a penny difference is 0.002, which squares to 0.000004. This, and a psychological effect that many people avoid prices so close to a prior day's price all led to the drop-off seen here. Since this is a pretty good fit to lognormality, other than the "penny effect", I call the distribution of stock price moves "Square Root of Lognormal."

Let's look at the squared Cauchy data. It is a continuous distribution with no pennies to worry about, so there is no drop-off. But the data are not lognormal either. The black dashed line is lognormal. The tails of this distribution are fatter than the tails of a lognormal distribution, much fatter. From experience I can tell you, that's pretty hard to do!

It so happens that this distribution is ill-behaved no matter how you transform it. In terms only a statistician could love, the Cauchy distribution has no "moments". That means that trying to measure anything other than its average value is meaningless. You can calculate a standard deviation for it, for example, but it has no meaning.

How does this relate to black swans? In stock market terms, if the daily moves were truly distributed according to Cauchy statistics, huge moves would be even more common than they are. However, the Square Root of Lognormal function that daily prices do seem to follow is not nearly so extreme, but it does have many more large moves, up or down, than a normal model would lead someone to believe. The standard deviation of the ATT data is about 1.5%. That means that, as seen in the chart above titled "Daily Change, all", none of the daily price moves should have been outside the bounds of [-4%, +4%]. However, on 250 occasions, a daily move was outside this range. That's only 2.7% of all the data. However, 250 times in 37 years means that six or seven times yearly, ATT's stock price moved by more than a normal model would predict is even possible during a run of nearly 9,200 market days.

Conclusion? To better predict the range of variation for a particular stock, list the daily moves over a period of a few years, square them, and plot as a CDF. From this you can determine a log-standard deviation, and its square root will yield a much better parameter of variation for the stock's behavior.

This last chart is a frequency diagram of the squared data for ATT, and the horizontal axis is logarithmic. The blip at the left is the "penny effect". Other than that, and a bit of skew, it looks a lot like a normal distribution; the choice of axis shows its affinity to a lognormal distribution.

Whether this will lead to a more robust way to set an investment strategy is anyone's guess!

Friday, November 23, 2018

Removing the straitjacket of non-causation in statistics

kw: book reviews, nonfiction, statistics, probability, causation, mathematics

In 1926, during the height of the eugenics movement in the U.S., a researcher who has been nearly forgotten studied the relationship between the intelligence of children and that of parents. This is the core debate, even today, regarding the "nature-nurture" dichotomy. Which is more important, upbringing or inheritance?

Step back a minute, and consider, with the current popularity of "big data", how this might be tackled. It is no longer difficult to gather enormous amounts of data regarding the IQ of numerous children, adults, and societal indicators such as neighborhood of residence. Do all the math you might wish, with regressions and correlation diagrams, and what might you find? No doubt some kind of correlation will show up, perhaps very obviously. But what does it mean? What has "caused" the greater intelligence of some children, and the lesser intelligence of others?

The word "cause" was forbidden in statistical monographs for decades. For many researchers even today, the mantra (I chose that word with malice aforethought) is, "Correlation does not imply causation." While this is indeed true, even a tautology, it is not all there is to it. We naturally think of nearly everything in cause-and-effect terms, and work done in the past couple of generations now makes it possible for researchers to discuss causes without losing tenure, grants, etc.

For the young researcher, Barbara Burks, the mantra was nonsense. She sought causes. To this end, she gave IQ tests to every member of 204 households that included foster children, and 105 households without foster children. For 1926, this was pretty big data. The choice of studying both foster children and natural children along with the adults was clever. Even more clever was the little diagram she used to analyze her results:

The arrows imply causation. Here, the "X" factor that might influence both the level of intelligence of the child, and the social status of the household, was thought to be the "heritage", including genetic inheritance, of the family. The parents, in whatever measure they benefit (or not) from "heritage", will have their own X factor, which could have been added as Y, off to the left perhaps.

Note that two of the arrows have heads at both ends. This indicates feedback effects between the social status and the intelligence of all members of the family (I imagine a family of "ordinary" intelligence having a very, very bright kid, and this leading to an improvement in social standing, for example).

Such a diagram embodies a "causal model", in the terminology of Judea Pearl, in The Book of Why: The New Science of Cause and Effect. Such a diagram, and mathematical processes invented by Pearl and his students, provide what is missing in non-causal statistics: the understanding that some things really do cause other things. By the way, Ms Burks's conclusion: genetics provides 35% of the observed differences in the intelligence of children. This was a disappointment to eugenicists, including Ms Burks. In particular, Louis Terman, an inventor of the Stanford-Binet IQ test, also famous for his "genius" studies, rejected it outright. He was quite certain that genetics were behind "nearly all" the differences in IQ. One might imagine him, upon seeing her results and conclusions, huffing, "Impossible!"

This reminds me of the first of Arthur Clarke's laws: "When a distinguished but elderly scientist states that something is possible, he is almost certainly right. When he states that something is impossible, he is very probably wrong." Dr. Pearl writes of his own Odyssey of discovery. He did not come to causal reasoning easily. But now he and his students have developed "causal calculus", which is introduced in The Book of Why, by Judea Pearl and Dana MacKenzie.

I must confess, though my long career as a scientific programmer led me into statistical work again and again, I never became comfortable with the formulas of probability. When I see the term P(Y|X), I have to think a moment to get my head around, "The Probability of Y occurring (or existing), given the occurrence (or existence) of X". In non-causal terms, you can freely substitute "daybreak" and "rooster crowing" for X and Y, either way: "The probability of daybreak, given that the rooster crowed" and "The probability of a rooster crowing, given that day is breaking." Hold that thought.

Dr. Pearl has added the "do" operator, which implies an intervention, so that P(Y|do(X)) means "The probability of Y occurring, given that the intervention X was made, compared to X not being done". This is the reasoning behind the randomized controlled trial (RCT) in medicine, but it was not stated in a formula before. Indeed, in older medical journals the authors use all kinds of locutions and verbal gymnastics to avoid saying, "Medicine X caused a Z% reduction in death rate due to disease Y". Many still do so.

Thus far, I can follow along. Dr. Pearl freely ignores the folklore that each mathematical expression used in a book reduces its audience by half. Now, I like math, but it would take a great deal of study for me to become conversant with causal calculus. In an example called "DO-CALCULUS AT WORK" on page 236, we find expressions such as
Σt P(c|do(s),do(t))P(t|do(s))
This is the second of seven formulas in a derivation. At that point I realized I probably ought to devote my few remaining years to something besides learning how to not only parse such statements, but to create and perform them!

Dr. Pearl's work has great benefits for those researchers who can wrap their minds around these concepts and formalisms. For example, the decades-long struggle to determine to what extent smoking causes lung cancer, the subject of a major chapter, was undertaken in the face of determined and well-funded opposition to the concept, but might have been shortened to a year or a few years if causal language had been allowed. This stricture was as if the scientists studying smoking and cancer, those who were not in the pay of the tobacco companies, tied both hands behind their backs and had to perform their work with their toes and tongues. Now causal language is out of the closet.

An early chapter discusses the Ladder of Causation, from Association (what we observe), to Intervention (what we do to see what happens), to Counterfactuals (what we imagine might happen if X were not so). It appears that only humans can perform counterfactual reasoning, such as, "Will the day break if we get rid of all the roosters?" or, as a song says, "What if we gave a War and nobody came?"

We can't always figure out what is a cause and what is an effect. But where we can, the language of causation helps us model an event, such as by the use of a diagram such as the one above. Also, Do-Calculus now provides a mathematical way to treat cause and effect in a meaningful and quantitative  way. It adds power to Design of Experiments logic, so that a researcher is more likely to correctly determine the appropriate set of causative factors and winkle out just how important each is, in producing the effect being studied.

As difficult as the reading was, due only to my unfamiliarity with the jargon and formulas, reading the book was very enjoyable. The winding path Dr. Pearl took to get past the hamstrung statistical reasoning of half a century ago, on through Bayesian analysis, and on to develop causal reasoning in a formal way, with the appropriate formalisms of the mathematical language of Do-Calculus, make for a quest saga every bit as gripping as the search for a hidden city.

Friday, October 31, 2014

The Stats are out to get you

kw: book reviews, nonfiction, statistics, logical fallacies

I reckon there are a few hundred books with subjects similar to the classic How to Lie With Statistics by Darrell Huff. They are really self-help aimed at helping us resist arguments made using flawed, or fraudulent, statistics. Now I find a book aimed at those who might use statistics to make an argument, to avoid fooling themselves: Standard Deviations: Flawed Assumptions, Tortured Data, and Other Ways to Lie With Statistics by Gary Smith.

As I began to read, I remember thinking, "He ought to title it Nonstandard Deviations", but I soon realized that proper statistical thinking is so rare, even among scientific writers, that the deviations the book presents are indeed standard practice. It is trouble enough that cynical marketers and politicos are using statistics fraudulently to deceive us; the larger problem is how many different ways proponents can lie to themselves!

The key chapter is #2: "Garbage In, Gospel Out". Although there are 16 more chapters exposing at least as many errors of statistical logic, and a great summary titled "When to Be Persuaded and When to Be Skeptical", those 16 chapters show all the common ways of using numbers to create nonsense. Several are based on faulty assumptions about trends.

We live in a world with two kinds of time. We are embedded in the cycles of the seasons: days, weeks, months, years, decades and centuries. Every day the sun rises, crosses the sky, and sets (unless you live in the high Arctic or on Antarctica). Every year the seasons come and go in sequence. Our most basic, gut-level experience of time is cyclic. But we also have linear time. Plant a tree and it grows taller every year. Some trees keep that up for a thousand years or more. We see continual population growth in most countries and in the whole world (Germany, France and a few other countries have reducing populations, but we don't think about that much). We have ancestors in the past, going all the way back to Noah or Adam or whatever progenitor we believe in; we also expect to have descendants going pretty much forever into the future, or at least "until Kingdom come".

We are less familiar with linear time, though, and tend to think linear trends can continue without limit. The key to unlocking this quandary is to realize that time itself is linear, but things that happen in time have a beginning and an end, and typically rise and fall in between. An evangelical "young-Earth" Christian believes in a strictly limited span of time, beginning about 6,000 years ago, maybe as much as 10,000 years, and ending within the next hundred or so. A purely agnostic scientist who knows cosmology believes time, or at least the current phase of phenomena in time, began 13.8 billion years ago, but there are a few hundred competing theories about when or whether it will end. Nonetheless, the end of life on Earth is pretty well understood to be a billion years from now, because the Sun is slowly heating up, and the end of the Earth itself will follow 3-4 billion years later, when the planet is crisped and perhaps evaporated by the Sun's red giant phase.

A few billion years is plenty of time enough for some trends to go along and go along for a long, long time. The human population of Earth has been steadily increasing for at least the last 50,000-70,000 years. The hope of many "zero population growth" advocates is that human population will stabilize within the coming 50-100 years, and even begin to shrink. However, if you want to start a business that requires population growth to continue, and you're satisfied with a run of 20-40 years, go for it. It'll take at least that long for growth to slow to the point you'd have a hard time keeping the business going. But the usual business cycle is about 6 years. Plan on some kind of downturn in the next few years. If you survive that into the next cycle, you just might keep that business going until your kids are grown.

The author exhorts us, again and again, to think. The motto of IBM used to be "THINK". Statistical reasoning doesn't come naturally, even for statisticians. He uses humorous stories of "experts" who ran afoul of their own wishful thinking. It takes a lot of data to prove a statistical inference. A key concept of statistics is "significance". Scientific journals are filled with articles that employ statistical tests and declare that some finding is "significant to the x% level". That "x%" is typically 95%, which is frequently stated as 0.95. That means that there is at least a 95% chance that the "significant" finding is true. But there's a 5% chance that it is not true.

Let's suppose that every scientific experiment resulted in a publication telling the results. Further, let's suppose that only one in ten reported "significant" results. Think a minute: why do scientists use statistics? It is because they don't get a clear-cut result. If using widget A was always lots better than using widget B, statistics would not be needed. The article could be very short: "In 100 trials, widget A always did a better job than widget B". Then you'd question whether the scientist were sane: after about 10 trials, you can stop already! That depends on just how much A was better.

More typically, there is overlap. Suppose that some scoring method showed that A is better 64% of the time. If that was 64 out of 100, it is probably a significant result, but if it was 16 out of 25, you could be in trouble with the law of small numbers. This is analogous to flipping a coin 25 times to see if it is a fair coin. You get 16 heads. How likely is that? Many people think there ought to be a nearly exact even split, either 12 or 13 heads. Here is how to analyze it:

  • For 25 coin flips, there are 33,554,432 possible outcomes, from all heads to all tails, but in 33,554,430 out of 33,554,432 cases, it'll be some mix. 
  • An outcome of 12 heads occurs 5,200,300 different ways, as does an outcome of 13 heads. Together they total 30.1% of all outcomes. That is, intuition is correct less than 1/3 of the time!
  • An outcome of exactly 16 heads occurs 2,042,875 different ways. Thus, the chance you'll get 16 heads is 6.1%. 
  • There is thus a 6.1% probability that this outcome indicates there is no difference between the two widgets. The result is not sufficiently "significant".

This analysis was done using Pascal's Triangle, and there is plenty of software out there that can do such an analysis. You just have to know enough to set it up. By the way, if this were the result of 50 trials, with 32 heads, you'd have a different conclusion. Firstly, getting exactly 32 heads in 50 throws occurs 1.6% of the time. You could also say that getting at least 64% occurs 3.2% of the time by chance alone. Thus, the "significance level" is 96.8%, which is better than 95%, so there is support to say that widget A is actually better than widget B.

This is not a lock. Remember, I posited a world in which every result is published, whether favorable or unfavorable to the initial conjecture. Do you think negative results are published? Nearly never!! So in a world of "publish everything", if 1/10th report "significant" results, some of those are likely to be due to chance alone. Perhaps one in 20, or 2 of the original 100 articles. But in the real world, the proportion may be quite a bit higher. It is certain to be at least 1 in 20.

OK, that's a long-winded excursion into just one item that struck my fancy. As in most endeavors, there is a very short list of ways to do it right, and a near-infinite number of ways to go wrong. That's why we need to expose our ideas to a great variety of folks with different backgrounds and viewpoints. Many times, though, the proponent(s) of an idea will circulate only among those who think alike.

It is also shown that wanting a certain result is the most powerful enemy of truth. I recall an old story of someone seeking a simple answer, because he didn't know how to figure it for himself. He got a variety of answers from people he knew, until he asked a political lobbyist, who responded, "What do you want it to be?" Well, that joke may be more political than statistical, but it is sobering. No matter how much we may want this or that to be true, the actual case is the actual case, the truth is the truth, and will outlive you and your most heartfelt desire.

Wednesday, March 05, 2014

The Market is People

kw: book reviews, nonfiction, statistics, physics, stock markets

A financial market behaves like a small collection of quantum particles. This is my conclusion after decades of investing (sometimes lucky, sometimes not), and reading about them, from The Emergence of Probability and The Taming of Chance by Ian Hacking, to The Black Swan by Nassim Taleb and Beat the Market by Ed Thorp and Sheen Kassouf, and now The Physics of Wall Street: A Brief History of Predicting the Unpredictable by James Owen Weatherall. The shine isn't quite off Dr. Weatherall's first PhD yet—it is but half a decade—but already he exhibits a breadth of vision that sets him apart. He actually has two doctorates, in physics and in philosophy, so he has the kind of mind I like, not just thinking outside the box, but leaving all the boxes behind.

So why would he be interested in market analysis? For the same reasons a ton of physicists have had already: that is where the money is. Plus it has the un-ignorable allure of a challenge that is almost impossible, yet not quite. Given that many thousands of smart people have been trying to "beat the market" for, oh, half a millennium at least, a few have gotten rich, at least by chance, but rare indeed are those persons or funds who managed to stay ahead of the pack and get rich by actually betting on predictions that panned out, again and again. A physicist-run hedge fund called Renaissance is claimed to be one of them.

The bulk of the book is a history of statistical thought, as it developed over the past few hundred years, frequently in response to the desire to understand price fluctuations in markets for currency, commodities, or stocks and options of various kinds. The tools used for this began with the Normal (AKA Gaussian) distribution, the familiar Bell Curve. All kinds of additive phenomena obey Gaussian statistics, such as average height for men or women of a given ethnicity, or most famously, IQ. A particular Normally distributed population is completely described by a Mean (µ) and a Standard Deviation (σ). The shape is scalable, wider for large σ and narrow for small σ, but is otherwise fixed, so that 68% of the population is found within the range µ-σ to µ+σ, called the 1-sigma range; and the 2-sigma range encompasses 95% of the population. So, for IQ, at least among Euro-Americans, µ is standardized at 100 and σ at 15. Thus the range [70-130] includes 95% of these folks, and 68% are found in [85-115]. Public education was originally aimed at the 1-sigma group, and the rest were left to fend for themselves, until Special Education and Gifted Education movements arose to help out those in the "tails", whether duller or brighter.

Is the Normal distribution a good model of market fluctuations? Not at all. First, we must realize that human perception is involved. A $1 change in a $10 stock feels just as large as a $5 change in a $50 stock, particularly if you have 500 shares of the first one or 100 shares of the other. Both changes are 10% of your $5,000 investment. The chart below shows the day-to-day change of closing price for Coca-Cola common stock, since the beginning of 1986, expressed as a % of the prior day's closing price.


If we sort these numbers and plot them against a "Probability Ordinate", really an inverse Normal ordinate (I use the NORM.S.INV function in Excel 2010), we would get a scatter plot that closely follows a straight line if the distribution were Normal. But here is what we get instead:


If we extend a line tangent to the central part of the distribution, to -4σ or +4σ, it strikes at a 5% change, indicating that variations greater than this ought to be rare indeed (there are 7,101 daily changes plotted here). But what do we see instead? Going back to the original data sheet, I find 42 days on which the stock increased by 5% or more, up to nearly +20%, and 33 days on which it fell 5% or more, to nearly -25%. How'd you like to own a million shares of this stock and have it lose 1/4 of its value on a single day? So early on, the Normal distribution was found wanting.

Normal analysis was based on the concept of a random walk, also called the drunkard's walk. Its additive nature will always result in a distribution of final locations, say after ten staggers, that is Normal. So a different distribution with extra-wide excursions is needed. In an entertaining section, Dr. Weatherall describes a drunken firing squad. They have a target upon a very long wall, but being too drunk to point well, might shoot in any direction at all. Give them lots of ammunition (and hide somewhere until they run out), and the pattern of bullet holes will follow a Cauchy distribution. It looks a little like the Normal distribution, but has a pointier top, and most importantly, "fat tails"; that is, many points that are farther—or much farther—from the middle than a Normal distribution would predict. The distribution above is also fat-tailed, having lots of numbers outside the range we'd expect from a Normal distribution. To test a distribution for Cauchy behavior, plot it against a Tangent function evenly distributed in the range -π/2 to +π/2. For my 7,101 points, the Tangent function ranges from nearly -5,000 to +5,000, so the chart is thus:


The Cauchy distribution is clearly a bit too much, its tails are "too fat", compared to the tails of Coca-Cola daily price fluctuations. This kind of conundrum was tackled by many bright people, from Fischer Black to Benoit Mandelbrot. Mandelbrot probably came closest with fractal analysis, which wasn't wedded to integer exponents. But I got another thought as I read along.

The Normal and Cauchy distributions are related, being examples of Stable distributions. In one formulation, a parameter called α has a value of 2.0 for a Normal distribution, and a value of 1.0 for a Cauchy distribution. The fattest tails possible are at α=0, the Uniform distribution of infinite width. Mandelbrot had used fractal analysis to calculate a distribution with α of 1.7, closer to Normal but still with a fat tail. I realized that the most familiar distribution with at least one long tail is Lognormal, but it is confined to positive only values. Does it have a complex square root, perhaps? I sorted the squares of the KO daily changes and charted them against a Normal ordinate on a logarithmic scale. But there was a problem. On 245 days there was no change in price. Prior to the 1970s stocks were valued in 8ths of a dollar (12.5¢), and in pennies thereafter, though dividend allocations can be calculated to 0.0001¢ increments. Trades are reported to the nearest cent. Anyway, you can't take the logarithm of zero, so in my spreadsheet I used a value a little smaller than the smallest calculated nonzero value for those 245. They form the line at bottom left on this chart:


If trades were made with a continuous range of values, not limited by the minimum value, I would expect the left portion of the chart to be as linear as the rightmost. Quantization errors have artificially depressed the daily motion for about 15% of the trades. In the other charts, the two extreme values, -25% and +20%, seemed like outliers. Here, as the two rightmost data points, they are seen to be at most slightly larger than one might expect.

So, all you quants out there, working out the best formula for calculating risk. Give a little attention to the square root of the Lognormal distribution! Now, back to the book.

A key theme of the book is both the value and the danger of numerical models. A physicist understands that a model is always simplified, and cannot be appropriately used outside its range of application. When the people using models of financial systems, to set option prices and other instruments, are physicists, they will know this and avoid over-extending the model. People without physics education will not. When you have a black box program that seems to work magic, it is easy to use it everywhere (the parable of the man whose only tool was a hammer comes to mind).

There have been several major crashes in the past century, and only the one in 1929 was free of the influence of sophisticated statistical modeling tools. I say "sophisticated" because there were statistical tools in use a century earlier, but they were back-of-the envelope estimates at best. All of the more recent ones show at least traces of "broken model" influence, but the October 1987 crash was an overt "robo-trading" crash. This brings up another principle that physicists, at least, ought to keep in mind: the observer effect.

I am not just talking about Heisenberg Uncertainty. Rather, most observations of physical phenomena disturb the system being measured. I remember my father telling me not to check the air pressure in my bike tires so often, because each measurement caused some air to be lost. Later, working in electronics (in a time when the components were visible and manipulable by hand) I learned how to use a Wheatstone Bridge to measure DC voltage the most accurately, because it uses a counter-voltage to keep from bleeding extra current from the circuit. It is only good for very steady DC, of course. Thermometers change the temperature of the pot roast, but only a tiny bit; still the effect is not zero. But now imagine that you have half a million people whose livelihood depends on knowing the temperature in your pot roast, and they all insist on using their own thermometer. There won't be much left of the roast! THAT's what happened in October 1987.

The use of new tools changes the way markets work. What worked in September 1987 doesn't work today; what worked in 1997 or 2007 doesn't work now, and so forth. This pretty much negates the notion of an efficient market. It can only be efficient under two conditions:
  1. The traders have no supercomputers available.
  2. All traders are coldly rational.
Fat chance, right?

The "efficient market" works like this: In comparatively quiet times, the asking price of a stock or whatever incorporates all the current knowledge about things that might affect its value in the future. To profit from trading that instrument, you either guess it might be underpriced, because of unknown or little-known information, or you try to learn something nobody else knows. The most common source of such knowledge is cadging or coercing it out of an insider, which happens a lot even though it is illegal. Quantitative analysis attempts to find patterns in price fluctuations that signal a change you can profit from. When someone finds a useful pattern, he or his company will profit from it for a while, until others catch on, then pretty soon everyone can do it, and the market is "efficient" again. So quants' work is a continual arms race. Thus, the tools used to test the market change the market.

But the markets are not that efficient. In the medium term they might be, but the momentary trading picture is much more emotional, and tiny bits of information or rumor disguised as information can sway a trader's estimate of value. If that trader is influential, and others see him (usually male) make a move they didn't contemplate before, some will follow. It can cascade into a large market move, that might last a matter of an hour or less, but might last a day or more, and then there is the potential for quite a swing, either towards a bubble or a crash.

The fragility of any market lies in the tendency for all the quantitative trading firms to use the same models, or models based on the same math, with the same or very similar trigger points. Certain rules instituted after 1987 can calm the flurry to some extent, but the events of 2007 to early 2009 present a case in which the agony was simply drawn out over the space of more than a year, rather than taking place in a month or less.


With the contents of this book under my belt, I ask myself, "What is the ordinary investor to do?" We don't have supercomputers and armies of physics PhD's running sophisticated options evaluation software, trading 10-a-minute on our behalf. Dr. Weatherall doesn't tell us what to do. It isn't his business to do so. He is instead advocating for a kind of financial Manhattan Project to set an appropriate, physics-based replacement for the Consumer Price Index, whose flaws are politically grounded, very much on purpose (Oh, you thought it was objective?). As I said, his PhD's are still shiny and new. His next PhD needs to be in human nature, particularly the nature of the political human.

In the meantime, if you dare to invest in stocks, the advice of Will Rogers is still the best:
  1. Buy a stock.
  2. When it goes up, sell it.
  3. If it isn't going to go up, don't buy it.

Sunday, October 06, 2013

Looking too hard, and not looking

kw: book reviews, nonfiction, forecasting, prediction, statistics

We are remarkably good at cutting through the clutter in many situations. For example, we can talk to someone at a crowded party and pick out what they are saying in spite of the noise all around; and we can often spot a familiar face in a crowd. However, we sometimes see (or hear, etc.) things that are not there. When I was a child we would look for faces or other shapes in clouds. In a few minutes of looking, something suggestive is bound to appear. And there is a painting by my father of waves breaking on a rocky seashore. One of the big rocks looks like a leopard's head, and once I'd seen it, ever since I always see that leopard's head whenever I glance at the painting.

My father had no intention to hide faces in his paintings. Seeing the leopard's head is an example of a Type 1 error. If my father did actually hide faces in all his paintings, and I have noticed only this one (I have several others), then missing the faces that are there would be Type 2 errors. If I become so rapt in searching clouds for faces that I don't notice a friend approaching until he taps me on the shoulder, I have fallen victim to both kinds of error! We lazy, sedentary Westerners tend to do this frequently. Not so someone living hand-to-mouth in the woods.

For nearly everyone, through all the one or two million years of our evolution as brainy apes, hyper-alertness was required. Where it matters most, a Type 1 error does no harm, but a Type 2 error might be fatal. Running from a rock that looks like a leopard can make you look silly, but not running from a leopard that looks like a rock will probably get you eaten. Strangely, though we have kept our strong propensity to make Type 1 errors, as the risk of not noticing a real leopard has fallen, we are more and more likely to make Type 2 errors. In our modern world, in which we increasingly rely on forecasts and predictions, this leads to trouble.

Nate Silver, in his new book The Signal and the Noise: Why So Many Predictions Fail – But Some Don't, presents a number of similar examples that display our modern tendency to pick faces out of clouds while ignoring the approaching friend (or foe). I'll simplify matters and mention that he finds successful forecasting in only two areas: weather and baseball. Politics and stock picking and a number of other areas come in for a drubbing.

This simple diagram tells me all I need to know about "technical analysis" of stock prices. The data are the day-to-day percent change in the price of DuPont stock, from 1962 to mid September of this year. That's just over 13,000 data points. The X axis is the change on any particular day, and the Y axis is the change on the following day. This diagram shows perfect non-correlation! It is a 2-D bell curve, though with thicker tails than a Gaussian bell curve.

During those 51 years, the stock rose nearly 4,200%. That averages out to 7.7% per year but only 0.032% daily. Someone who bought $1,000 of DD stock in early January 1962 would have $43,000 today. Now, there's been a lot of inflation. That $1,000 in 1962 had the buying power of $7,740 today. So a half-century of waiting produced an effective multiplier of 5.5. That's 3% yearly after adjusting for inflation. Better than the bank.

The most extreme daily jumps are -20% and +10%. Stock speculators, particularly day traders, dream of taking advantage of the many days that a stock's price changes more than a percent or two. And such days are more common than if the distribution were strictly Gaussian. DuPont stock moves up at least 2.5% in a day about 5% of the time, and downward with similar frequency. That means, if you could pick just those up days, about 12 days each year, you could earn at least a 20% return yearly. That's 2-3 times what a buy-and-hold strategy will earn. Then, look at this:


The chart shows the historical record of DuPont stock, adjusted for splits. Focus on late 1974, late 1987, and late 2008 to early 2009. These show DD following the herd during market crashes, and represent downturns of 50%, 41% and 65%, respectively. If you could have avoided them, by selling just at the peak and buying back in at the bottom, your final return would be 9.69 times greater, for a total value of $416,000! Adjusted for inflation, that's over 8% return yearly (12.5% dollar-for-dollar yearly return).

Such figures stoke the dreams of day traders. But the first chart, showing no day-to-day correlation, dashes those dreams. Day traders work very hard for little return, and most lose. Some lose, big time, and some gain, but it is by accident either way. There are millions of day traders and other stock speculators. As Churchill wrote, "Even a fool is right once in a while."

Now we must differentiate prediction from forecasting. A prediction is a flat statement that a specific happening will or will not occur at some time or in some time horizon. For example, "There will be a magnitude 7 earthquake in Fremont within the coming year." A proper forecast includes the forecaster's uncertainty and is stated in probabilistic terms, as, "Projecting the trend of earthquakes in Fremont indicates that an earthquake of magnitude 7 or greater occurs about 3 times every 200 years." [Fremont was the imaginary State in the novel Space by James A. Michener]. One might add to such a forecast, a hybrid statement such as, "Fremont has not experienced an earthquake of magnitude greater than 6 in the past 100 years," which implies that "the big one" may be overdue. But it may indicate that conditions deep down may also be changing.

Earthquake prediction is the poster child of unpredictable phenomena. Intense study and research over decades, even centuries, have failed to yield a single valid prediction. Sports betting is close behind, except in the arena of baseball. Nate Silver once created a system he calls PECOTA, that rates the strength of teams against one another according to the past statistics of their players, and a well-known "aging curve" of the way performance changes over a player's career. Because baseball has such a rich data set, going back a century, and the principles needed to make useful forecasts are also well known, PECOTA and similar systems can evaluate players and teams at a level nearly equal to the best scouts. The computer can't quite replicate the humans, but it does give 'em a run for the money!

Why are forecasting and prediction so hard? Even though we have randomness at the deepest level of atomic phenomena, that randomness is constrained by the statistics of large numbers, and physics works very accurately to predict many systems, such as planetary orbits. Thus, though the path of an electron after passing through a hole may be uncertain, the distribution center of the paths of trillions of electrons (say, a millionth of an ampere for 0.1 second or so) will be very sharply defined and can be accurately measured, and the shape of the distribution tells you additional facts: the hole's size and shape. The much larger "distribution" consisting of the atoms making up a baseball mean that its flight, once thrown or batted, will be easily predicted.

The geological setting of an earthquake is not as simple as an electron. Perhaps this year, an earthquake might occur, large enough that the two sides of a fault will slip by each other by half a meter. That may be enough to put two kinds of rock in contact, that were not in contact before, which changes the likelihood of the next earthquake.

What about the weather? Air is in constant motion; its humidity and temperature, and thus its density, change constantly. How can anyone make a useful weather forecast? In some ways, we are still dependent on the "signs in the sky" that Jesus mentioned. In modern (18th Century) terms, "Red sky at morning, sailor take warning. Red sky at night, sailor's delight." Lore such as this is a compilation of patterns that happen over and over, so that generations of our ancestors took note and remembered. Yet now we can get a forecast up to a week or two ahead, complete with expected high and low, precipitation chances and intensity, and wind strength.

It's all done in a computer. Air may have complex behavior, but the physics of air motion and how it changes with temperature, pressure and humidity are well known. The 3D-gridded-cell models that run in supercomputers use surprisingly simple physics to determine how a 3D cell is influenced by the 6 cells it is in facial contact with, and the 8 cells at its corners. The reason supercomputers are used is that Earth is big. The surface area of the planet is 4πr², where r is 6,370 km: about 510 million km². Cells of half a km on a side, plus 0.1 km in depth (up to 12 km altitude) result in a Global Circulation Model (you'll see the acronym GCM in some weather web sites) with 1/4 trillion cells. It takes a lot of calculation to determine what will happen in the next quarter hour. There are 96 quarter hours in a day, and 672 in a week. To do all those trillions and quadrillions of calculations in only an hour or two requires today's largest computers. And the forecasters' computer gurus don't do it once, they run it several times with very small variations (the formal practice of selecting the variations is called Design of Experiments), to test the stability and sensitivity of the forecast to perturbations.

Weather forecasters have an incentive to get it right that others don't have. The reality is going to arrive tomorrow or the next day, it is visible to all, and it is no fun getting a call such as, "I have ten inches of 'partly cloudy' that I need to shovel off my driveway. Want to come over and help?" They also get a ton of research money from the Dept. of Defense, because good forecasts are crucial to military activities. Earth dynamic studies are different. Students of earthquakes can't observe the day-to-day conditions of a fault line. Its active zone is typically 8-15 km deep, and we can't yet drill a well that deep. Earthquakes are also rare. Sure, there are thousands of little ones, at the bottom of "measurable", every day, but there are trillions of weather events around the globe, every few minutes.

Mr. Silver entertains us with many, many stories of the vagaries of forecasts of all types. In the end, most phenomena are too difficult to forecast appropriately. Some involve living things. The cardinal rule of animal studies is, "Given any particular set of temperature, lighting, food availability and ambient noise, the rat will do whatever the rat wants to do." And this is in spite of lab rats being so inbred that their genetics are practically identical. The statistics of playing poker yield a few big winners, who work hard for the kind of edge they need to beat their fellow experts. But they love to be in a game that is well supplied with "fish": overconfident amateurs. A well-written computer package might tell a poker player the optimum betting strategy, but only if it is betting against other computers. The social aspects of the game, bluffing and speed or slowness of a bet for example, often provide a lot more of an edge than the math does. Carefully crafted intimidation works wonders. I don't expect a computer to master these aspects of the game for a number of decades (that's my forecast!).

The book's final example is the climate, particularly "global warming" or "climate change" or "greenhouse effect" or whatever the next buzzword will be. Climate is not weather. It is the setting in which weather happens. Climate changes unfold over multiple decades or centuries or millennia. Weather changes take seconds. In numerical analysis, this is the Stiffness problem. When something changes suddenly, it takes time for the effects to either move elsewhere or to die down. If you are interested in something with a 5-year cycle, such as El Niño (also called ENSO), the exact location and timing of today's sudden thundershower will not matter one tiny bit. If your interest is in human-induced greenhouse warming that began in the late 1700s, ENSO is an irritation at best. In fact, weather and medium-scale cycles such as ENSO are "noise" in the context of this book's thesis. Another researcher, later on, made clear a different view, that noise is really signals, but about stuff you aren't interested in at the moment.

This is like the crystal radio I made as a kid. It initially consisted of a long wire, running to a treetop, a piece of germanium crystal, and a "whisker", a wire that formed a diode with the germanium; and earphones attached to the whisker and the ground connection on the back of the germanium crystal. The diode "detected" the audio signal by separating it out of the radio frequency "hash". There was just one strong station nearby, so I could hear them pretty clearly. But later, as more stations came on the air (this was the 1950s), I could hear all of them at once. So, following a diagram in Mechanix Illustrated, I made a coil and paid a dime for a small capacitor and a piece of copper, to make a rough tuner. It could be tuned to resonate with one AM station at a time, so I could "tune out" the "noise" of the other stations. They were actually signals, just signals I didn't want right then.

The global greenhouse has warmed about 0.5°C (0.9°F) in a century, and perhaps 1°C (1.8°F) since 1750. Some of that may be warming since the Little Ice Age, which some consider a regional phenomenon, not a global one. But the current "ForecastFox for Mozilla" forecast for the next 24 hours indicates we'll have a 20°F swing tomorrow, from 75 in midafternoon to 55 overnight. You have to average out a lot of daily temperatures to see a change of a degree over 250 years. When you want weather, that is your signal. When you want climate, weather is noise, and lots of it.

The science of greenhouse warming is partly very well known, and partly not so well known. I learned to replicate the Arrhenius calculations from 150 years ago, when I was a pre-teen. Actual warming since his day has been about twice what he expected, because there seem to be amplifying factors. These are very poorly known. Does more cloud cover cool the atmosphere by reflecting more sunlight, or warm it by acting as a further thermal blanket? Or does it do one thing at a certain latitude and another elsewhere? If we do have a further warming by 2 to 4°C, will it shift the Hadley Cell north, or south, or not at all? (The northern edge of the Hadley Cell is a range of latitudes characterized by dry, descending air that form all the world's great deserts.) I've thought of buying land in central Canada, that is currently too cold to farm. Perhaps in 20 years it will be arable…unless the Hadley Cell shifts north and dries out Canada. Then maybe the Mojave would become a tropical paradise!

Y'know how to make a complex system into a positively unsolvable mess? Make it political. Both sides of the Climate debate are so politicized that they can only talk past each other. The tiniest proposal to set any policy is vigorously fought by every vested interest, even those who might benefit (the devil you know…). Heaven help us if weather forecasting ever gets politicized! It is already true that most forecasters err on the wet side: a 20% chance of rain is reported as a 40% or even 50% chance, because the ones rained on are less likely to complain, and those that aren't will feel they dodged a bullet. What if some "weather outcomes" become more politically correct than others?

By the way, I take issue with Silver's definition of statistical rain forecasts. He writes that if 40% of the computer models indicate rain in Chicago, and the rest don't, it is reported as a 40% chance of rain. Sounds logical, but it is quite different than that. The "chance of rain" has different meanings in spring (plus summer) and autumn (plus winter). Spring and summer squall lines pass through areas that are well predicted by most GCM programs. But a squall line is not a solid front of rain. It is a line of thunderstorms. A light squall line may have storms half a mile wide, spaced 2-3 miles apart, giving 20% of the area a 100% chance of rain. The forecasters just don't know which 20%, so the whole area is given a 20% chance of rain. A heavy squall line will have larger storms with closer spacing, and maxes out at about 80% coverage (though this will probably be reported as "near certain"). Fall and early winter storms tend to be solid and widespread, but subject to ripples several miles wide in the upper atmosphere. As a system rides up a ripple, it drops rain along a solid band dozens of hundreds of miles long but only about a mile wide or so. As it rides down, it dries out. The height of the ripples determines whether the overall chance of rain is 30% or 70% or somewhere between. The ripples drift along as system after system rides through, so it is very hard to tell exactly where the rain will fall. Timing is everything. Then, a lower-level storm that just dumps (ignoring the ripples) leads to those 100% forecasts, which are generally accurate.

In most arenas, Silver advocates using Bayesian analysis rather than "frequentist" simulations or estimations. These allow individualized forecasts for particular cases. An example is the probability of breast cancer in a woman in her 40s, who has just had the unwelcome news that a mammogram is "positive". The factors of a Bayesian calculation are:
  • x - Prior Estimate: the chance that a proposition is true.
  • y - Type 1 analysis: the chance that new data which indicates "Yes" is actually correct.
  • z - Type 2 analysis: the chance that the proposition is not true, in spite of the new data.
The data are usually noted as percents. The formula for a new estimate (a new x) is xy/(xy+z(1-x)). For this example, we find:


In this case, the woman may wish for a needle biopsy, but a bit of blood chemistry may be in order first. Enzymes in the blood can indicate whether a new cancer is likely to be slow growing, or faster. Is it slower (the most likely case)? She can wait a year for another mammogram. If the next mammogram is positive, re-do the analysis, replacing the 1.4% with 9.6%. Now the "new x" is just over 44%, and at the very least a biopsy is indicated. Most other forecasting methods don't use multi-step refinement. And by the way, if the next mammogram is negative (and no palpation can detect a lump, or any growth in an earlier lump), running the analysis with 9.6%, 10% and 75%, in that order, reverts to 1.4% as the "new x".

Those who follow this blog may wonder why it took me 3 weeks to read such a fascinating book. The writing is good and the examples are interesting, so that didn't slow me down. We have a lot going on, however, so I have had much less time for reading than usual. Retirement has been good to me so far, but I have to be careful not to take on too many projects at once. I completed a Real Estate course and passed the test in July. However, I will probably not seek a license or become a Realtor®, because there are simply too many other things I'd prefer to do. The change of style and reduced frequency with which I post is a similar effect. I used to post almost every lunch hour, doing research in off hours. I think I am working longer days than when I worked! Better busy than bored. Since retiring in February, I have put 24 items in my "job jar" file. Half of them, mostly the bigger ones, have been completed. One major item is awaiting an event that is at least a year in the future, but the preparations are nearly all completed. Others are smaller so I can take an odd half day to perform one. All things in their own time. In the meantime, I read when I can, and report what I read.

Tuesday, July 23, 2013

You don't have to hate statistics

kw: book reviews, nonfiction, mathematics, statistics, popular treatments

Measure something. Say, take a yardstick and measure the width of the kitchen counter. In my kitchen, I get 24 inches. That is an observation. Guess what? You can't do statistics using one observation. Not because you are somehow incompetent, but because of the way statistics is defined. A common definition is:
Statistics is the practice or science of collecting and analyzing numerical data in large quantities.
Note the final qualifier: "in large quantities". It is possible to do a certain amount of statistical inference using just a few items—and we'll do some momentarily—but you typically need lots of data to produce a robust inference. However, a few principles can be seen by analyzing just a few observations. I measured my counter in five more locations. Here are all my observations:

24
24 1/8
23 7/8
23 7/8
23 3/4 (= 23 6/8)
23 5/8

We can do a few things with these six numbers. First, comparing the largest with the smallest, we see that the range is 3/8 (just under 1cm). I can take the average, which comes to 23 7/8. Hmm; if the building plans specified a 24 inch counter top, this one averages an eighth inch too narrow. Then there is a trend. These are in order, from one end of the counter to the other. The largest measurement is the second one, the smallest is the last, and the rest of the measurements follow a decreasing trend. In angular terms, a "tilt" of 3/8" in about 10 feet is only a sixth of a degree, but I'd expect a builder to do better than have "nearly a half inch" of variation over ten feet. Oh, well. One of my projects for later this year is to replace the counter tops anyway. I hope quality control has improved since these were installed in the 1970s!

Now for just a little terminology. The "average" I figured is known at the "mean". It is not the only way to determine "central tendency". Another is the "median", which means, the one in the middle (or the average of the central two if the sample has an even number of observations).  For example, if I sort these six numbers (in this case, just move the 24 1/8 above the 24), it happens that there are three that are 23 7/8 or larger, and three that are 23 7/8 or smaller. So the median is 23 7/8. This is not always the case, and perhaps it is not even usually the case. For example, if I have the seven numbers 1, 2, 3, 5, 8, 14, 30, the mean is 9 but the median is 5. Note that only 2 of these numbers are greater than 9.

Another such measure is the "mode", which means the most likely value. Mode is really not too meaningful when there are only six observations, but for these data, the mode is also 23 7/8. Suppose instead that I had measured that fourth width as 23 3/4. This would have very little difference on the mean (23 6.8/8) or the median (23 13/16 or 23 6.5/8), but the mode would now be 23 3/4, because that number arose the most frequently (twice).

This illustration shows how these are related (Image from The Daily Dongle). A frequency plot of a very regular set of measurements such as shown in (a) will have mean, median and mode that are equal or nearly equal. Sometimes we make measurements that have more than one "hump" (their distribution is called bimodal) as in (b). But (c) and (d) show two ways that a series of measurements may reveal a skewness, in which case the three measures will be quite different.

Each has its uses. Average height of Euro-American males is best described as the mean, the numerical average of all measurements. We might also surmise that the median and mode will be very similar to the mean. But if you include Euro-American women, the bimodality may not be too evident, but it is there. At the very least a frequency plot will be flatter on top and have a wider total range. If the average male is 70" tall and the average woman is 64" tall (for Euro-A's, anyway), the grand average will be 67", but that single number tells you less than the two numbers, segregated by sex.

What about yearly income, or prices of homes in a city or county, or the whole country? When you hear a Real Estate report on the radio, you will hear, for example, "Median home price has risen by $5,000 in the past month". Why not use the mean? Because the distribution is skewed. There might be a few homes with very small values, and a few with very high values, but where do you put the "middle"?

Example: Broken Arrow, OK (I know someone there). The least expensive houses on the market, as I find from Realtor.com, are in the $25,000-$50,000 range. The most expensive, in the range between $1.2 million and $1.4 million. Do you think it likely that home prices are evenly distributed between these limits, producing a "middle" value of about $700,000? Not likely! In this market, this moment, 684 homes are for sale. Houses # 341 and 342 on the sorted list the web site provides are both priced at $170,000. That is our median for this market (today). Quite a bit different from 700k, isn't it? Half the houses' owners are asking $170,000 or less, and the other half are asking more. If you can afford a $200,000 house, at the most, you have a lot to choose from. Wherever the larger values in a distribution are a big multiple of the smaller values, the median is usually the best measure of "average".

This is my simple attempt to explain a few statistical principles. Charles Wheelan does a superb job of explaining these and a goodly number of others in Naked Statistics: Stripping the Dread From the Data. In the middle of the book, for example, he dwells quite a bit on the Central Limit Theorem. This has to do with sampling.

Above, I took six measurements of my kitchen counter. I could have taken a lot more, perhaps spaced every inch, or even closer. Suppose I sent my wife into the kitchen with a yardstick and asked her to make six measurements, with the same yardstick, in locations of her choosing. Then perhaps we could grab some of our neighbors and have them repeat the experiment. Now I will have several sets of numbers, and each set will have its own average. Do you think any of the averages will be close to, say 22, or 27? Not unless there are some BIG wiggles in the counter's shape, that I avoided with my measurements. If I could get a lot of my neighbors to make sets of measurements, the Central Limit Theorem (CLT) predicts that they will be distributed a lot like section (a) of the illustration above, clustering about some average value that is close to the "real" mean for all possible measurements of my counter.

As the author goes on to show, with marvelous examples, this is the source of the power of polling. Not only can a poll yield very useful results about all 180 million American adults by polling 1,000 or 2,000 people (properly chosen!), marketers (who pay the most for such data) can predict some of our preferences based on what we have already bought or even searched for (Google sells its search results, don'tcha know). My wife and I have "loyalty cards" from a few local grocers and other stores. We get discounts on certain items for scanning the card when checking out. In a sense, the store is paying us for the right to keep track of our purchases. Something else we get during checkout is a series of spot-printed coupons (the more we buy the more coupons they print). Some coupons are for more of the things we often buy. Others are for similar items of competing brands (the brands' owners are in on this also). And there will usually be a few "wild card" coupons that show up over time, for things we might not usually buy. Why? Because other people whose purchasing habits are similar to ours buy those things, and the store is betting that we are more likely to try those items if we get a coupon to prod us, compared to giving the same coupon to random shoppers. They have also figured out that we are "on the edge of elderly", so some of the coupons are for things like Ensure (an energy drink for old folks) or Depends (adult diapers).

Think about it. A typical supermarket has tens of thousands of customers that visit regularly. If 25% have the store card, they can slice and dice that population a dozen or a hundred ways, to target their coupon campaign. And, since coupons cost almost nothing to print, they can throw in 30%-50% off-the-wall coupons so we don't realize how precisely we have been targeted!

If you get nothing else out of this book, read it carefully for the author's explanation of the CLT and his stories of how it is used (such as how Target "helped" a father learn that his teen daughter was pregnant). He reveals all sorts of tricks of the trade, such as the numerical way to handle binary differences such as male/female. I am a math junkie, so of course I love a book like this. But I think the math-averse will also find it very entertaining and informative.

Friday, July 06, 2012

Rechecking autism prevalence

kw: autism, statistics

This morning I heard an ad I have heard many times, in which Curt Schilling ends by saying, "The chances of having a son with autism: one in 110." I decided to look into it. My sources for the statistical information below are this publication by C.J. Newschaffer et al and this Wikipedia article.


The public perception of autism was greatly improved by Dustin Hoffman's portrayal of a high-functioning autistic man in The Rain Man. The character of Raymond was based on Kim Peek, who died in 2009. However, not many autistic persons have genius-level abilities of memory or calculation. Kim Peek is an example of an autistic savant, and he was somewhere in the middle of the "autistic spectrum".


The autistic spectrum comprises three defined syndromes:
  • Classical Autism describes people who cannot form relationships, may never recognize other people, and communicate poorly or not at all.
  • PDD-NOS, Pervasive Developmental Disorder Not Otherwise Specified, is a wide-ranging slush bucket of syndromes that include some, but not all, characteristics of classical autism. Some cases are more profoundly affected than others, and many or most instances of autistic savants are found here.
  • Asperger Syndrome is the highest-functioning type. A few AS persons are savants, but the syndrome as a whole describes people who are very focused on one or a few interests, may become very expert in those interests, are able to converse and form relationships, but have certain autistic behaviors.
I know three men with Asperger syndrome. One in particular is a young man who went to high school with my son. He and I were assigned to collect tickets for a football game. He is extremely interested in weather, and when he learned I had lived in Oklahoma, he quizzed me extensively about my experiences of powerful thunderstorms and tornadoes. My son said he is the school weatherman, always glad to provide a forecast, though perhaps in greater detail than you have time for. Unlike people with classical autism, he could look you in the eye.


I also know two young men who are profoundly autistic. One is the son of a couple who belong to the local Mensa chapter. He is barely able to recognize his parents, cannot speak, and the slightest frustration causes him to roar like an animal. It seems his mother can guide his behavior a little, in that he understands a few of the words she says. His perception of English is probably below that of a house pet.


In spite of several decades of study, the actual incidence of these disorders is only poorly known:
  • Autism: 1 or 2 per 1,000. The average of a number of studies is near 1.7 per 1,000.
  • PDD-NOS: 3.7 per 1,000.
  • Asperger: About 0.6 per 1,000.
These add up to very nearly 6 per 1,000, or 1 out of 167. There is an added fact. Across the spectrum, males outnumber females 4.3 to 1. In my reading I didn't get a clear feeling whether one of the three areas hits females harder or less hard than males, in proportion, so I just analyzed these data based on the 4.3:1 figure for all. I find that the totals for the entire spectrum are 1/440 for girls and 1/103 for boys. The ad with Curt Schilling is correct in this regard, as it specifically says "son". The spectrum splits up this way:
  • Autism: 28%, or 1/1580 for girls and 1/370 for boys.
  • PDD-NOS: 62% or 1/710 for girls and 1/165 for boys.
  • Asperger: 10% or 1/4,400 for girls and 1/1,030 for boys.
If there is a silver lining in this dark cloud of affliction, it is that nearly 3/4 of children on the autism spectrum will be able to relate to others at a useful level. The tragedy is that of 4.1 million babies born yearly in the U.S., nearly 25,000 will be afflicted with an autistic spectrum disorder, and nearly 6,900 of them will be classically autistic and require lifelong care.