Showing posts with label prediction. Show all posts
Showing posts with label prediction. Show all posts

Monday, January 16, 2023

The irresistible impulse to predict

 kw: book reviews, nonfiction, speculation, prediction, futurism

The second chapter of The Skeptics' Guide to the Future: What Yesterday's Science and Science Fiction Tell Us About the World of Tomorrow, by Dr. Steven Novella and his brothers Bob Novella and Jay Novella, begins with a quote by the sage of the ball park, Yogi Berra: "The future ain't what it used to be."

I can think of three facts about the 1969 landing on the Moon by astronauts of Apollo 11, that were not anticipated by science fiction writers, nor by futurists in general:

  1. Between 10% and 30% of Americans believe it was all a hoax.
  2. The sponsor was not a corporation nor a "rich industrialist" but the Federal government.
  3. The world watched Neil Armstrong climb down the ladder on color TV.

It took only 50 years for a few "rich industrialists" to gather both the financial muscle and the personal will to start spaceship companies, and at present there are three! In the golden age of space fiction, mainly pre-1960, nobody dreamed the combined Mercury-Gemini-Apollo programs would cost about a trillion dollars (in current dollars).

What did science fiction get right? Arthur Clarke foretold telecommunications via synchronous satellites in 1945, 12 years before Sputnik 1 surprised the world. On the dystopian side, George Orwell wrote in 1948 of a pervasive surveillance state (he was mixing Stalin with his expectation of immense progress in video technology); now it is upon us. More prosaically, in 1911 Hugo Gernsback wrote of video chatting; of course the inimitable Hugo made dozens of predictions in everything he wrote, so getting it right once in a while is simple statistics in this case.

What approach do the Novella brothers take? I'd call Skeptics' Guide a catalog of futurist ideas sorted by plausibility. I am amused by the title of Chapter 3, "The Science of Futurism." To me, that's an oxymoron. Science is about what we know, or think we know. Stuff that hasn't happened cannot be a scientific subject. To me, futurist speculation is fun, and when folks in a position to know more than most of us about technological and sociological trends speculate that one or another trend "has legs" and is likely to affect all of us, it's worth taking note. Some decent fraction of the time, the speculation turns into fact.

The authors sagely note that it's hard to distinguish robust trends from fads. I look back at a number of things I once thought were "neaty-keen!", that have vanished from the scene, with decidedly mixed feelings. I tried out a Segway at a science park once. I like them. The business model was fatally flawed, though, and my insurance company would probably have charged me more than the cost of the machine...per year. So as much as I hope for the rapid development and spread of self-driving autos, I suspect I won't live to see them "take over" (but I'll get one of they do!).

Or maybe I'll live long enough for many more trends to play out. The next-to-last chapter deals with indefinite life extension. Scale that back to "let's make it possible for most people to live 100-120 years in good health", and I have a decent chance to use or own (if we still own cars in 2040) a self-driving car, and to have not a robot butler but instead several (affordable!) robotic appliances, such as one that can make a decent omelet. Maybe I'll live to see actual, true, artificial general intelligence (AGI). My goal post for AGI is that a mechanism, entirely without "handlers and minders", does research, produces something new, and obtains a patent. A secondary goal post: one that can carry on a discussion on numerous subjects, and hold its own at a meeting of an intelligent, well-read Book of the Month Club.

Of course, artificial narrow intelligence (ANI) is beginning to do some very interesting things. One is the language understanding behind "art generators" such as Dall-E and "text generators" such as ChatGPT, and the object recognition behind image sharpeners such as Sharpen AI. Would life in a settlement on Mars such as the one shown here appeal to you? I prompted Dall-E to create this image. It's a portion of a larger image I "created" with "outpainting". ANI is beginning to push our culture in new directions.

After the introductory part of the book, each of the four sections deals with matters that are less and less likely. The last section discusses the impossible, at least until new laws of physics are discovered: time travel, transporters, and faster-than-light travel for example. My, how prone we are to wrestle against the constraints of the universe!

The point is often made that fiction writers and futurist writers cannot escape putting today's people into their future scenarios. A small number of people are fanatical about "living forever", but do they have a way to predict the social disruption that would result, if even 0.1% of humans became effectively immortal? The authors discuss both positive and negative views of this, and are careful not to draw conclusions. It would be a Black Swan moment for the human race. While I'd like to "not die", I am not willing for evil people to live long enough to get really, really evil. Imagine if Stalin or Mao Zedong were still alive...

As good as the writing is, the book is rather insipid. The authors didn't want to be sensationalists. They succeeded at that.

Yogi Berra also said, "It's tough to make predictions, especially about the future."

Friday, June 04, 2021

Linear thinking in a non-linear world

 kw: book reviews, nonfiction, prediction, forecasting, disciplines

My colleague at work had a poster: "Life is uncertain. Eat dessert first." We hate uncertainty, which is why we want to find out what is going to happen. We hate lack of control, which is why we try to control the future. Unfortunately, the future is unknowable and beyond our control. We aren't even very good at self-control. How could we hope to control a world composed of billions of others who are just as uncontrollable as ourselves?

What do you see here? Suppose this represents the general state of someone's health, projected over a long lifetime. Where is its end? If this shows the life of someone who is "fated" to live 85 years, they had a near-death experience around age 35.

Sliding toward the bottom of that valley—was it an illness, an accident?—, it may seem like the end is near and that good health is irretrievable. But, hey, things get a little better almost immediately, and later, a lot better. But that's not where these data came from.

Suppose instead it represents some factor of the economy. Is this part of a bigger picture? What comes after, and what came before? Did you notice that the end of the diagram is close to the level of the beginning, though a little bit lower? There are a couple of big, sustained "UPs", and a couple of big, sustained "DOWNs". In the middle of one of those sustained runs, whether up or down if you had to bet on the future of the rest of the diagram, how would you bet? Would you have any idea what the value would be at any point on this chart?

As a matter of fact, this is a financial chart, and adding the axes shows more context:

This is the Google Finance chart of the weekly closing price for Apple, Inc (AAPL), from 1/1/2021 to date (6/3/2021). It looks pretty dramatic! But check the left axis: the low for the year, in early March, was 116, the high in late January was 143, and the June 3 closing was 123.5 (pretty close to the average for the period of 128.6). If you knew just those figures, you could make a "bet with sideboards" that, for the short term, the stock price would be within ±11%. A final chart will show a little more context:

I downloaded the YTD historical data and charted it in Excel. When a stock price chart includes a zero, it looks less dramatic. From this perspective, this stock has been "flat" for five months.

How long will it stay "flat"? I couldn't hazard a guess. Do this: pick a time frame, such as a year, or five years, or (being more canny) the Midterm Election in late 2022 or the Presidential Election in 2024, and call that "the end of AAPL flatness." Write it down (or put a memo in your phone's calendar). Check it on that date and see if AAPL is still flat, or if it entered a period of serious gyration.

Now think about possible "black swans", such as someone buying Apple, or a competing product blowing the iPhone out of the market, or more globally, a new war, a recession, or a suddenly booming economy and the Dow Jones goes to 100,000. I don't have any idea how possible any of these things are, but if the possibility of any of them is more than a few percent, it would call any prediction about this stock or any other into serious question.

Consider this. The moment-by-moment price of a stock is influenced partly by the market as a whole, made up of investors that represent only a couple % of the population, and partly by the emotions of investors in that particular stock, which number a few dozen to a few hundred people. Wouldn't you expect that predicting the outcome of larger questions might be even harder? A year or five years from now, what will be the health of the economy of the nation or of many nations; what is the likelihood of war between two rival nations; or what of the effects of offshoring or onshoring a major segment of the work force? All of these depend on the decisions of thousands to millions of people!

About two months ago I reviewed a book about forecasting and "superforecasters", the kind of people who are better at making predictions that are at least a little more accurate than guesses. The author of that book referred to another, Future Babble: Why Expert Predictions are Next to Worthless, and You Can Do Better, by Dan Gardner. I spent the past two weeks reading this book, with some care. I think it best to begin with a few statements that caught my eye:

On Behavioral Economics, "In the 1950s, Solomon Asch, Richard Crutchfield, and other psychologists conducted … experiments that revealed an unmistakable tendency to abandon our own judgments in the face of a group consensus—even when the consensus is blatantly wrong. …three-quarters of test subjects did this at least once." (p. 99 in Chapter 4)

On "predicting 'more of the same' ", "…[such] predictions are more likely to be right when current trends continue and least likely to be right when there is a drastic change. That's most unfortunate, because when the road is straight, anyone can see where it is going. It's the curves and corners that cause crashes. …predictions are most likely to be right when they are least needed and least likely to be right when they are essential." (p. 105 in Chapter 4)

On Chaos, or the unexpected influence of small differences, "Like most of what's interesting in life, weather is subject to chaos and all sorts of nonlinear weirdness that limits how far we can peer into the future. Those limits will never be eliminated." (p. 105 in Chapter 8)

The first item above refers to "groupthink", which led to the disastrous Bay of Pigs "invasion" of Cuba early in the Kennedy administration, and to the under-planning of a tragic rescue attempt in Iran during the Carter administration. Do note the successful rescue operation in 1979, of two EDS employees being held in Iran, an operation ordered and sponsored by H. Ross Perot: his decision, carried out under the command of one experienced "hired gun". No groupthink.

Allied to this is the need for certainty, which means the most confident (and loud) voice typically carries the day. But turn the last phrase of the around: one-fourth of the experimental subjects did not succumb to groupthink. Their voice may not have prevailed to "turn" the others (who were all confederates of the experimenters), but they did not give up their view. In the terminology used by Gardner, they were probably "foxes", as in the proverb, "The fox knows many things, but the hedgehog knows one important thing." As one might expect, the hedgehog's "important thing" may or may not be correct, but when it is wrong, the hedgehog will not change his mind. Foxes are more flexible. They are willing to consider whether their thinking is wrong. Using the same analogy, the "superforecasters" I wrote about in April are foxes.

Let us consider the idée fixe, or fixed idea: an idea or desire that occupies one's mind, often to the point of obsession. This is the hedgehog's specialty. It takes work to break free from an obsession, but successful forecasters are willing to do the work, and the extra work to gather knowledge from numerous sources (the fox's "many things"). However, foxes have a hard time getting through to the public, which prefers the certainty of hedgehogs, no matter what their track record (which is uniformly abysmal except for the occasional lucky guess).

I recall that the root word for "fool" in the book of Proverbs is "self confident". It is used about 70 times. Not only are hedgehogs such fools, so are all those who listen to them uncritically.

For the second quote, I'd call it self-evident, but we all have "hindsight bias", which is its basis. There is a road I sometimes take to work. It goes over a rather steep hill. Just at the top, the road jogs just a little to the right, which means if you aren't paying close attention, you will suddenly be halfway into the other lane. A great place for head-on collisions. It would be better if the roadway on either side of the hill were curvy, forcing drivers to pay better attention. Even better if the highway department tore out fifty feet on either side of the crest and smoothed that transition.

Life doesn't usually provide smooth transitions. Look again at the stock charts above. Any two points more than a tenth of an inch apart would seem to have nearly no relationship to each other. In a few cases, a tenth of an inch is enough for a rather dramatic swing. The basic rule of thumb we need to guide us is, "The more confident an expert sounds, the less we should trust the prediction." Let me repeat that:

The more confident an expert sounds,
the less we should trust the prediction.

As fun as it would be to belabor the third quote, about Chaos, I'll forbear. I realize that only the "choir" would take such a word, and everyone else would ignore it. Instead let's realize that Chaos is a manifestation of Nonlinearity. We have linear minds. From many years as a working mathematician, I know how hard it is to think nonlinearly. It isn't natural to us. Even when we know something is cyclical, like the march of the seasons each year, the day-to-day variations of sunlight, clouds, rain and high or low temperature are tough to predict. The best weather forecasting computers do OK for a day or two. After that all bets are off.

Mark Twain wrote about a period of around 100 years in which the Mississippi River's length was reduced by dozens of miles, because a few long, loopy bends in the riverbed were cut off when the river flooded and cut shortcuts. He said that if the trend continued, in another couple of centuries St. Louis would be practically a coastal town on the Gulf of Mexico, and that millions of years ago the river must have been sticking out over the Gulf "like a fishing rod." 

We know that is illogical, and we laugh at it. But we neglect to consider that seemingly regular trends don't persist. Any measure we might look at is more like the stock chart above, with ups and downs that nobody could predict. Over the entire five months charted, the stock price went from 129.4 to 123.5. That's about a 5.5% reduction. But there were a couple of spots in there where a day trader could have gained 10% or more in a week's time, and a couple of others where a day trader would have had to ride out a 10% loss.

I apologize if I keep returning to the stock market as an example. It is an area I'm familiar with. It is seldom referred to in Future Babble. The author's interest is in the way we react to predictions and the way we almost always forget that they very rarely pan out. Most pundits are so wrong so frequently, it is amazing that anybody buys their books. So I'll end with a prediction: 

The pundits and "experts" who are the best at sounding confident will get rich on the gullibility of a public that craves certainty where there is none to be had.

Friday, August 07, 2015

Our life in bits and bytes

kw: book reviews, nonfiction, algorithms, prediction, sociology

What would life be like if the atoms that make us up were just big enough to see, if we could witness directly how they slide, merge and separate? How complex could our life be if the sum total of our lives could be described by, say, 1,000 characteristics, or perhaps 100? How about 10?

Yet how quick we are to pigeonhole people according to one or two, or at most five, distinguishing items! What do most of us now about, for example, Yo Yo Ma? Male, Chinese, famous musician (maybe you know he is a cellist), … anything else? How about that he is French born, a Harvard graduate, and has earned 19 Grammys? That's six items, more than most people probably now about him.

To what extent do you think you could predict his tastes and buying habits from these six items? If another person shares these six characteristics, to what extent will he also share certain tastes in clothing or food or books to read? Some people wish us to think, "to a great extent". In The Formula: How Algorithms Solve All Our Problems and Create More by Luke Dormehl, some of the people he interviewed claim to do just that. (Maybe you've made a profile on a dating site that starts matching you up when you've entered no more than four or five items. And how fully have you completed your FaceBook profile?). But some go to quite an extreme in another direction, using "big data" to pry inside our skulls.

What kind of big data? All your searches on Google, Yahoo, Alta Vista, Bing, or whatever; every click, Twitter text, FaceBook, LinkedIn, blog post, or online chat. We create tons of data about our day-to-day, even moment-by-moment activities. There was recently an item on the noon radio news about a company that aggregates such data and sells "packages" to companies, who pay $1 to $2 million dollars on some periodic basis for it (That's all I remember, I was listening with half an ear while folding laundry). Why is all that data so valuable? Because businesses believe they can better predict which products will sell to what kind of people if they crunch it.

A few months ago a handle on a drawer broke. Naturally, the cabinet is decades old and nothing even remotely similar in style could be found at Home Depot or a decorator's salon. So of course I looked online for something with the right spacing of mounting holes, with an appearance that would be compatible with the cabinet, in a set of four, so the handles would all match. It took a few days. I bought a set I liked, online, and installed them. For the next several months, however, ads about cabinet door handles appeared everywhere I went online: Google, FaceBook, Amazon, eBay. They all knew I'd been looking for door hardware. None of them knew I was done looking! (Google, are you listening? Do, please, close the loop and collect purchase data also.)

What is The Formula? Luke Dormehl calls it an Algorithm. What is an algorithm? To anyone but a mathematician it is a Recipe or a Procedure. I used to have a book, which I used into unusability: How to Keep Your Volkswagen Alive: A Manual of Step-by-Step Procedures for the Compleat Idiot by John Muir and Richard Sealey. With its help I kept my 1966 Bug alive into Moon Unit territory. The "procedures" were recipes, or algorithms, for things like setting valve clearances, changing a wheel bearing, or overhauling an engine. In computer science, an algorithm is the detailed instructions to a computer to direct it what you want it to do, very, very exactly.

Here is the kicker. A traditional algorithm is carried out in a procedural manner (don't pay attention to claims of non-procedural, object-oriented computer language gurus. At the root, a computer CPU carries out a series of procedural instructions), according to a "computer code" or "program", written in one or more formal languages. Some time ago I looked at the internal release notes for the Android OS used in many cell phones. That version, at least, released in 2009, had modules written in 40 computer languages. No matter how complex the program or program system, the instructions are written by a person, or perhaps by many persons, and no matter how many, their knowledge is finite. There are also time constraints, so that the final product will be biased, firstly by the limitations of the programmer(s), secondly by tactical decisions of what to leave out for the sake of time or efficiency, and thirdly by the simplifications or shortcuts this or that programmer might have made so that some operation was easier to write the code for. They may also be biased by inner prejudices of the programmer(s).

Another kicker: A kind of start-stop-start process had been going on around Neural Networks. They try to mimic the way our brains are wired. There are two kinds, hardware and software. Hardware neural nets are difficult to construct and more difficult to change, but they have much greater speed, yielding almost immediate results. Because people who can wire up such hardware are quite rare compared to people who can write computer software, hardware nets are also rare, and nearly all the research being done with them is being done using software simulations. "Machine learning" by neural nets can be carried out by either hard- or software nets, but I'll defer remarks on one significant difference for the moment.

A neural network created for a specific task—letter recognition in handwritten text, for example—is trained by providing two kinds of inputs. One is a series of target images to "view", perhaps in the form of GIF files, or with appropriate wiring, a camera directly attached. The other is the "meaning" that each target image is to have. A training set may have five exemplars of the lower-case "a", along with five indicators meaning "that is an a", five of "b" and their indicators, and so forth. The innards of the net somehow extract and store various characteristics of the training data set. Then it is "shown" an image to identify, and it will produce some kind of output, perhaps the ASCII code for the letter.

The inner workings of neural nets are pretty opaque, and perhaps unknowable without extremely diligent enumeration of all the things happening at every connection inside. But at the root, in a software neural network there is a traditional algorithm that describes the ways that the network connections will interact, which ones will be for taking input or making output, which ones will store things worth "remembering", and so forth. This is one reason that software nets are rather slow, even on pretty fast hardware. The simulation program cannot produce the wholly parallel processing that a hardware net uses (brains use wholly parallel processing, and are hard-put at linear processing, the opposite of computer CPU's). If the net is small, with only a few dozen or a few hundred nodes, the node-by-node computations can be accomplished rapidly, but a net that can recognize faces, for example, has to be a lot bigger than that. It will be hundreds of times slower.

Now for the other significant difference. The computer running the simulation is digital, while a hardware network is analog. I remember the first time I used a computer, that I was quite impressed to see calculations with 7-8 digits of significance, and if I used double precision, 15 digits. That sounds very precise, and for many uses, it is. Fifteen digit precision means one can specify the size of something about the size of a continent to the nearest nanometer. That is about the size of five or 10 atoms. However, a long series of calculations will not maintain such a level of precision. For many practical uses, calculations of much lower precision are sufficient. Before computers came along, buildings and bridges were built, and journeys planned; a slide rule was accurate enough to do the calculations. My best precision using a slide rule was 3-4 digits. But "real life"systems are typically nonlinear, and the sums tend to partly cancel one another out. You might start with very accurate measurements (but it's quite unlikely they are more accurate than 4-6 digits). Run a simulation based upon those figures a few dozen steps, and somewhere along the line there might have been a calculation similar to this:

324.871 659 836 648 - 324.860 521 422 697 → 0.011 138 413 951 016 4

If you've been counting digits, you might notice that the digits 0164 (which I colored red) are superfluous...where did they come from? That is the rounding error, both that which arose from representing the two numbers above in binary format, and that from the conversion of the result back into decimal form for display. But the bigger problem is that, counting only the black digits, only 11 are useful. Four have been lost. Further, if you were to start with decimal numbers that can be represented exactly in binary form, such as 75/64 = 1.171 875 and 43/128 = 0.335 937 5, multiplying them results in 3,225/8,182 = 0.393 676 757 812 5, which has 13 digits of precision, whereas the original numbers had seven each. Thus it typically takes twice as many digits to represent the result of a multiplication, as were needed to represent the two multiplicands.

I could go on longer, but an interested person can find ways to determine error propagation in all kinds of digital systems, many of which have long been studied already. By contrast, an analog system is not limited by rounding errors. Rather, real wires and real electronic components have thermal noise, which can trouble systems that run at temperatures we might find comfortable. Further, Extracting the outputs in numerical form takes delicate equipment, and the more accurately you want those output numbers to be, the more delicate and expensive the equipment gets. However, until readout, the simulation runs with no errors due to subtraction or multiplication, other than gradual amplification of thermal noise.

Suffice it to say, both direct procedural algorithms and neural network machine-learning systems are in use everywhere, trying to predict what the public is going to do, be it buying, voting, dating, relocating, or whatever. That is the main reason for science, after all: predicting the future. Medical science in the form of a doctor (or more than one) looks at a sick person and first tries to find a diagnosis, an evaluation of what the problem is. The next step is a prognosis, a prognostication or prediction; it is the doctors' expectation of the progress of the disease or syndrome, either under one treatment or another, or under none. A chemist trying to determine how to make a new polymer will use knowledge of chemical bonding to predict what a certain mixture of certain chemicals will produce. Then the experiment is carried out to either confirm the expectation (the prediction), or if it does not, to learn what might have gone against expectation and why. The experiments that led to the invention of Nylon took ten years. But based upon them, many other kinds of polymers later proved easier and quicker to develop. It is even so in biological science. Insect or seashell collecting can be a fun hobby, but a scientist will visit a research museum (or several) to learn all the places a certain animal lives, and when various specimens were collected, and then determine if there is a trend such as growing or shrinking population. Is the animal going extinct? Or is it flourishing and increasing its range worldwide?

In the author's view, The Formula represents the algorithms used in the business world, broadly construed, to predict what you might like, and thus present you with advertising to trigger your desire for that thing. My experience with cabinet handles shows that they often get their timing wrong. Many cool and interesting ads showed up, but it was too late. However, that isn't the author's point. The predictive methods find what ads to show us for products, or prospective dating partners on eHarmony or OK Cupid, or those that manage a politician's image, all tend to narrow our choices. A case in point from the analog world: one of the best jobs I had before going into Engineering came about because an Employment Agent, leafing through job sheets, muttered, "You wouldn't be interested in that," but I quickly said, "Try me!"

Try making some Google searches while logged in to Google, and then (perhaps using a different browser, and if you're really into due diligence, on a different computer network such as a library), making the same searches while not logged in. The "hits" in the main column will be similar, or possibly the same. But the ads on the right are tailored to your own search history and other indicators that Google has gathered.

Is all this a bad thing? Maybe. You can game the system a little, but as time goes on, your history will more and more outweigh things you do differently today. Sure, I got a sudden influx of ads about cabinet handles after searching for same, but if I had a history as a very skilled handyman (I don't!), the exact ads I saw might have been quite different. And I might have also seen ads about certain power tools intended to make the mounting of new cabinet handles even easier.

The author has four concerns and spends a chapter on each.

  1. Are algorithms objective? They cannot be. Programmers are not objective, and machine learning is dependent on the training set, which depends on the persons who create it, and they are not objective.
  2. Can an algorithm really predict human relationships? We have proverbs that give us pause, such as, "Opposites attract", and "If you're not near the one you love, you'll love the one you're near".
  3. Can algorithms make the law more fair? I was once asked by a supervisor if I thought he was fair. I replied, "Too much concern for fairness can result in harshness. We (his 'direct reports') wish to be treated not just fairly but well. We'd like a little mercy with our justice." Mr. Dormehl cites the case of an experiment with an inflexible computer program, given the speed records from a car on a long-distance trip. It issued about 500 virtual tickets. A different program, that averaged speed over intervals just a little longer, issued one ticket.
  4. Can an algorithm create art? Since all the programs created to date operate by studying what makes existing artworks more or less popular, they can only copy the past. True creation means doing what has not been done. Picasso and others who developed Cubism did so against great opposition. Now their works sell for millions. It was art even before it was popular, but the "populace" didn't see it that way for a couple decades.

The book closes with a thoughtful section titled "How to Stay Human in the World of the Formula." While he has some suggestions, I think the best way is to avoid being totally predictable. In many ways, that is hard for me, because I am a man of regular habits. I'm quite happy eating the same meat-and-cheese sandwich for lunch day after day, taking the same route to a work place (or these days, a place I volunteer), eating at a certain kind of restaurant and eschewing most "fine dining" places, wearing a certain kind of garb depending on the season, playing (on acoustic instruments, not electronic devices) certain kinds of music to the exclusion of others, and so forth. But I am also the kind of guy, when I make a mobile, it will be quite different from any other I have ever made: different materials, different color schemes, and different numbers of hanging objects clustered—or not—in various ways. I made one out of feathers once; not my most successful mobile. When I write a formal document or a letter for sending via snail mail, though I type it because handwriting is so slow, I usually pick a new typeface in which to print it; I have a collection of nearly 2,000 font files, carefully selected either for readability or as specialized drop caps (I love drop caps, though I am careful in their use). I haven't bothered to try alternate typefaces for this blog, because there are only 7 available anyway, and the default is as good as any.

The author proposes that we "learn more about the world of The Formula". Sure. But as long as Google's Edge Rank (formerly Page Rank) is a black box, and as long as everyone out there from FaceBook and LinkedIn to Amazon and NetFlix keep tweaking their own black box "recommendation engines", it will be a kind of arms race between the cleverest consumers and the marketers. But, hasn't that always been true?

Sunday, October 06, 2013

Looking too hard, and not looking

kw: book reviews, nonfiction, forecasting, prediction, statistics

We are remarkably good at cutting through the clutter in many situations. For example, we can talk to someone at a crowded party and pick out what they are saying in spite of the noise all around; and we can often spot a familiar face in a crowd. However, we sometimes see (or hear, etc.) things that are not there. When I was a child we would look for faces or other shapes in clouds. In a few minutes of looking, something suggestive is bound to appear. And there is a painting by my father of waves breaking on a rocky seashore. One of the big rocks looks like a leopard's head, and once I'd seen it, ever since I always see that leopard's head whenever I glance at the painting.

My father had no intention to hide faces in his paintings. Seeing the leopard's head is an example of a Type 1 error. If my father did actually hide faces in all his paintings, and I have noticed only this one (I have several others), then missing the faces that are there would be Type 2 errors. If I become so rapt in searching clouds for faces that I don't notice a friend approaching until he taps me on the shoulder, I have fallen victim to both kinds of error! We lazy, sedentary Westerners tend to do this frequently. Not so someone living hand-to-mouth in the woods.

For nearly everyone, through all the one or two million years of our evolution as brainy apes, hyper-alertness was required. Where it matters most, a Type 1 error does no harm, but a Type 2 error might be fatal. Running from a rock that looks like a leopard can make you look silly, but not running from a leopard that looks like a rock will probably get you eaten. Strangely, though we have kept our strong propensity to make Type 1 errors, as the risk of not noticing a real leopard has fallen, we are more and more likely to make Type 2 errors. In our modern world, in which we increasingly rely on forecasts and predictions, this leads to trouble.

Nate Silver, in his new book The Signal and the Noise: Why So Many Predictions Fail – But Some Don't, presents a number of similar examples that display our modern tendency to pick faces out of clouds while ignoring the approaching friend (or foe). I'll simplify matters and mention that he finds successful forecasting in only two areas: weather and baseball. Politics and stock picking and a number of other areas come in for a drubbing.

This simple diagram tells me all I need to know about "technical analysis" of stock prices. The data are the day-to-day percent change in the price of DuPont stock, from 1962 to mid September of this year. That's just over 13,000 data points. The X axis is the change on any particular day, and the Y axis is the change on the following day. This diagram shows perfect non-correlation! It is a 2-D bell curve, though with thicker tails than a Gaussian bell curve.

During those 51 years, the stock rose nearly 4,200%. That averages out to 7.7% per year but only 0.032% daily. Someone who bought $1,000 of DD stock in early January 1962 would have $43,000 today. Now, there's been a lot of inflation. That $1,000 in 1962 had the buying power of $7,740 today. So a half-century of waiting produced an effective multiplier of 5.5. That's 3% yearly after adjusting for inflation. Better than the bank.

The most extreme daily jumps are -20% and +10%. Stock speculators, particularly day traders, dream of taking advantage of the many days that a stock's price changes more than a percent or two. And such days are more common than if the distribution were strictly Gaussian. DuPont stock moves up at least 2.5% in a day about 5% of the time, and downward with similar frequency. That means, if you could pick just those up days, about 12 days each year, you could earn at least a 20% return yearly. That's 2-3 times what a buy-and-hold strategy will earn. Then, look at this:


The chart shows the historical record of DuPont stock, adjusted for splits. Focus on late 1974, late 1987, and late 2008 to early 2009. These show DD following the herd during market crashes, and represent downturns of 50%, 41% and 65%, respectively. If you could have avoided them, by selling just at the peak and buying back in at the bottom, your final return would be 9.69 times greater, for a total value of $416,000! Adjusted for inflation, that's over 8% return yearly (12.5% dollar-for-dollar yearly return).

Such figures stoke the dreams of day traders. But the first chart, showing no day-to-day correlation, dashes those dreams. Day traders work very hard for little return, and most lose. Some lose, big time, and some gain, but it is by accident either way. There are millions of day traders and other stock speculators. As Churchill wrote, "Even a fool is right once in a while."

Now we must differentiate prediction from forecasting. A prediction is a flat statement that a specific happening will or will not occur at some time or in some time horizon. For example, "There will be a magnitude 7 earthquake in Fremont within the coming year." A proper forecast includes the forecaster's uncertainty and is stated in probabilistic terms, as, "Projecting the trend of earthquakes in Fremont indicates that an earthquake of magnitude 7 or greater occurs about 3 times every 200 years." [Fremont was the imaginary State in the novel Space by James A. Michener]. One might add to such a forecast, a hybrid statement such as, "Fremont has not experienced an earthquake of magnitude greater than 6 in the past 100 years," which implies that "the big one" may be overdue. But it may indicate that conditions deep down may also be changing.

Earthquake prediction is the poster child of unpredictable phenomena. Intense study and research over decades, even centuries, have failed to yield a single valid prediction. Sports betting is close behind, except in the arena of baseball. Nate Silver once created a system he calls PECOTA, that rates the strength of teams against one another according to the past statistics of their players, and a well-known "aging curve" of the way performance changes over a player's career. Because baseball has such a rich data set, going back a century, and the principles needed to make useful forecasts are also well known, PECOTA and similar systems can evaluate players and teams at a level nearly equal to the best scouts. The computer can't quite replicate the humans, but it does give 'em a run for the money!

Why are forecasting and prediction so hard? Even though we have randomness at the deepest level of atomic phenomena, that randomness is constrained by the statistics of large numbers, and physics works very accurately to predict many systems, such as planetary orbits. Thus, though the path of an electron after passing through a hole may be uncertain, the distribution center of the paths of trillions of electrons (say, a millionth of an ampere for 0.1 second or so) will be very sharply defined and can be accurately measured, and the shape of the distribution tells you additional facts: the hole's size and shape. The much larger "distribution" consisting of the atoms making up a baseball mean that its flight, once thrown or batted, will be easily predicted.

The geological setting of an earthquake is not as simple as an electron. Perhaps this year, an earthquake might occur, large enough that the two sides of a fault will slip by each other by half a meter. That may be enough to put two kinds of rock in contact, that were not in contact before, which changes the likelihood of the next earthquake.

What about the weather? Air is in constant motion; its humidity and temperature, and thus its density, change constantly. How can anyone make a useful weather forecast? In some ways, we are still dependent on the "signs in the sky" that Jesus mentioned. In modern (18th Century) terms, "Red sky at morning, sailor take warning. Red sky at night, sailor's delight." Lore such as this is a compilation of patterns that happen over and over, so that generations of our ancestors took note and remembered. Yet now we can get a forecast up to a week or two ahead, complete with expected high and low, precipitation chances and intensity, and wind strength.

It's all done in a computer. Air may have complex behavior, but the physics of air motion and how it changes with temperature, pressure and humidity are well known. The 3D-gridded-cell models that run in supercomputers use surprisingly simple physics to determine how a 3D cell is influenced by the 6 cells it is in facial contact with, and the 8 cells at its corners. The reason supercomputers are used is that Earth is big. The surface area of the planet is 4πr², where r is 6,370 km: about 510 million km². Cells of half a km on a side, plus 0.1 km in depth (up to 12 km altitude) result in a Global Circulation Model (you'll see the acronym GCM in some weather web sites) with 1/4 trillion cells. It takes a lot of calculation to determine what will happen in the next quarter hour. There are 96 quarter hours in a day, and 672 in a week. To do all those trillions and quadrillions of calculations in only an hour or two requires today's largest computers. And the forecasters' computer gurus don't do it once, they run it several times with very small variations (the formal practice of selecting the variations is called Design of Experiments), to test the stability and sensitivity of the forecast to perturbations.

Weather forecasters have an incentive to get it right that others don't have. The reality is going to arrive tomorrow or the next day, it is visible to all, and it is no fun getting a call such as, "I have ten inches of 'partly cloudy' that I need to shovel off my driveway. Want to come over and help?" They also get a ton of research money from the Dept. of Defense, because good forecasts are crucial to military activities. Earth dynamic studies are different. Students of earthquakes can't observe the day-to-day conditions of a fault line. Its active zone is typically 8-15 km deep, and we can't yet drill a well that deep. Earthquakes are also rare. Sure, there are thousands of little ones, at the bottom of "measurable", every day, but there are trillions of weather events around the globe, every few minutes.

Mr. Silver entertains us with many, many stories of the vagaries of forecasts of all types. In the end, most phenomena are too difficult to forecast appropriately. Some involve living things. The cardinal rule of animal studies is, "Given any particular set of temperature, lighting, food availability and ambient noise, the rat will do whatever the rat wants to do." And this is in spite of lab rats being so inbred that their genetics are practically identical. The statistics of playing poker yield a few big winners, who work hard for the kind of edge they need to beat their fellow experts. But they love to be in a game that is well supplied with "fish": overconfident amateurs. A well-written computer package might tell a poker player the optimum betting strategy, but only if it is betting against other computers. The social aspects of the game, bluffing and speed or slowness of a bet for example, often provide a lot more of an edge than the math does. Carefully crafted intimidation works wonders. I don't expect a computer to master these aspects of the game for a number of decades (that's my forecast!).

The book's final example is the climate, particularly "global warming" or "climate change" or "greenhouse effect" or whatever the next buzzword will be. Climate is not weather. It is the setting in which weather happens. Climate changes unfold over multiple decades or centuries or millennia. Weather changes take seconds. In numerical analysis, this is the Stiffness problem. When something changes suddenly, it takes time for the effects to either move elsewhere or to die down. If you are interested in something with a 5-year cycle, such as El Niño (also called ENSO), the exact location and timing of today's sudden thundershower will not matter one tiny bit. If your interest is in human-induced greenhouse warming that began in the late 1700s, ENSO is an irritation at best. In fact, weather and medium-scale cycles such as ENSO are "noise" in the context of this book's thesis. Another researcher, later on, made clear a different view, that noise is really signals, but about stuff you aren't interested in at the moment.

This is like the crystal radio I made as a kid. It initially consisted of a long wire, running to a treetop, a piece of germanium crystal, and a "whisker", a wire that formed a diode with the germanium; and earphones attached to the whisker and the ground connection on the back of the germanium crystal. The diode "detected" the audio signal by separating it out of the radio frequency "hash". There was just one strong station nearby, so I could hear them pretty clearly. But later, as more stations came on the air (this was the 1950s), I could hear all of them at once. So, following a diagram in Mechanix Illustrated, I made a coil and paid a dime for a small capacitor and a piece of copper, to make a rough tuner. It could be tuned to resonate with one AM station at a time, so I could "tune out" the "noise" of the other stations. They were actually signals, just signals I didn't want right then.

The global greenhouse has warmed about 0.5°C (0.9°F) in a century, and perhaps 1°C (1.8°F) since 1750. Some of that may be warming since the Little Ice Age, which some consider a regional phenomenon, not a global one. But the current "ForecastFox for Mozilla" forecast for the next 24 hours indicates we'll have a 20°F swing tomorrow, from 75 in midafternoon to 55 overnight. You have to average out a lot of daily temperatures to see a change of a degree over 250 years. When you want weather, that is your signal. When you want climate, weather is noise, and lots of it.

The science of greenhouse warming is partly very well known, and partly not so well known. I learned to replicate the Arrhenius calculations from 150 years ago, when I was a pre-teen. Actual warming since his day has been about twice what he expected, because there seem to be amplifying factors. These are very poorly known. Does more cloud cover cool the atmosphere by reflecting more sunlight, or warm it by acting as a further thermal blanket? Or does it do one thing at a certain latitude and another elsewhere? If we do have a further warming by 2 to 4°C, will it shift the Hadley Cell north, or south, or not at all? (The northern edge of the Hadley Cell is a range of latitudes characterized by dry, descending air that form all the world's great deserts.) I've thought of buying land in central Canada, that is currently too cold to farm. Perhaps in 20 years it will be arable…unless the Hadley Cell shifts north and dries out Canada. Then maybe the Mojave would become a tropical paradise!

Y'know how to make a complex system into a positively unsolvable mess? Make it political. Both sides of the Climate debate are so politicized that they can only talk past each other. The tiniest proposal to set any policy is vigorously fought by every vested interest, even those who might benefit (the devil you know…). Heaven help us if weather forecasting ever gets politicized! It is already true that most forecasters err on the wet side: a 20% chance of rain is reported as a 40% or even 50% chance, because the ones rained on are less likely to complain, and those that aren't will feel they dodged a bullet. What if some "weather outcomes" become more politically correct than others?

By the way, I take issue with Silver's definition of statistical rain forecasts. He writes that if 40% of the computer models indicate rain in Chicago, and the rest don't, it is reported as a 40% chance of rain. Sounds logical, but it is quite different than that. The "chance of rain" has different meanings in spring (plus summer) and autumn (plus winter). Spring and summer squall lines pass through areas that are well predicted by most GCM programs. But a squall line is not a solid front of rain. It is a line of thunderstorms. A light squall line may have storms half a mile wide, spaced 2-3 miles apart, giving 20% of the area a 100% chance of rain. The forecasters just don't know which 20%, so the whole area is given a 20% chance of rain. A heavy squall line will have larger storms with closer spacing, and maxes out at about 80% coverage (though this will probably be reported as "near certain"). Fall and early winter storms tend to be solid and widespread, but subject to ripples several miles wide in the upper atmosphere. As a system rides up a ripple, it drops rain along a solid band dozens of hundreds of miles long but only about a mile wide or so. As it rides down, it dries out. The height of the ripples determines whether the overall chance of rain is 30% or 70% or somewhere between. The ripples drift along as system after system rides through, so it is very hard to tell exactly where the rain will fall. Timing is everything. Then, a lower-level storm that just dumps (ignoring the ripples) leads to those 100% forecasts, which are generally accurate.

In most arenas, Silver advocates using Bayesian analysis rather than "frequentist" simulations or estimations. These allow individualized forecasts for particular cases. An example is the probability of breast cancer in a woman in her 40s, who has just had the unwelcome news that a mammogram is "positive". The factors of a Bayesian calculation are:
  • x - Prior Estimate: the chance that a proposition is true.
  • y - Type 1 analysis: the chance that new data which indicates "Yes" is actually correct.
  • z - Type 2 analysis: the chance that the proposition is not true, in spite of the new data.
The data are usually noted as percents. The formula for a new estimate (a new x) is xy/(xy+z(1-x)). For this example, we find:


In this case, the woman may wish for a needle biopsy, but a bit of blood chemistry may be in order first. Enzymes in the blood can indicate whether a new cancer is likely to be slow growing, or faster. Is it slower (the most likely case)? She can wait a year for another mammogram. If the next mammogram is positive, re-do the analysis, replacing the 1.4% with 9.6%. Now the "new x" is just over 44%, and at the very least a biopsy is indicated. Most other forecasting methods don't use multi-step refinement. And by the way, if the next mammogram is negative (and no palpation can detect a lump, or any growth in an earlier lump), running the analysis with 9.6%, 10% and 75%, in that order, reverts to 1.4% as the "new x".

Those who follow this blog may wonder why it took me 3 weeks to read such a fascinating book. The writing is good and the examples are interesting, so that didn't slow me down. We have a lot going on, however, so I have had much less time for reading than usual. Retirement has been good to me so far, but I have to be careful not to take on too many projects at once. I completed a Real Estate course and passed the test in July. However, I will probably not seek a license or become a Realtor®, because there are simply too many other things I'd prefer to do. The change of style and reduced frequency with which I post is a similar effect. I used to post almost every lunch hour, doing research in off hours. I think I am working longer days than when I worked! Better busy than bored. Since retiring in February, I have put 24 items in my "job jar" file. Half of them, mostly the bigger ones, have been completed. One major item is awaiting an event that is at least a year in the future, but the preparations are nearly all completed. Others are smaller so I can take an odd half day to perform one. All things in their own time. In the meantime, I read when I can, and report what I read.